# Filebeat include\_lines performance v.s. grep

**URL:** <https://discuss.elastic.co/t/filebeat-include-lines-performance-v-s-grep/151751>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [October 10, 2018, 5:59am UTC](https://discuss.elastic.co/t/filebeat-include-lines-performance-v-s-grep/151751 "2018-10-10T05:59:02Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![dontscrambleme](https://avatars.discourse-cdn.com/v4/letter/d/a4c791/32.png) [@dontscrambleme](https://discuss.elastic.co/u/dontscrambleme)\
**Post date:** [October 10, 2018, 5:59am UTC](https://discuss.elastic.co/t/filebeat-include-lines-performance-v-s-grep/151751/1 "2018-10-10T05:59:02Z")

</div>

I use filebeat to harvest lines including keywords and send it to logstash for post processing.  
But the time filebeat searching for the string is much longer than running grep in Ubuntu shell  
I don't have number to show up but I can definitely 'feel' it  
Did the beat team compare the filebeat inlcude\_lines performance v.s. grep?

Here is my environment -

- filebeat 6.4.1 in my Ubuntu docker container
- ELK and Ubuntu filebeat are under the same network (created through docker-compose) running in the PC
- each of message files size is around 1.1MB, around 12000 lines in it

filebeat configuration -  
filebeat.inputs:

- type: log  
enabled: true  
paths:

output.logstash:  
hosts: ["logstash:5044"]  
index: 'filebeat\_sit73'

path.data: /filebeat/data

---

<div class="post-metadata">

**Author:** ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)\
**Post date:** [October 12, 2018, 5:36pm UTC](https://discuss.elastic.co/t/filebeat-include-lines-performance-v-s-grep/151751/2 "2018-10-12T17:36:28Z")

</div>

`inculde_lines` is a full regular expression, while `grep` by default does not use a regular expression.

In go the regex engine is not the most efficient. This is a known issue. Plus using a regular expression to match for a sub-string (+ back-tracking by the regex engine) is effectively similar to the naive string matching algorithm (which has about quadratic time complexity).

Assuming you use GNU grep, you will find it using quite some tricks and a much more efficient string matching algorithm. See here: [https://lists.freebsd.org/pipermail/freebsd-current/2010-August/019310.html](https://lists.freebsd.org/pipermail/freebsd-current/2010-August/019310.html)

Beats also need to create events, plus the input is copied into an internal buffer for dealing with potential file truncation. Lines are turned into events and only the final event will be matched. Such that it works correctly with

In Beats we do some analysis and optimisation of matchers if applicable. In this case the inefficient regex of a plain string should be turned into a more efficient sub-string match using different algorithms depending on the size of the input pattern.

Also keep in mind, grep only has to print found lines to the console. Filebeat creates and buffers events into batches. Only if the buffer is full or after some timeout will the buffer be published and events be forwarded to the outputs. Then it does some additional processing, encodes events with meta data to JSON, forwards them to Logstash. Logstash and Elasticsearch can add additional back-pressure, slowing down reading in filebeat. Once queues are full in filebeat due to back-pressure, it will stop processing any more input until Logstash/Elasticsearch have finished processing.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 9, 2018, 5:36pm UTC](https://discuss.elastic.co/t/filebeat-include-lines-performance-v-s-grep/151751/3 "2018-11-09T17:36:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
