# Recommendations for parsing 1000's ~10MB files to backfill elasticsearch

**URL:** <https://discuss.elastic.co/t/recommendations-for-parsing-1000s-10mb-files-to-backfill-elasticsearch/174114>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [March 27, 2019, 11:54am UTC](https://discuss.elastic.co/t/recommendations-for-parsing-1000s-10mb-files-to-backfill-elasticsearch/174114 "2019-03-27T11:54:18Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![drwg](https://avatars.discourse-cdn.com/v4/letter/d/ecccb3/32.png) [@drwg](https://discuss.elastic.co/u/drwg)\
**Post date:** [March 27, 2019, 11:54am UTC](https://discuss.elastic.co/t/recommendations-for-parsing-1000s-10mb-files-to-backfill-elasticsearch/174114/1 "2019-03-27T11:54:19Z")

</div>

I'm currently using filebeat -\> logstash -\> elastic to back fill elastic search with exit codes from several thousand text output files each 10's MB in size and \>10k lines. ~200GB in total. The server can push 5GB/s read/write so that's not a bottleneck.

I'm configuring filebeat to search for specific keyworks which I am specifying in the "include\_lines: ['keyword1', .......,'keywordN'], where the number of keywords could be as high as 20 but at present my problems (performance) are showing with a single keyword. I also wish to use exclude\_lines at some point.

Performance is incredibly slow, any recommendations for improving the performance?

I suspect parsing the files is one aspect, how exacly does the include/exclude\_lines work?

But also the number of harvesters which is started in parallel possibly? - I have tried limiting the number of harvesters by setting the `harvester_limit`to the number of cores.

filebeat.inputs:

- type: log  
enabled: true  
paths:
  - /pathto/symlinksdir/\*  
symlinks: true  
tags: ["some\_value"]  
fields: {log\_type: "some\_value2"}  
include\_lines: ['keyword']

Thanks

---

<div class="post-metadata">

**Author:** ![ruflin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ruflin/32/3116_2.png) [@ruflin](https://discuss.elastic.co/u/ruflin)\
**Post date:** [March 29, 2019, 3:46pm UTC](https://discuss.elastic.co/t/recommendations-for-parsing-1000s-10mb-files-to-backfill-elasticsearch/174114/2 "2019-03-29T15:46:03Z")

</div>

I'm a bit suprised that you hit a bottleneck here on the Filebeat side. Could you share a bit more on how your `include_lines` statement looks like? Regexp?

What results do you see if you just ship everything? Much higher throughput?

---

<div class="post-metadata">

**Author:** ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)\
**Post date:** [March 29, 2019, 3:53pm UTC](https://discuss.elastic.co/t/recommendations-for-parsing-1000s-10mb-files-to-backfill-elasticsearch/174114/3 "2019-03-29T15:53:53Z")

</div>

What is the ratio between lines published and lines filtered out. The registry file keeps track of the file offset, but needs some IO to be written. If the ratio is somewhat 'bad', then the registry writes will slow down filebeat, as it also requires some fsync when writing the registry. Setting `filebeat.registry_flush: 1s` helps in this case (See [registry\_flush docs](https://www.elastic.co/guide/en/beats/filebeat/current/configuration-general-options.html#_literal_registry_flush_literal)).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 26, 2019, 4:01pm UTC](https://discuss.elastic.co/t/recommendations-for-parsing-1000s-10mb-files-to-backfill-elasticsearch/174114/4 "2019-04-26T16:01:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
