# How to slow down large amount of data coming from filebeat?

**URL:** <https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970>\
**Category:** Logstash\
**Created:** [March 25, 2020, 10:09am UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970 "2020-03-25T10:09:51Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sagar\_Mandal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sagar_mandal/32/57073_2.png) [@Sagar\_Mandal](https://discuss.elastic.co/u/Sagar_Mandal)\
**Post date:** [March 25, 2020, 10:09am UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970/1 "2020-03-25T10:09:51Z")

</div>

Hi Team,

How will you slow down large amount of data streaming from filebeat to logstash so that it can be processed accurately in filter section.

Thanks and Regards,  
Sagar Mandal

---

<div class="post-metadata">

**Author:** ![A\_B](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/a_b/32/17104_2.png) [@A\_B](https://discuss.elastic.co/u/A_B)\
**Post date:** [March 25, 2020, 4:43pm UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970/2 "2020-03-25T16:43:50Z")

</div>

Hi @Sagar_Mandal,

that should be mostly automatic. Of course depends on how much data would have to be cached...

Can't find this in any official Elastic documentation but as far as I remember, it is part of the lumberjack protocol that is used between Filebeat and Logstash.

From [Send Your Data | Logz.io Docs](https://logz.io/blog/beats-tutorial/)

> One of the facts that make Filebeat so efficient is the way it handles _backpressure—_ so if Logstash is busy, Filebeat slows down it’s read rate and picks up the beat once the slowdown is over.

But from personal experience, if the Logstash filter section is very process heavy, you can still get into trouble. I have killed my Logstash instances with sub-optimal filters, especially GROK filters with poorly written patterns and no anchoring.

---

<div class="post-metadata">

**Author:** ![Sagar\_Mandal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sagar_mandal/32/57073_2.png) [@Sagar\_Mandal](https://discuss.elastic.co/u/Sagar_Mandal)\
**Post date:** [March 25, 2020, 5:24pm UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970/3 "2020-03-25T17:24:21Z")

</div>

okay so the thing is about 100GB of data comes everysingle day to logstash from filebeat and then it goes to a filter section where a lot of conditioning is being done so....yeah.

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [March 25, 2020, 6:38pm UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970/4 "2020-03-25T18:38:21Z")

</div>

100GB per day is a lot of data. I would recommend something like:

filebeat --\> kafka --\> multiple logstash instances --\> elasticsearch

Bursts of messages can then be queued in Kafka and multiple Logstash instance can be used to scale the post-processing of that data.

Rob

[![GitHub](https://us1.discourse-cdn.com/elastic/original/3X/6/f/6f8ae834f16b1a02d31607317669716807844d84.png)](https://github.com/robcowart) [![YouTube](https://us1.discourse-cdn.com/elastic/original/3X/4/3/43b9b81a8c93786219985aeb5335c1c323449053.png)](https://www.youtube.com/channel/UCivWvTx1DwrWNcDLV58kmOg) [![LinkedIn](https://us1.discourse-cdn.com/elastic/original/3X/6/7/674f3370d0f0542ddc5e408516beb1b7edd6c1bf.png)](https://www.linkedin.com/in/robertcowart/)  
**[How to install Elasticsearch & Kibana on Ubuntu - incl. hardware recommendations](https://www.youtube.com/watch?v=gZb7HpVOges)**  
**[What is the best storage technology for Elasticsearch?](https://www.youtube.com/watch?v=nKUpfJCBiS4)**

---

<div class="post-metadata">

**Author:** ![A\_B](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/a_b/32/17104_2.png) [@A\_B](https://discuss.elastic.co/u/A_B)\
**Post date:** [March 26, 2020, 7:20am UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970/5 "2020-03-26T07:20:19Z")

</div>

We are doing about 300GB of logs (about 400M documents) for 4 x Logstash with 12 CPU cores each. To be fair, load is \< 1 at the moment. I did spend a lot of time optimizing our Logstash filters.

We are working on adding Kafka to the mix, not so much to deal with spikes but to be able to queue all messages during maintenance or if for some reason Logstash or Elasticsearch breaks completely.

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [March 26, 2020, 9:37am UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970/6 "2020-03-26T09:37:25Z")

</div>

@A_B as you make the move to Kafka, a few things that will really boost throughput...

1. increase `pipeline.batch.size` from the default of 125 to at least 1024 (1280 was best in my environment)
2. increase `pipeline.batch.delay` from the default of 50 to at least 500 (1000 was best in my environment)
3. in the `kafka` input, set `max_poll_records` to the same value as `pipeline.batch.size`
4. each thread defined by `consumer_threads` in the `kafka` input will be an instance of a consumer. So if you have 4 instances with 2 threads, that is 8 consumer instances. Your Kafka topics must have at least 8 partitions for all consumer threads to ingest data. You will want more partitions than your current needs so you can easily scale in the future.
5. the number of `pipeline.workers` should be at least equal to `consumer_threads`.
6. the kafka output should set `batch_size` to at least 16384

You may end up tweaking some of the buffer settings as well, but the above will give you a good starting point.

Rob

[![GitHub](https://us1.discourse-cdn.com/elastic/original/3X/6/f/6f8ae834f16b1a02d31607317669716807844d84.png)](https://github.com/robcowart) [![YouTube](https://us1.discourse-cdn.com/elastic/original/3X/4/3/43b9b81a8c93786219985aeb5335c1c323449053.png)](https://www.youtube.com/channel/UCivWvTx1DwrWNcDLV58kmOg) [![LinkedIn](https://us1.discourse-cdn.com/elastic/original/3X/6/7/674f3370d0f0542ddc5e408516beb1b7edd6c1bf.png)](https://www.linkedin.com/in/robertcowart/)  
**[How to install Elasticsearch & Kibana on Ubuntu - incl. hardware recommendations](https://www.youtube.com/watch?v=gZb7HpVOges)**  
**[What is the best storage technology for Elasticsearch?](https://www.youtube.com/watch?v=nKUpfJCBiS4)**

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 23, 2020, 9:37am UTC](https://discuss.elastic.co/t/how-to-slow-down-large-amount-of-data-coming-from-filebeat/224970/7 "2020-04-23T09:37:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
