# How does logstash use pipeline.batch.size to execute pipeline?

**URL:** <https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772>\
**Category:** Logstash\
**Created:** [January 9, 2023, 9:28pm UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772 "2023-01-09T21:28:29Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![rickfish](https://avatars.discourse-cdn.com/v4/letter/r/e480ec/32.png) [@rickfish](https://discuss.elastic.co/u/rickfish)\
**Post date:** [January 9, 2023, 9:28pm UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/1 "2023-01-09T21:28:29Z")

</div>

I am looking for documentation on how Logstash processes a pipeline when pipeline.batch.size is more than 1.

This is my assumption:

1. The input plugin generates several events, not dictated by batch size
2. While there are events left, Logstash takes a batch at a time, and for each batch, gets events and calls filter() for each
3. if the output plugin implements multi\_receive, then the output plugin is called once with the whole filtered batch.
4. If the output plugin implements receive(), then the output plugin is called for each filtered batch event, one at a time.

Is this correct? Is there documentation that discusses this?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 9, 2023, 9:55pm UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/2 "2023-01-09T21:55:20Z")

</div>

That matches my understanding. I do not think it is documented anywhere.

I believe an entire batch of events moves from plugin to plugin within the filter section. In all the years I have used logstash I have only been bitten by that once.

---

<div class="post-metadata">

**Author:** ![rickfish](https://avatars.discourse-cdn.com/v4/letter/r/e480ec/32.png) [@rickfish](https://discuss.elastic.co/u/rickfish)\
**Post date:** [January 9, 2023, 10:02pm UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/3 "2023-01-09T22:02:15Z")

</div>

Thanks @Badger, I didn't even think about the possibility that the batch would be passed to each plugin within the filter section. I was envisioning each event going through all of the filter plugins before the next event would be processed. I appreciate the clarification.

---

<div class="post-metadata">

**Author:** ![Rios](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rios/32/95745_2.png) [@Rios](https://discuss.elastic.co/u/Rios)\
**Post date:** [January 10, 2023, 12:09am UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/4 "2023-01-10T00:09:57Z")

</div>

There is some documentation.

> Each input stage in the Logstash pipeline runs in its own thread. Inputs write events to a central queue that is either in memory (default) or on disk. Each pipeline worker thread takes a batch of events off this queue, runs the batch of events through the configured filters, and then runs the filtered events through any outputs.

If I am correct, 125 events (default) will be push to the filter, processed all 125 as a transaction and push to output as bulk. If there is 8 workers, there will be processed 1000 events at the same time inside filter+output.

> - The `pipeline.batch.size` setting defines the maximum number of events an individual worker thread collects before attempting to execute filters and outputs. Larger batch sizes are generally more efficient, but increase memory overhead. Some hardware configurations require you to increase JVM heap space in the `jvm.options` config file to avoid performance degradation. (See [Logstash Configuration Files](https://www.elastic.co/guide/en/logstash/current/config-setting-files.html) for more info.) Values in excess of the optimum range cause performance degradation due to frequent garbage collection or JVM crashes related to out-of-memory exceptions. Output plugins can process each batch as a logical unit. The Elasticsearch output, for example, issues [bulk requests](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-bulk.html) for each batch received. Tuning the `pipeline.batch.size` setting adjusts the size of bulk requests sent to Elasticsearch.

[Here](https://www.elastic.co/blog/logstash-persistent-queue) is an interesting pic. Queues are between input and filter, and inside filter there is the most likely a loop for processing, 125 received events.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 10, 2023, 4:23am UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/5 "2023-01-10T04:23:37Z")

</div>

> [@Rios](#):
>
> If I am correct, 125 events (default) will be push to the filter, processed all 125 as a transaction and push to output as bulk.

Let me expand a little on the "push to output as bulk" bit.

The base output plugin is [here](https://github.com/elastic/logstash/blob/main/logstash-core/lib/logstash/outputs/base.rb). The code suggests that the plugin must define a [receive](https://github.com/elastic/logstash/blob/e4dc82a9b319625f79582730e173c8736d5a1ba6/logstash-core/lib/logstash/outputs/base.rb#L95) method, but I suspect that that is not true. I think the receive\_multi method of the plugin is called in preference to the receive method, so that exception never gets thrown if the plugin implements receive\_multi\_encoded or a receive\_multi that actually processes the batch.

The base receive\_multi will send the batch to receive\_multi\_encoded if it exists, otherwise it will loop over the batch calling receive for each event (and throwing that exception if the plugin does not implement it). Note that a plugin can override the base receive\_multi if it wants access to the batch without encoding (and not that long ago _with encoding_ was not even an option).

Some examples ... the http output defines multi\_receive, because it wants to send the batch to elasticsearch in a single \_bulk request if it can. The udp output just sends off each event in a UDP packet, so it defines receive. The s3 output defines [receive\_multi\_encoded](https://github.com/logstash-plugins/logstash-output-s3/blob/f893dae88ee60c2edb67608ab6a0f42f34a17235/lib/logstash/outputs/s3.rb#L243).

If I understand it correctly (which I may well not) then if the output plugin has a receive\_multi\_encoded method then the base output plugin will call the codec to encode each event before passing it to the output plugin. Looking at the PRs on github this appears to be a baby step towards the synchronization required for multi-threaded outputs. Maybe.

The old logstash 6.6 how-to on writing an output plugin no longer exists (which suggests that it is no longer accurate), but it can be found in the Wayback Machine [here](https://web.archive.org/web/20190324203423/https://www.elastic.co/guide/en/logstash/current/_how_to_write_a_logstash_output_plugin.html). That suggests that only multi\_receive is required.

---

<div class="post-metadata">

**Author:** ![rickfish](https://avatars.discourse-cdn.com/v4/letter/r/e480ec/32.png) [@rickfish](https://discuss.elastic.co/u/rickfish)\
**Post date:** [January 10, 2023, 10:54am UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/6 "2023-01-10T10:54:14Z")

</div>

Wow! Thanks to both of you so much for this great description. @Rios , would tou mind posting the URL where that doc is? I am having a tough time finding this type of info.

---

<div class="post-metadata">

**Author:** ![Rios](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rios/32/95745_2.png) [@Rios](https://discuss.elastic.co/u/Rios)\
**Post date:** [January 10, 2023, 1:36pm UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/7 "2023-01-10T13:36:05Z")

</div>

[How Logstash Works](https://www.elastic.co/guide/en/logstash/8.5/pipeline.html)  
[Tuning and Profiling Logstash Performance](https://github.com/elastic/logstash/edit/8.5/docs/static/performance-checklist.asciidoc)  
[Logstash Persistent Queue](https://www.elastic.co/guide/en/logstash/current/tuning-logstash.html)

---

<div class="post-metadata">

**Author:** ![rickfish](https://avatars.discourse-cdn.com/v4/letter/r/e480ec/32.png) [@rickfish](https://discuss.elastic.co/u/rickfish)\
**Post date:** [January 10, 2023, 2:00pm UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/8 "2023-01-10T14:00:03Z")

</div>

Thank you. Much appreciated.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 7, 2023, 2:00pm UTC](https://discuss.elastic.co/t/how-does-logstash-use-pipeline-batch-size-to-execute-pipeline/322772/9 "2023-02-07T14:00:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
