# Logstash output load balancing and workers settings

**URL:** https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176
**Category:** Beats
**Created:** [April 22, 2016, 1:48pm UTC](https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176 "2016-04-22T13:48:52Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![Bruno\_Lavoie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_lavoie/32/8408_2.png) [@Bruno\_Lavoie](https://discuss.elastic.co/u/Bruno_Lavoie)
#### Post date: [April 22, 2016, 1:48pm UTC](https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176/1 "2016-04-22T13:48:53Z")

</div>

Hello,

I have a filebeat agent that will be sending quite a lot of access log data and our current setup have 4 receiving logstash hosts. I would like to benefit from these to gain a maximum throughput.

When I read the «workers» setting description [here](https://www.elastic.co/guide/en/beats/filebeat/current/logstash-output.html#_worker_2), I'm a little puzzled:

> The number of workers «per configured host» publishing events to Logstash. This is best used with load balancing mode enabled. Example: If you have 2 hosts and 3 workers, in total 6 workers are started (3 for each host).

The first part : number of workers **per configured hosts** publishing events to Logstash.  
Which host does the «per configured hosts» means?

- The host running the agent?
- The hosts in the provided destination list?

Sorry for this one, maybe it's already stated clearly, but english is not my native...

So, if I have 4 target hosts and want to benefit from this, should I specify 4 workers?

Also, I saw [here](https://discuss.elastic.co/t/filebeat-client-failed-to-connect/47745/5) that logstash output [bulk\_max\_size](https://www.elastic.co/guide/en/beats/filebeat/current/logstash-output.html#_bulk_max_size_2) should be sized wisely with the global [spool\_size](https://www.elastic.co/guide/en/beats/filebeat/current/configuration-filebeat-options.html#_spool_size) ?

To recap my understandings, to benefit from load balancing to a maximum I should have a config key performance settings like this:

```
filebeat:
  spool_size: 2048 # Default 2048
  
output:
  logstash:    
    hosts: ["logs1.domain.com", "logs2.domain.com", "logs3.domain.com", "logs4.domain.com"]
    loadbalance: true
    worker: 4
    bulk_max_size: 512 # Default 2048

```

Very thanks  
Bruno Lavoie

---

<div class="post-metadata">

### Author: ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)
#### Post date: [April 25, 2016, 11:11am UTC](https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176/2 "2016-04-25T11:11:34Z")

</div>

Hi,

best advice is measure, measure, measure. [This post](https://discuss.elastic.co/t/filebeat-sending-data-to-logstash-seems-too-slow/37596/13) contains a python script + instructions how you can get throughput info right from filebeat. For testing have a log file prepared (I often use [NASA HTTP logs](http://ita.ee.lbl.gov/html/contrib/NASA-HTTP.html)) and delete registry files between runs. (optionsl) In addition use the null (or stdout with `dotted` codec and `pv` tool) output plugin in logstash to not generate any back-pressure from logstash.

The `worker: ...` config is really per host. The default value is 1. That is if `H=# of hosts` and `W=# worker`, then `H*W` workers doing output will be spawned. In your sample config it means you're spawning like 16 workers pushing data in total.

There are 2 options to try in your case:

1. set `filebeat.publish_async: true`. This will push batches as soon as batches are ready into the publisher pipeline. In this case I'd set `spool_size between [bulk_max_size, bulk_max_size * worker * (# of hosts)]`. If one logstash instance is not responding (or slowing down), filebeat continues publishing events using the other workers (due to async/pipelines publishing).

2. set `filebeat.spool_size = output.logstash.bulk_max * output.logstash.worker * len(output.logstash.hosts)`. This will split bathces into `output.logstash.worker * len(output.logstash.hosts)` batches when publishing, so every host gets it's share. Drawback is, if one logstash instance slows down it first takes some timeout to detect it's not responding and transmitting the sub-batch via another logstash instance, basically blocking output until all sub-batches have been ACKed.

My gut feeling tells me option 1 would have better throughput (despite gut feelings being often wrong), but in the end it's up to you to run experiments and measure your setup to figure some good configurations matching your requirements. And don't forget, the higher your throughput, the more resources will be required by filebeat and logstash to process your data.

---

<div class="post-metadata">

### Author: ![Bruno\_Lavoie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_lavoie/32/8408_2.png) [@Bruno\_Lavoie](https://discuss.elastic.co/u/Bruno_Lavoie)
#### Post date: [April 25, 2016, 12:38pm UTC](https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176/3 "2016-04-25T12:38:04Z")

</div>

Thanks a lot for this complete and clear answer...  
Should the official doc be more clear on these settings?

---

<div class="post-metadata">

### Author: ![Bruno\_Lavoie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_lavoie/32/8408_2.png) [@Bruno\_Lavoie](https://discuss.elastic.co/u/Bruno_Lavoie)
#### Post date: [April 25, 2016, 12:42pm UTC](https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176/4 "2016-04-25T12:42:37Z")

</div>

> [@steffens](#):
>
> 1. set filebeat.pool\_async: true.

did you mean publish\_async ?

---

<div class="post-metadata">

### Author: ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)
#### Post date: [April 25, 2016, 1:03pm UTC](https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176/5 "2016-04-25T13:03:39Z")

</div>

Thanks, I fixed my post. (Also fixed spool\_size in option 1)

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 5, 2017, 9:52pm UTC](https://discuss.elastic.co/t/logstash-output-load-balancing-and-workers-settings/48176/6 "2017-07-05T21:52:49Z")

</div>


