# Filebeat insert logs from multiple hosts in parallel

**URL:** <https://discuss.elastic.co/t/filebeat-insert-logs-from-multiple-hosts-in-parallel/219952>\
**Category:** Elasticsearch\
**Created:** [February 19, 2020, 11:35am UTC](https://discuss.elastic.co/t/filebeat-insert-logs-from-multiple-hosts-in-parallel/219952 "2020-02-19T11:35:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ice3man543](https://avatars.discourse-cdn.com/v4/letter/i/dfb087/32.png) [@ice3man543](https://discuss.elastic.co/u/ice3man543)\
**Post date:** [February 19, 2020, 11:35am UTC](https://discuss.elastic.co/t/filebeat-insert-logs-from-multiple-hosts-in-parallel/219952/1 "2020-02-19T11:35:52Z")

</div>

We use elasticsearch for storing logs in our infrastructure. The entire infrastructure is composed of multiple workers that write the logs to a central elasticsearch server. The number of workers can increase as the rate increases of amount of data. I was wondering what would be an optimal configuration for writing data to elasticsearch with filebeats? Should each worker get it's own log writer that uses filebeats or should i use some third party mechanism to ship the logs to somewhere from where they could be indexed to elastic?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [February 19, 2020, 2:09pm UTC](https://discuss.elastic.co/t/filebeat-insert-logs-from-multiple-hosts-in-parallel/219952/2 "2020-02-19T14:09:15Z")

</div>

Having a single filebeat instance on each node, that reads all the different logfiles sounds like a good plan to me.

I do not fully understand the difference between worker and writer and their relationship to a filebeat (or a node). Maybe you can explain.

From an architectural perspective, there are a couple of ways you can go. Have the beats write directly to Elasticsearch. For filebeat this is fine, as even in case of outages a filebeat can just remember where it stopped reading data in a file and go from there. If you run other beats that collect statistics they can spool some data to disk, but will ultimately drop data if Elasticsearch is not available - you could another component in between that collects data from several beats and persists it temporarily on disk, like logstash with a persistent queue or kafka (which would also allow to not be reliant on the indexing speed of elasticsearch in case of a surge of logs, think DoS). But as with everything there comes the cost of maintaining it, so I would always start small and then grow over time.

Hope this helps!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 18, 2020, 2:09pm UTC](https://discuss.elastic.co/t/filebeat-insert-logs-from-multiple-hosts-in-parallel/219952/3 "2020-03-18T14:09:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
