# Scaling beats(filebeat at this moment) ~5k servers

**URL:** <https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006>\
**Category:** Beats\
**Created:** [July 20, 2016, 4:57pm UTC](https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006 "2016-07-20T16:57:21Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![lightkz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lightkz/32/12386_2.png) [@lightkz](https://discuss.elastic.co/u/lightkz)\
**Post date:** [July 20, 2016, 4:57pm UTC](https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006/1 "2016-07-20T16:57:21Z")

</div>

Hi Beats folks,

We just recently started playing with filebeats and see a lot of usefulness. But we got a bit stuck with scaling, configs management, and some limitation(by design [https://github.com/elastic/beats/issues/1112](https://github.com/elastic/beats/issues/1112)) for outputs.

Our pipelines look like these.

server -\> kafka -\> logstash -\> elasticsearch

or

server -\> kafka -\> samza -\> elasticsearch

For delivering and deploying we are using puppet. So it make sure it pushes and installs filebeats. As for config management we have inhouse developed framework(cover all our apps), which require substantial changes to accommodate filebeat deployment. We POC it and it is working but has some limitation on metadata to support multiple topics for multiple filebeat processes on the server.

I was wondering how others scale beats in env, assuming we have multiple files on each server that we need to filebeat and deliver to multiple different topics(lets say only to one broker for now)?

---

<div class="post-metadata">

**Author:** ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)\
**Post date:** [July 20, 2016, 7:49pm UTC](https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006/2 "2016-07-20T19:49:34Z")

</div>

I don't fully understand the actual configuration problems you're facing. More important, what exactly you want filebeat todo.

Sending to kafka you want to push to multiple topics? One case use `output.kafka.use_type` and `filebeat.prospectors.X.document_type`, to configure different topics per prospector-type. Support for choosing topics might be enhanced in future versions.

Filebeat supports environment variable for changing settings + some configurable 'config'-directory. The directory support allows you to put multiple prospector configurations into one directory. e.g. use puppet on machine type to put some config per service into config directory. After restarting filebeat the prospector configs are merged with main filebeat config.

Recent nightly builds support:

- load multiple config files by using `-c <file>` option multiple times
- overwrite any config setting from command line using `-E <setting>=<value>`.

---

<div class="post-metadata">

**Author:** ![lightkz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lightkz/32/12386_2.png) [@lightkz](https://discuss.elastic.co/u/lightkz)\
**Post date:** [July 20, 2016, 8:57pm UTC](https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006/3 "2016-07-20T20:57:25Z")

</div>

I was not sure how to push different events to different topics per prospector. For example app generates three different files and I would like to push those files to different topics. From what you are saying we need to do this.

filebeat.config\_dir - would have all our prospectors configuration.

main config file for output would have something:

> output.kafka:  
> hosts: ["broker"]  
> use\_type: true  
> compression: snappy

and our prospectors would be look like this:

> filebeat.prospectors:  
> document\_type: **topic\_name?**
> 
> - input\_type: log  
> paths:
> - "/blah/blah/log\_file"

and then star would look like **filbeat -c main\_config.yml**?

or another way I will run multiple times filebeat with different config files.

---

<div class="post-metadata">

**Author:** ![andrewkroh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrewkroh/32/3784_2.png) [@andrewkroh](https://discuss.elastic.co/u/andrewkroh)\
**Post date:** [July 21, 2016, 10:55pm UTC](https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006/4 "2016-07-21T22:55:36Z")

</div>

@lightkz, sounds right to me. But the `document_type` option is part of each individual prospector in the array. So just specify it like so:

```auto
filebeat.prospectors:
  - input_type: log
    paths:
      - "/blah/blah/log_file"
    document_type: topic_name

```

---

<div class="post-metadata">

**Author:** ![lightkz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lightkz/32/12386_2.png) [@lightkz](https://discuss.elastic.co/u/lightkz)\
**Post date:** [July 22, 2016, 7:19pm UTC](https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006/5 "2016-07-22T19:19:56Z")

</div>

@andrewkroh, ehh... I see. Thank you so much!

That will solve a lot of our deployment strategy.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 10, 2016, 4:58pm UTC](https://discuss.elastic.co/t/scaling-beats-filebeat-at-this-moment-5k-servers/56006/6 "2016-08-10T16:58:08Z")

</div>

This topic was automatically closed after 21 days. New replies are no longer allowed.
