# Suggestion improving filebeat performance

**URL:** <https://discuss.elastic.co/t/suggestion-improving-filebeat-performance/105508>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [October 27, 2017, 4:28am UTC](https://discuss.elastic.co/t/suggestion-improving-filebeat-performance/105508 "2017-10-27T04:28:40Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![max\_dodo](https://avatars.discourse-cdn.com/v4/letter/m/c77e96/32.png) [@max\_dodo](https://discuss.elastic.co/u/max_dodo)\
**Post date:** [October 27, 2017, 4:28am UTC](https://discuss.elastic.co/t/suggestion-improving-filebeat-performance/105508/1 "2017-10-27T04:28:41Z")

</div>

My ELK setup goes like :  
Filebeats --\> Logstash(ingest nodes) --\> Elasticsearch ( master + Datanodes ) --\> Kibana

Recently we are observing a huge amount of delay in logfile ingestion ( 4 ~ 5 hr) . From the capacity perspective we have added enough horsepower ( large-machines: 10+ ingestnodes, 15+ datanodes ) . Per day total log size reaches upto 900gb. Multiple applications generating huge amount of logs.

Note- on a daily basis around 100+ logfiles are generated on a single server, each of 500mb size.

Our filebeat configuration is as below. Please suggest if anything can be modified or added to take care of this performance issue.

* * *

## filebeat.prospectors:

paths:  
- /\<log\_path\>/application\*.log  
fields:  
level: debug  
review: 1  
json.keys\_under\_root: true  
json.overwrite\_keys: true  
harvester\_buffer\_size: 16384  
scan\_frequency: 5s  
document\_type: \<document\_type\_name\>  
registry\_file: .filebeat  
spool\_size: 20480  
tail\_files: false  
idle\_timeout: 5s  
input\_type: log  
max\_backoff: 10s  
max\_bytes: 10485760

logging:  
files:  
keepfiles: 5  
name: filebeat.log  
path: /var/log/filebeat-logs  
rotateeverybytes: 10485760  
level: info  
to\_files: true  
to\_syslog: false

output.logstash:  
hosts:  
- "VIP-address:port"

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [October 27, 2017, 9:57am UTC](https://discuss.elastic.co/t/suggestion-improving-filebeat-performance/105508/2 "2017-10-27T09:57:29Z")

</div>

Start by analyzing the whole pipeline and figuring out where the bottleneck is. Looking at the CPU load on the machines should give you an indication. Is the load distributed reasonably evenly across the Logstash hosts?

---

<div class="post-metadata">

**Author:** ![max\_dodo](https://avatars.discourse-cdn.com/v4/letter/m/c77e96/32.png) [@max\_dodo](https://discuss.elastic.co/u/max_dodo)\
**Post date:** [October 27, 2017, 3:29pm UTC](https://discuss.elastic.co/t/suggestion-improving-filebeat-performance/105508/3 "2017-10-27T15:29:44Z")

</div>

Thanks for the response Magnus, Checked the whole ELK pipeline, there is no concern about the resource utilization on any of the nodes ( LS, DN, MN ) - CPU utilization hovers from 5-10% and MEM utilization is also normal.

We notices few things

1. there are multiple json parsefailures which we see in the logstash nodes
2. the number of logfiles opened by the filebeat is large on each client nodes ( Harvester started for file )

Can any of these cause the delay in ingestion of logfiles ?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 24, 2017, 3:29pm UTC](https://discuss.elastic.co/t/suggestion-improving-filebeat-performance/105508/4 "2017-11-24T15:29:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
