# Filebeat to Elasticsearch log shipping is very slow

**URL:** <https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570>\
**Category:** Elasticsearch\
**Created:** [June 12, 2018, 4:04pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570 "2018-06-12T16:04:02Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![rajisankar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rajisankar/32/43250_2.png) [@rajisankar](https://discuss.elastic.co/u/rajisankar)\
**Post date:** [June 12, 2018, 4:04pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/1 "2018-06-12T16:04:02Z")

</div>

Hi,

I have found some posts on this before. But, have not got a definitive answer with respect to Beats 6.x as spool\_size has been removed.

The following tests have been performed on  
Filebeat 6.x  
Elasticsearch 6.2.4 with 16GB Heap

Filebeat config with File Output gives 80,000 events/s

output.file:  
path: "/opt/CCURfilebeat"  
filename: filebeat  
number\_of\_files: 7  
permissions: 0600

queue:  
mem:  
events: 40000  
flush.min\_events: 20000

Filebeat config with elasticsearch output gives 14,000 events/s  
output.elasticsearch:  
hosts: ["elastic-server:9200"]  
bulk\_max\_size: 20000  
username: "elastic"  
password: "elasticpassword"

queue:  
mem:  
events: 40000  
flush.min\_events: 20000

The elasticsearch indexing rate is 13,800 events/s and this seems to the bottleneck.

What i dont understand is Elasticsearch **CPU Utilization is 10%** and JVM Heap Used is **6GB/16GB**. Then why is the indexing rate still so low? What other factors should we consider to stress the elasticsearch system?

Any suggestions on improving this performance would be highly appreciated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 12, 2018, 4:18pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/2 "2018-06-12T16:18:57Z")

</div>

Have you optimised Elastichsearch for [indexing speed](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/tune-for-indexing-speed.html)? What type of storage do you have? What is disk I/O and iowait looking like? How many nodes in the cluster?

---

<div class="post-metadata">

**Author:** ![rajisankar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rajisankar/32/43250_2.png) [@rajisankar](https://discuss.elastic.co/u/rajisankar)\
**Post date:** [June 21, 2018, 6:10pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/3 "2018-06-21T18:10:46Z")

</div>

Hi,

After a brief benchmark, it appears that the problem was with Elasticsearch indexing speed. Thanks, Christian!

For Benchmarking purposes, we have a single node Elasticsearch. This has a single index with one shard and no replicas. We use a 400GB SSD , 32GB RAM in which 16GB is allocated for Heap, 12 Core Processor.  
we are using a single thread and a 2GB flush threshold and 30s refresh interval.  
Index settings are as follows.

> "index.merge.scheduler.max\_thread\_count" : "1",  
> "index.translog.flush\_threshold\_size" : "2gb",  
> "index.refresh\_interval": "30s",  
> "index.mapping.total\_fields.limit":"30000"

i am unable to index more than 60,000 document/s from Filebeat.

Using X-Pack monitoring, the CPU Utilisation is 60%, JVM Heap Utilization is 72% and disk I/O is 130 MB/s. Clearly none of these factors is the bottleneck.

I am not sure what else might be the bottleneck. Is there a way to find what factors might attribute to this?  
Thanks in advance.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 21, 2018, 6:44pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/4 "2018-06-21T18:44:57Z")

</div>

If you are using dynamic mappings (I am guessing this may be the case based on the number of fields you have specified) and are adding fields as indexing progresses, each change will require the cluster state to get updated, which can slow indexing down.

---

<div class="post-metadata">

**Author:** ![rajisankar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rajisankar/32/43250_2.png) [@rajisankar](https://discuss.elastic.co/u/rajisankar)\
**Post date:** [June 21, 2018, 7:22pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/5 "2018-06-21T19:22:28Z")

</div>

Sorry, that parameter was not needed. We use static mappings. Any other factors that might affect this performance?

I am wondering why its set at 60,000 documents/s when Elasticsearch can do much more. Not knowing what the bottleneck is bothering me. I am sure i am missing something here.

---

<div class="post-metadata">

**Author:** ![rajisankar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rajisankar/32/43250_2.png) [@rajisankar](https://discuss.elastic.co/u/rajisankar)\
**Post date:** [June 22, 2018, 4:22pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/6 "2018-06-22T16:22:40Z")

</div>

From Elasticsearch Benchmarking for HTTP Logs at [https://elasticsearch-benchmarks.elastic.co/index.html#tracks/http-logs/nightly/30d](https://elasticsearch-benchmarks.elastic.co/index.html#tracks/http-logs/nightly/30d) , the number of documents indexed seems to be 171,000 docs/s for 3-node Elasticsearch.

With one Node elasticsearch, i am able to, 60,000 docs/s. This seems to be fine though. But, is this comparison valid?

Apart from the usual resources, like CPU, Memory, Disk I/O, Network , what other factors could limit elasticsearch performance?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 22, 2018, 4:28pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/7 "2018-06-22T16:28:00Z")

</div>

Documents per second is not really a very good measurement of indexing performance as it will depend a lot on the size and complexity of the documents being indexed. You will get a better comparison if you run the same Rally track on your hardware.

---

<div class="post-metadata">

**Author:** ![rajisankar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rajisankar/32/43250_2.png) [@rajisankar](https://discuss.elastic.co/u/rajisankar)\
**Post date:** [June 22, 2018, 6:46pm UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/8 "2018-06-22T18:46:26Z")

</div>

Yeah, that makes sense. But, this benchmark is for HTTP logs and i am importing raw logs from a HTTP server as well. Hence, was hoping it would be close enough.

I am still struggling with finding what else could be the bottleneck. Any pointers to that is highly appreciated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 25, 2018, 6:57am UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/9 "2018-06-25T06:57:57Z")

</div>

The standard HTTP logs track uses very small documents, so it may or may not be comparable. I created a track that simulates events that are a bit larger and probably is closer to what you would get out of Filebeat. We talked about it [here](https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right) and it is available [on GitHub](https://github.com/elastic/rally-eventdata-track).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 23, 2018, 6:58am UTC](https://discuss.elastic.co/t/filebeat-to-elasticsearch-log-shipping-is-very-slow/135570/10 "2018-07-23T06:58:08Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
