# ELK to slow

**URL:** <https://discuss.elastic.co/t/elk-to-slow/79312>\
**Category:** Logstash\
**Created:** [March 20, 2017, 7:50pm UTC](https://discuss.elastic.co/t/elk-to-slow/79312 "2017-03-20T19:50:39Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 20, 2017, 7:50pm UTC](https://discuss.elastic.co/t/elk-to-slow/79312/1 "2017-03-20T19:50:39Z")

</div>

I have elk installation:  
`logshash -> rmq -> logstash(2 instances) -> Elasticsearch`  
Elasticsearch cluster with 3 nodes. elasticsearch-5.2 (8 cores, 32 Gb, SSD) 12Gb HEAP  
`"number_of_replicas": 0`

> cluster.name: logs  
> node.name: el1,2,3  
> network.host: 0.0.0.0  
> transport.host: 172.30.30.180  
> http.port: 9200  
> indices.recovery.max\_bytes\_per\_sec: 150mb  
> http.cors.enabled: true  
> http.cors.allow-origin: "\*"  
> discovery.zen.ping.unicast.hosts: ["192.168.0.1", "192.168.0.2", "192.168.0.3"]  
> discovery.zen.minimum\_master\_nodes: 2  
> action.destructive\_requires\_name: true

Logstash-5.2

> pipeline.workers: 8  
> pipeline.batch.size: 500  
> -Xms4g  
> -Xmx4g

And my cluster can't consume more than 1.5 k events per sec from rmq.  
My application produce more than 2.5k events per sec.  
Almost all events are logs in json format so logstash config looks like this:

> filter { json { source =\> "message" } }

I believe my cluster can work better. How cat I find a bottleneck? LA on severs about 0.2

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [March 20, 2017, 7:53pm UTC](https://discuss.elastic.co/t/elk-to-slow/79312/2 "2017-03-20T19:53:49Z")

</div>

Have you tried using the json\_lines codec in the input instead of using the json filter?

---

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 20, 2017, 8:24pm UTC](https://discuss.elastic.co/t/elk-to-slow/79312/3 "2017-03-20T20:24:34Z")

</div>

Actually I think issue in Elastic configuration.  
I tried to remove all filters, left only input and output and result almost the same - 1,7-1,8k events.

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [March 20, 2017, 9:02pm UTC](https://discuss.elastic.co/t/elk-to-slow/79312/4 "2017-03-20T21:02:45Z")

</div>

> [@permyakovsv](#):
>
> Actually I think issue in Elastic configuration.

You have a 12G heap on each Elasticsearch node?

> [@permyakovsv](#):
>
> Elasticsearch cluster with 3 nodes. elasticsearch-5.2 (8 cores, 32 Gb, SSD) 12Gb HEAP

How many indices do you have? How many shards per index do you have?

---

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 21, 2017, 8:06am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/5 "2017-03-21T08:06:28Z")

</div>

Only 1 logstash index, cluster is almost emty.  
Shards 5.

---

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 21, 2017, 9:32am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/6 "2017-03-21T09:32:47Z")

</div>

If I set "number\_of\_replicas": 0 indexing speed increase to 10-12K

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 21, 2017, 2:50pm UTC](https://discuss.elastic.co/t/elk-to-slow/79312/7 "2017-03-21T14:50:43Z")

</div>

What is the average size of your events?

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [March 21, 2017, 4:04pm UTC](https://discuss.elastic.co/t/elk-to-slow/79312/8 "2017-03-21T16:04:52Z")

</div>

> [@permyakovsv](#):
>
> If I set "number\_of\_replicas": 0 indexing speed increase to 10-12K

This does suggest your cluster is at least somewhat I/O bound. Removing replicas takes away 1/2 of the writes if you have 1 replica per primary shard. However, it seems you have a 5x speed boost. I'm with @Christian_Dahlqvist, here, in wondering how large your documents/event are.

---

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 22, 2017, 9:38am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/9 "2017-03-22T09:38:45Z")

</div>

~ 1.4 KB

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 22, 2017, 9:49am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/10 "2017-03-22T09:49:32Z")

</div>

The increase in performance due to disabling replicas is larger than I would expect. How many replicas did you have configured before? Apart from disk I/O, network usage should also drop with reduced number of replicas. What type of networking do you have in place? Is the cluster deployed on bare-metal hardware or VMs with shared networking?

---

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 24, 2017, 9:34am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/11 "2017-03-24T09:34:46Z")

</div>

Bare metal, 1000Mb/s network.

'{ "number\_of\_replicas": 0 }' - \> index.rate: 8010.0  
'{ "number\_of\_replicas": 1 }' - \> index.rate: 2048.2

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 24, 2017, 9:44am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/12 "2017-03-24T09:44:37Z")

</div>

> [@permyakovsv](#):
>
> ~ 1.4 KB

Is this the raw size or the size of the JSON documents you are ingesting? How many fields? Do you use nested mappings?

Do you have X-Pack monitoring installed so you can provide graphs around node metrics?

---

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 24, 2017, 10:04am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/13 "2017-03-24T10:04:31Z")

</div>

1.4 RAW size.  
~ 40 fields.

Yes, I have X-Pack, what graphs can help you?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 24, 2017, 10:10am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/14 "2017-03-24T10:10:07Z")

</div>

What is the average size of the documents once converted to JSON (I find this a better measure than raw size as it accounts for any enrichment being performed)?

Graphs showing CPU and heap usage as well as indexing and query rates during indexing would be useful. Full screenshots of the `Nodes/Overview` and `Nodes/Advanced` screens would give us a better idea.

---

<div class="post-metadata">

**Author:** ![permyakovsv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/permyakovsv/32/16521_2.png) [@permyakovsv](https://discuss.elastic.co/u/permyakovsv)\
**Post date:** [March 24, 2017, 11:30am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/15 "2017-03-24T11:30:39Z")

</div>

1569 bytes, I count it like this: `"store"."size_in_bytes"/"docs"."count"`  
Graphs, now {"number\_of\_replicas": 0 } :

 ![](https://us1.discourse-cdn.com/elastic/original/3X/d/8/d8391efa2fe0e26ae8e39a156875aa5b79039df6.png) ![](https://us1.discourse-cdn.com/elastic/original/3X/e/2/e24c192c97dcb45c8c1ed97bce8340ca62ed7d82.png)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 21, 2017, 11:31am UTC](https://discuss.elastic.co/t/elk-to-slow/79312/16 "2017-04-21T11:31:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
