# Need help diagnosing slow indexing speed

**URL:** <https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123>\
**Category:** Elasticsearch\
**Created:** [March 11, 2020, 12:04pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123 "2020-03-11T12:04:40Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [March 11, 2020, 12:04pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/1 "2020-03-11T12:04:41Z")

</div>

Hi guys!  
I need some help diagnosing my cluster performance.  
Indexing is very slow and the whole cluster as well as Kibana is pretty unresponsive.  
Here is my cluster:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/b/7bc1b1ce8b32a9848e358ea28635351aed1ba378.png)  
We index about 1,5TB or 1.5 bn events every day.  
We have multiple indices indexing at the same time with the most load originating from the logstash-\* index. (38 Shards, no replicas)  
Our cluster is split into two tiers.

- T1: SSD, high CPU, high RAM
- T2: HDD, medium CPU, high RAM

Here are the index settings for the logstash-\* index:  
[https://gist.github.com/xenoid/9c705ee8b2a5f8b7989be312b887a2dd](https://gist.github.com/xenoid/9c705ee8b2a5f8b7989be312b887a2dd)  
This is the config of a typical T1 Data node:  
[https://gist.github.com/xenoid/c35d7d4a5f39f040abdaa1b5b4bd4282](https://gist.github.com/xenoid/c35d7d4a5f39f040abdaa1b5b4bd4282)

A screenshot of the Monitoring page for the node:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/9/a/9ae1998b0640bc2ec3851084ae215532c405183e.png)

A screenshot of the Monitoring page for the logstash-\* index:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/2/0/20a07c7d70be8801038c2424828287fc73394ce2.png)

Here are a few minutes of logs from the data node:  
[https://gist.github.com/xenoid/2278a059b0500e22ad6bcaaa86da7874](https://gist.github.com/xenoid/2278a059b0500e22ad6bcaaa86da7874)

Here you can see the bulk indexing queue of several nodes:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/d/b/dbcfbd7aca9919c5a4643c46bbf65348c76a05f3.jpeg)

Thanks for reading this far.  
Please hit me up if you need any more information.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 11, 2020, 12:18pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/2 "2020-03-11T12:18:24Z")

</div>

What is the size of your documents? Do you use parent-child or nested documents? Do you perform any updates of existing data or just index new documents? Are you letting Elasticsearch assign the document id or providing one externally?

---

<div class="post-metadata">

**Author:** ![wifi](https://avatars.discourse-cdn.com/v4/letter/w/f05b48/32.png) [@wifi](https://discuss.elastic.co/u/wifi)\
**Post date:** [March 11, 2020, 12:21pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/3 "2020-03-11T12:21:04Z")

</div>

To have everything running smooth your system load needs to be below 1

Have you set up some sort of ingest pipeline that performs heavy operations?

How do you work with the data? If you perform aggregations you should rather create a rollup index and use it instead of raw data.

---

<div class="post-metadata">

**Author:** ![wifi](https://avatars.discourse-cdn.com/v4/letter/w/f05b48/32.png) [@wifi](https://discuss.elastic.co/u/wifi)\
**Post date:** [March 11, 2020, 12:37pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/4 "2020-03-11T12:37:05Z")

</div>

> <https://gist.github.com/xenoid/9c705ee8b2a5f8b7989be312b887a2dd#file-logstash_index_config-L140>

Maybe the codec should be set to something else? I assume that compression needs CPU. I believe that you have too many tasks that get scheduled back and forth.

Can you run `GET _tasks` and post the result?

---

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [March 11, 2020, 12:57pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/5 "2020-03-11T12:57:57Z")

</div>

Our data consists mainly of log files (webserver, firewall, proxy, ...). Is there a way i can measure the exact size of an average document?  
No updates on existing data is performed.  
Document IDs are set automatically.  
CPU should be good enough. 64 to 128 Cores on every T1 machine.  
Here is a CPU Load graph:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/3/532a3879c6facb135179626f707840461d8130cd.jpeg)

No special index pipeline is set up. Processing is done in Logstash.

Here is a list of GET \_tasks:

> <https://gist.github.com/xenoid/2d4a9de9aa390d5f5d6ca047d5fd9653>

---

<div class="post-metadata">

**Author:** ![wifi](https://avatars.discourse-cdn.com/v4/letter/w/f05b48/32.png) [@wifi](https://discuss.elastic.co/u/wifi)\
**Post date:** [March 11, 2020, 1:27pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/6 "2020-03-11T13:27:24Z")

</div>

Looking at the tasks i see two issues

1. The tasks are distributed unequally

2. You have two machines that get a lot of indexing work:

and

```
   "LKveL5ccQHqXN2cP41U9wg" : {
    "name" : "server-0259",
    "transport_address" : "10.2.0.237:9301",
    "host" : "10.2.0.237",
    "ip" : "10.2.0.237:9301",
    "roles" : [
      "ingest"
    ],
    "attributes" : {
      "zone" : "VM",
      "xpack.installed" : "true"
    },

```

Maybe this helps somehow? I have no clue about elasticsearch performance diagnosis and just taking a wild guess.

---

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [March 11, 2020, 1:36pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/7 "2020-03-11T13:36:31Z")

</div>

Good catch! server-0258 and server-0259 are the "indexing" nodes. They get the requests from logstash.  
They have slightly less CPUs (24) but I noticed server-0259 crashing several times a day without any notice in the log.

---

<div class="post-metadata">

**Author:** ![wifi](https://avatars.discourse-cdn.com/v4/letter/w/f05b48/32.png) [@wifi](https://discuss.elastic.co/u/wifi)\
**Post date:** [March 11, 2020, 1:48pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/8 "2020-03-11T13:48:06Z")

</div>

The server-258 is pretty busy with indexing and on top of that it also gets requests like `data/read/search` and `admin/mappings/get`

maybe you should take a look at the routing of these requests

Hmmm is it possible that this server also holds kibana?

In the config of kibana you could try to set up all of your nodes instead of just one. If you set up only one server it will act like a loadbalancer or proxy. If it is the overloaded ingest node kibana might be slow

---

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [March 11, 2020, 2:53pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/9 "2020-03-11T14:53:58Z")

</div>

server-0258 and server-0259 are the "outward-facing" nodes. They take care of Kibana, REST queries and connection to logstash.  
This never caused issues, but some time ago the cluster started performing badly.  
I suspect some old setting as the cause.

As the cluster has been upgraded from ES 0.\* to now ES6.8.1 , it is inevitable that some deprecated setting will persist somewhere.

But in the next days I will probably try to move Kibana to another server and try again.

EDIT: @wifi  
I just changed the ES-hosts in the Kibana configs. Performance did not change.  
Kibana 6 takes ages to load the discover page when the logstash-\* index is selected. I put the blame on some sort of mild mapping explosion in the logstash-\* index.

---

<div class="post-metadata">

**Author:** ![wifi](https://avatars.discourse-cdn.com/v4/letter/w/f05b48/32.png) [@wifi](https://discuss.elastic.co/u/wifi)\
**Post date:** [March 11, 2020, 4:18pm UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/10 "2020-03-11T16:18:50Z")

</div>

What did you put into the kibana setting `elasticsearch.hosts:`

---

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [March 12, 2020, 7:57am UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/11 "2020-03-12T07:57:33Z")

</div>

I put a few of our fastest data nodes in there.  
I also tried reducing the number of replicas on the .kibana index down to 5 from 44.  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/7/77e849b00afe554a873f5fcecf8b579baceb2337.png)  
It did not change performance.

---

<div class="post-metadata">

**Author:** ![wifi](https://avatars.discourse-cdn.com/v4/letter/w/f05b48/32.png) [@wifi](https://discuss.elastic.co/u/wifi)\
**Post date:** [March 12, 2020, 9:10am UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/12 "2020-03-12T09:10:21Z")

</div>

The bad performance you experience is only in Kibana right? Maybe you can get some clues if you use F12 - debugging Tools in the browser. The network monitor in Chrome could be interesting. BTW which Browser do you use?

---

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [March 12, 2020, 9:22am UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/13 "2020-03-12T09:22:14Z")

</div>

It's definitely not just in Kibana.  
I just installed a separate server for Kibana and it's decently fast. As soon as ES is involved it gets slow. I tried IE, Firefox, Chrome, Edge and Brave.  
Also here's todays indexing queue:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/b/3bfb6e0b9d99f9ea4d4e6bb0fab675321263ea7f.jpeg)  
(Different colors represent different nodes)  
If I only could diagnose what is making these bulk queries so slow.  
They got fast SSDs, big CPUs, and good heap.

EDIT:  
Here is the graph after I changed the refresh interval from 30s to 180s:  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/a/3ae469e06bbbf1d1a5d759c49911113afd545671.png)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 9, 2020, 9:22am UTC](https://discuss.elastic.co/t/need-help-diagnosing-slow-indexing-speed/223123/14 "2020-04-09T09:22:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
