# Elasticsearch high CPU usage on a mostly bulk indexing use case

**URL:** <https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196>\
**Category:** Elasticsearch\
**Created:** [May 18, 2020, 8:31pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196 "2020-05-18T20:31:25Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![ssouris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ssouris/32/68657_2.png) [@ssouris](https://discuss.elastic.co/u/ssouris)\
**Post date:** [May 18, 2020, 8:31pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/1 "2020-05-18T20:31:26Z")

</div>

Hi, we are running `Elasticsearch 7.6.2` on our own in AWS.

We have:

- 6 data nodes on `i3en.xlarge` instances
- 3 dedicated master nodes
- 200 indices storing ~800million documents.
- Total size ~5TB
- 30s refresh interval

We are observing high CPU usage most of the time (above 60-80% with spikes at 100%).  
Although the performance of the cluster seems acceptable:

- ingesting ~500 docs/second
- no rejections on bulk thread pool
- query times look good

I'm a bit concerned about the CPU usage, because we expect more data from our customers, and we are going to add more nodes or even increase the specs of the existing ones to handle all the additional load.

This is the output from `Hot Threads`:

> <https://gist.github.com/ssouris/9475140109fa12990d58c2ff97f40371>

It seems like Elasticsearch is spending a lot of time merging.

I would like to understand if such a high CPU usage is expected on a well-balanced cluster, especially when it's mostly bulk indexing.  
I tried to increase `refresh_interval` but the CPU usage didn't change at all. GC activity changed a bit for sure, and the merge sizes increased but the CPU remained high.

Thanking in advance.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 19, 2020, 5:17am UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/2 "2020-05-19T05:17:33Z")

</div>

What is the size of the documents? Are nested documents used to a large extent? Are you updating documents or just indexing new ones? Are you by any chance forcing a refresh after each bulk request?

---

<div class="post-metadata">

**Author:** ![ssouris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ssouris/32/68657_2.png) [@ssouris](https://discuss.elastic.co/u/ssouris)\
**Post date:** [May 19, 2020, 3:36pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/3 "2020-05-19T15:36:59Z")

</div>

Thanks for answering so fast.

_What is the size of the documents?_  
3-4KB per document. Also, most of the documents have more than 1K fields.

_Are nested documents used to a large extent?_  
Yes, 1/6th of the documents have nested documents.

_Are you updating documents or just indexing new ones?_  
Only indexing new ones.

_Are you by any chance forcing a refresh after each bulk request?_  
I am using Jest Client to connect to the cluster and I am not passing any parameter to force a refresh.  
I, also, debugged the code and I didn't see any parameter being passed to refresh immediately after the bulk insert by the client library.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 19, 2020, 4:06pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/4 "2020-05-19T16:06:35Z")

</div>

How can a document of 4kB have 1000 fields? Does not sound right.

Large and complex documents require Elasticsearch to do as lot of work per document so will be a lot slower and use more CPU than smaller documents.

---

<div class="post-metadata">

**Author:** ![ssouris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ssouris/32/68657_2.png) [@ssouris](https://discuss.elastic.co/u/ssouris)\
**Post date:** [May 19, 2020, 5:37pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/5 "2020-05-19T17:37:47Z")

</div>

Actually I sent the average from a daily index.

Some of the documents can go over the 1K field limit and can be more than 100KB in size.

I guess if I manage to reduce the document size somehow and/or avoid indexing fields that are not needed I can get the CPU usage down.

---

<div class="post-metadata">

**Author:** ![ssouris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ssouris/32/68657_2.png) [@ssouris](https://discuss.elastic.co/u/ssouris)\
**Post date:** [May 20, 2020, 2:29pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/6 "2020-05-20T14:29:40Z")

</div>

_Related to the high CPU._

I observed that all the daily indices are getting refreshed, which makes sense, but most of my indices especially older than 2-5 days, will never change, so refresh is not needed.

I will try to disable `refresh_interval` on older indices to see if it makes any difference.

---

<div class="post-metadata">

**Author:** ![ssouris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ssouris/32/68657_2.png) [@ssouris](https://discuss.elastic.co/u/ssouris)\
**Post date:** [May 22, 2020, 11:20pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/7 "2020-05-22T23:20:21Z")

</div>

In the end I ended up adding four more nodes to the cluster, and now the CPU usage appears to be normal.  
I don't know if I'm obsessing over the correct thing (high CPU usage),

but I feel that:

1. I might need to configure the bulk size of the writes. Right now we do continuous batches of 200 records straight from Kafka. Maybe I should try increasing that size.

2. Also in a bulk I might have 6-12 different indices, so maybe it puts more strain to the cluster while bulk indexing.

---

<div class="post-metadata">

**Author:** ![Rahul\_Kumar4](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rahul_kumar4/32/67369_2.png) [@Rahul\_Kumar4](https://discuss.elastic.co/u/Rahul_Kumar4)\
**Post date:** [May 23, 2020, 7:30pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/8 "2020-05-23T19:30:26Z")

</div>

I am curious as we have a similar situation. Did disabling refresh\_interval on older indices make any difference?

---

<div class="post-metadata">

**Author:** ![ssouris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ssouris/32/68657_2.png) [@ssouris](https://discuss.elastic.co/u/ssouris)\
**Post date:** [May 23, 2020, 8:11pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/9 "2020-05-23T20:11:37Z")

</div>

No difference at all in our case with `refresh_interval`.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 23, 2020, 8:21pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/10 "2020-05-23T20:21:26Z")

</div>

If you are indexing into many indices and shards it may result in small batches being processed by individual shards. This can lead to a lot of disk I/O so it would be good to look at disk utilisation and iowait, e.g. using iostat. What type of storage do you have? Local SSDs?

---

<div class="post-metadata">

**Author:** ![ssouris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ssouris/32/68657_2.png) [@ssouris](https://discuss.elastic.co/u/ssouris)\
**Post date:** [May 26, 2020, 9:59pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/11 "2020-05-26T21:59:35Z")

</div>

This is how IO operations look like from my dashboard.

_I/O operations_

 ![Screen Shot 2020-05-26 at 10.53.26 PM](https://us1.discourse-cdn.com/elastic/original/3X/5/2/52f214ddc967732d8200234d06d5033ab77ce186.png)

_Disk Average Wait Time_

 ![Screen Shot 2020-05-26 at 10.58.00 PM](https://us1.discourse-cdn.com/elastic/original/3X/6/c/6c7bb93bc4d3ec1c97db1fda2de82e53793fc469.png)

Yes I'm using Local SSDs.  
By adding 4 more nodes to the cluster the CPU usage went lower.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 23, 2020, 9:59pm UTC](https://discuss.elastic.co/t/elasticsearch-high-cpu-usage-on-a-mostly-bulk-indexing-use-case/233196/12 "2020-06-23T21:59:48Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
