# Help please with high CPU utilization on 1 node of cluster :)

**URL:** <https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027>\
**Category:** Elasticsearch\
**Created:** [October 12, 2015, 6:15pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027 "2015-10-12T18:15:36Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 12, 2015, 6:15pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/1 "2015-10-12T18:15:37Z")

</div>

Hi all,

I'm having one data node of 6 dedicated nodes have very high CPU utilization. I'm hoping to have some help figuring out what is causing it.

Here is the output from "top".

```
top - 17:51:19 up 91 days, 3:38, 1 user, load average: 25.01, 23.68, 17.83
Tasks: 671 total, 1 running, 670 sleeping, 0 stopped, 0 zombie
Cpu(s): 0.0%us, 40.7%sy, 0.0%ni, 59.2%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Mem: 65858520k total, 64972200k used, 886320k free, 233832k buffers
Swap: 8388600k total, 24064k used, 8364536k free, 26568692k cached

   PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND          
 32206 elastics 20 0 763g 38g 6.4g S 1304.9 61.8 2187:13 java    

```

You can see that the load average is crazy high, and CPU utilization is as well. The problem does not fix itself (I've seen it run over 2 days), and all I can do is cycle the JVM.

This is a 16 core machine.  
This is the Java version:

```
root@bdprodes08:[218]:~> java -version
java version "1.7.0_79"
OpenJDK Runtime Environment (rhel-2.5.5.3.el6_6-x86_64 u79-b14)
OpenJDK 64-Bit Server VM (build 24.79-b02, mixed mode)

```

I'm running ES 1.6.0.

Here is the (very snipped) output of the hot\_threads API for this node:

```
::: [elasticsearch-bdprodes08][MWIHO5NQSjChyPm196c5ug][bdprodes08][inet[/10.200.116.248:9300]]{master=false}
   Hot threads at 2015-10-12T17:51:45.866Z, interval=500ms, busiestThreads=10, ignoreIdleThreads=true:
   
   100.8% (504.2ms out of 500ms) cpu usage by thread 'elasticsearch[elasticsearch-bdprodes08][bulk][T#4]'
     9/10 snapshots sharing following 14 elements
       org.apache.lucene.index.IndexWriter.updateDocument(IndexWriter.java:1526)
       org.apache.lucene.index.IndexWriter.addDocument(IndexWriter.java:1252)
       org.elasticsearch.index.engine.InternalEngine.innerCreateNoLock(InternalEngine.java:356)
       org.elasticsearch.index.engine.InternalEngine.innerCreate(InternalEngine.java:298)
       org.elasticsearch.index.engine.InternalEngine.create(InternalEngine.java:269)
       org.elasticsearch.index.shard.IndexShard.create(IndexShard.java:483)
   
   100.8% (504.2ms out of 500ms) cpu usage by thread 'elasticsearch[elasticsearch-bdprodes08][[derbysoft-ihg-20151012][0]: Lucene Merge Thread #174]'
     9/10 snapshots sharing following 20 elements
       java.io.FileOutputStream.writeBytes(Native Method)
       java.io.FileOutputStream.write(FileOutputStream.java:345)
       org.apache.lucene.store.FSDirectory$FSIndexOutput$1.write(FSDirectory.java:390)
       java.util.zip.CheckedOutputStream.write(CheckedOutputStream.java:73)
       java.io.BufferedOutputStream.flushBuffer(BufferedOutputStream.java:82)
       java.io.BufferedOutputStream.write(BufferedOutputStream.java:95)
       org.apache.lucene.store.OutputStreamIndexOutput.writeByte(OutputStreamIndexOutput.java:45)

   100.8% (504.1ms out of 500ms) cpu usage by thread 'elasticsearch[elasticsearch-bdprodes08][[derbysoft-ihg-20151012][0]: Lucene Merge Thread #181]'
     2/10 snapshots sharing following 21 elements
       java.io.FileOutputStream.writeBytes(Native Method)
       java.io.FileOutputStream.write(FileOutputStream.java:345)
       org.apache.lucene.store.FSDirectory$FSIndexOutput$1.write(FSDirectory.java:390)

   83.7% (418.2ms out of 500ms) cpu usage by thread 'elasticsearch[elasticsearch-bdprodes08][[nuke_metrics][1]: Lucene Merge Thread #66]'
     unique snapshot
       org.apache.lucene.index.ConcurrentMergeScheduler$MergeThread.run(ConcurrentMergeScheduler.java:476)

   16.9% (84.4ms out of 500ms) cpu usage by thread 'elasticsearch[elasticsearch-bdprodes08][[derbysoft-bestwestern-20151012][0]: Lucene Merge Thread #125]'
     unique snapshot
       org.apache.lucene.index.ConcurrentMergeScheduler$MergeThread.run(ConcurrentMergeScheduler.java:476)
     2/10 snapshots sharing following 21 elements

```

It looks like a combination of merges and bulk inserts, but I'm not sure why this one node in particular is having issues, while the rest are not. Right now, I'm having to cycle this node about every 2 or three days, which on this cluster size is somewhat of a pain. 😄

Here are my cluster settings for merging:

```
 "index.merge.scheduler.max_thread_count": "1",
 "index.merge.policy.type": "tiered",
 "index.merge.policy.max_merged_segment": "5gb",
 "index.merge.policy.segments_per_tier": "10",
 "index.merge.policy.max_merge_at_once": "10",
 "index.merge.policy.max_merge_at_once_explicit": "10",
 "index.translog.flush_threshold_size": "1gb",

```

The data directory is made up of 8 2TB SATA disks in a Raid 0 stripe.

Can anyone see something misconfigured, or suggest something to try to fix this?

Thank you so much for your time and help!  
Chris

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [October 12, 2015, 6:55pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/2 "2015-10-12T18:55:06Z")

</div>

> [@Chris\_Neal](#):
>
> This is a 16 core machine.

If this is _real_ load 25 is high for a 16 core machine but its workable.

> [@Chris\_Neal](#):
>
> It looks like a combination of merges and bulk inserts, but I'm not sure why this one node in particular is having issues, while the rest are not.

Do you happen to use a daily or weekly indexing strategy or something? Elasticsearch's allocation mechanisms think of every shard as equal but in the typical time based use case the current day/week/whatever's index get much much more load than the others because its being written to. Whenever I have an index that I know I'm going to hammer I use [total\_shards\_per\_node](https://www.elastic.co/guide/en/elasticsearch/reference/current/_total_shards_per_node.html) to force the shards apart from eachother. You might find this helpful. But you should be cognizant that it is a "hard" limit on allocation - elasticsearch won't violate the rule even if the cluster would go red. Its safe if you understand how it works.

> [@Chris\_Neal](#):
>
> The data directory is made up of 8 2TB SATA disks in a Raid 0 stripe.

There is a hot/warm architecture where you have 3 nodes that you use for indexing with buff solid state disks and when you stop writing to the index you use allocation awareness attributes to shift the indexes onto nodes with larger, slower spinning disks. I dunno if that is a solution to your problem, but it is a general solution to a problem like yours.

---

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 12, 2015, 7:07pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/3 "2015-10-12T19:07:49Z")

</div>

Thanks Nik.

> [@nik9000](#):
>
> If this is real load 25 is high for a 16 core machine but its workable.

25 is "average" for the events that I've seen. Where things get nasty is when it gets to 40+, then indexing gets pretty dramatically backed up until I bounce the node. Yuck. Is there anything configuration-wise that looks suspicious to you by chance. Anything else of interest?

> [@nik9000](#):
>
> Do you happen to use a daily or weekly indexing strategy or something? Elasticsearch's allocation mechanisms think of every shard as equal but in the typical time based use case the current day/week/whatever's index get much much more load than the others because its being written to. Whenever I have an index that I know I'm going to hammer I use total\_shards\_per\_node to force the shards apart from eachother. You might find this helpful. But you should be cognizant that it is a "hard" limit on allocation - elasticsearch won't violate the rule even if the cluster would go red. Its safe if you understand how it works.

Yes, we use a daily indexing strategy. I've got 6 data nodes, with number\_of\_shares = 6, and number\_of\_replicas = 1. So each node has one primary and one replica for each of my (currently) 16 indices. Honestly, I am a little hesitant to play with the total\_shards\_per\_node setting right now, just because things are so evenly distributed, I'm a little curious to see what it might be about this particular node that _always_ has the issue. I have yet to see it on any of the other 5.

> [@nik9000](#):
>
> There is a hot/warm architecture where you have 3 nodes that you use for indexing with buff solid state disks and when you stop writing to the index you use allocation awareness attributes to shift the indexes onto nodes with larger, slower spinning disks. I dunno if that is a solution to your problem, but it is a general solution to a problem like yours.

Yeah....That would definitely be ideal. As of now, I have no SSD drives in these servers, but that could be another step in the architecture evolution. I've read up on it, and would love to implement it at some point.

Thanks again for the input.  
Chris

---

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 13, 2015, 1:13pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/4 "2015-10-13T13:13:38Z")

</div>

Again. :((

Marvel for the last 24 hours.

 ![](https://us1.discourse-cdn.com/elastic/original/2X/4/4b458f3db7d5acaacfb04929deeac951fee531e8.png)

Still looking for any suggestions, and still looking myself!

Many thanks.  
Chris

---

<div class="post-metadata">

**Author:** ![tinle](https://avatars.discourse-cdn.com/v4/letter/t/c77e96/32.png) [@tinle](https://discuss.elastic.co/u/tinle)\
**Post date:** [October 13, 2015, 4:19pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/5 "2015-10-13T16:19:43Z")

</div>

Is this node the master?

I am seeing similar issue with just one node experiencing high CPU and constant GC, but it's the master node.

Might be different issue....

---

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 13, 2015, 4:31pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/6 "2015-10-13T16:31:41Z")

</div>

Hi Tin,

For me, the node is a dedicated data node, and I'm not having any GC issues. Might be a different issue!

Chris

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 13, 2015, 4:42pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/7 "2015-10-13T16:42:27Z")

</div>

Are you distributing indexing and query load evenly across the cluster?

---

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 13, 2015, 7:37pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/8 "2015-10-13T19:37:37Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Are you distributing indexing and query load evenly across the cluster?

Well, AFAIK I am. 😉 I have one primary and one replica shard for each index on each of the 6 dedicated data nodes. I don't think querying is an issue, because there are none going on when these high CPU events happen, and the hot threads API shows indexing and merging events as the busy threads.

I'm thinking I might play with some of those settings on the cluster (see previous thread for my settings) and see if perhaps I have them too high...although it still bothers me that these are only happening on this one server.

I've also looked at potential hardware issues (failed disks, etc) to see if that could be contributing, but there are none.

Still looking. Thanks for all the replies!  
Chris

---

<div class="post-metadata">

**Author:** ![tinle](https://avatars.discourse-cdn.com/v4/letter/t/c77e96/32.png) [@tinle](https://discuss.elastic.co/u/tinle)\
**Post date:** [October 13, 2015, 8:06pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/9 "2015-10-13T20:06:01Z")

</div>

Yes, sound like different issue than mine. I believe you have replica turned on. I have a 9 node cluster that see fairly high CPU, when I turn off replica for current (active) index, the load went down. I have a script that turn on replica for previous day index.

FWIW, you might want to try that.

Tin

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:45pm UTC](https://discuss.elastic.co/t/help-please-with-high-cpu-utilization-on-1-node-of-cluster/32027/10 "2017-07-05T23:45:02Z")

</div>


