# Uneven node load

**URL:** <https://discuss.elastic.co/t/uneven-node-load/46245>\
**Category:** Elasticsearch\
**Created:** [April 4, 2016, 1:56pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245 "2016-04-04T13:56:32Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![akshat](https://avatars.discourse-cdn.com/v4/letter/a/22d042/32.png) [@akshat](https://discuss.elastic.co/u/akshat)\
**Post date:** [April 4, 2016, 1:56pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/1 "2016-04-04T13:56:32Z")

</div>

I have a 8 node cluster (3 master and 5 data nodes). There are 5 shards and 4 replicas, so each of the data nodes have identical shard information. There is one particular node which always shows high CPU usage. I have looked at other node stats like disk space used, queries processed etc and seem identical across all nodes. I tried hot spot threads and jstack but output appears similar across nodes. How can I debug why this node is misbehaving ?

---

<div class="post-metadata">

**Author:** ![MarkOStewart](https://avatars.discourse-cdn.com/v4/letter/m/898d66/32.png) [@MarkOStewart](https://discuss.elastic.co/u/MarkOStewart)\
**Post date:** [April 4, 2016, 3:24pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/2 "2016-04-04T15:24:26Z")

</div>

are you sure it is Elasticsearch causing the load?

Is the OS swapping?

Have you ran top and sort by CPU and then press C to show process information?

---

<div class="post-metadata">

**Author:** ![akshat](https://avatars.discourse-cdn.com/v4/letter/a/22d042/32.png) [@akshat](https://discuss.elastic.co/u/akshat)\
**Post date:** [April 4, 2016, 4:26pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/3 "2016-04-04T16:26:00Z")

</div>

Pretty sure it is ES. Nothing else runs on the machine and top shows high  
CPU usage by ES.

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [April 4, 2016, 5:32pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/4 "2016-04-04T17:32:31Z")

</div>

Do any of the shards have an unusually extra amount of documents on it?

Are you using any custom routing?

---

<div class="post-metadata">

**Author:** ![akshat](https://avatars.discourse-cdn.com/v4/letter/a/22d042/32.png) [@akshat](https://discuss.elastic.co/u/akshat)\
**Post date:** [April 4, 2016, 5:46pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/5 "2016-04-04T17:46:17Z")

</div>

I am not using custom routing. I haven't checked number of docs in each  
shard but all data is replicated in all the 5 nodes. Each node has 5 shards  
ensuring it has complete copy of entire data.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [April 4, 2016, 5:58pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/6 "2016-04-04T17:58:10Z")

</div>

> [@akshat](#):
>
> How can I debug why this node is misbehaving ?

I always use jstack. I usually run it a few times and dump the output to a file. I write a little bash script that tries to classify each stack trace with grep. Because I have a thing for silly bash scripts, I guess.

jstack really is the best way. If it doesn't say anything I check things like GC rates. It is probably also worth making sure that your problem node is running using the same configuration and that clients are pushing requests to the cluster randomly/round robin/whatever. Just so long as they aren't hammering that node in particular.

---

<div class="post-metadata">

**Author:** ![akshat](https://avatars.discourse-cdn.com/v4/letter/a/22d042/32.png) [@akshat](https://discuss.elastic.co/u/akshat)\
**Post date:** [April 8, 2016, 1:36pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/7 "2016-04-08T13:36:54Z")

</div>

I have 3 master nodes in the cluster which are behind a load balancer. I am  
assuming that they round robin the requests to distribute load in a  
reasonable manner. I tried multiple dumps using jstack but there doesn't  
seem any differences between loaded and other nodes.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:01pm UTC](https://discuss.elastic.co/t/uneven-node-load/46245/8 "2017-07-05T23:01:09Z")

</div>


