# All load is being concentrated on one node?

**URL:** <https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488>\
**Category:** Elasticsearch\
**Created:** [September 30, 2018, 9:20pm UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488 "2018-09-30T21:20:55Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [September 30, 2018, 9:20pm UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/1 "2018-09-30T21:20:55Z")

</div>

I have a problem with elastic search, all the load is concentrated on just one node, if I add new nodes they just sit there doing nothing.  
How can I make elastic search distribute the load around all the nodes on the cluster?

I'm running a 3 node cluster

The cluster is just serving search requests its not injesting or indexing data.  
the biggest index is 100 million records  
I'm searching product titles for keywords. nothing complex.

here's the output from:

curl localhost:9200/\_cat/nodes?v

```
ip heap.percent ram.percent cpu load_1m load_5m load_15m node.role master name
xxx.xxx.xxx.48 9 99 21 39.54 37.08 31.97 mdi - ES2
xxx.xxx.xxx.77 11 99 40 9.82 11.21 11.24 mdi * ES1
xxx.xxx.xxx.223 11 99 16 2.75 4.05 4.04 di - ES3

```

as you can see most of the load is on ES2 while ES3 is doing nothing.  
I'm sending all the search requests to ES3 in the hope that it would take over some of the work.

All these servers have: 24 cpus, 65gb ram

running: ES 6.4.1

I basically installed Elastic Search from scratch,  
change the /etc/elasticsearch/jvm.options to -Xms24g -Xmx24g  
set discovery.zen.minimum\_master\_nodes: 2  
all nodes are data nodes, with two masters  
I hooked up the nodes to the cluster as normal.

I've tried changing the number of primary shards from 5 to 10 with 1 replica but it doesn't make any difference.  
I've tried restarting the servers / Elastic Search multiple times.  
The cluster status is green.

i'm lost as what to try next!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 1, 2018, 6:10am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/2 "2018-10-01T06:10:57Z")

</div>

Are the shards evenly distributed across the nodes? Are all of the nodes sharing the same specification and configuration?

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 6:19am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/3 "2018-10-01T06:19:39Z")

</div>

Yes all shards are evenly distributed

all nodes are exactly the same with the same config and Os / memory / cpu etc...

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 1, 2018, 6:22am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/4 "2018-10-01T06:22:14Z")

</div>

Are you hitting all indices evenly? Are you using some feature that could cause an imbalanced load, e.g. custom routing or parent-child?

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 6:26am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/5 "2018-10-01T06:26:23Z")

</div>

I expect some indices are reciving a lot more search requests than others.

I haven't set up any custom routing or parent child features

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 1, 2018, 6:29am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/6 "2018-10-01T06:29:09Z")

</div>

Are you querying using preference, which would cause the same shards to be queried? Are the shards for most frequently queried indices evenly distributed?

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 6:37am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/7 "2018-10-01T06:37:10Z")

</div>

this is the search query i'm using..

\_search?q=title:$keyword&from=$start&size=50

this is the shard distribution on the busy index..

 ![shards](https://us1.discourse-cdn.com/elastic/original/3X/a/4/a480580d6903abb5ec95c5728781347994466d67.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 1, 2018, 6:38am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/8 "2018-10-01T06:38:10Z")

</div>

What is the output of the hot threads API on the busy node?

Although it is not related, I would also recommend making the third node master-eligible as well. You always want a minimum of 3 master eligible nodes in a cluster.

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 6:48am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/9 "2018-10-01T06:48:21Z")

</div>

Ok thanks will do.. btw thank you for your help with this!

I ran this command...

\_nodes/ES2/hot\_threads

it came back with...

[https://pastebin.com/8yCfeNsY](https://pastebin.com/8yCfeNsY)

(too much to paste here)

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 6:58am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/10 "2018-10-01T06:58:33Z")

</div>

here's an example of the load..

blue = es2  
green = es1  
purple = es3

![load](https://us1.discourse-cdn.com/elastic/original/3X/6/6/669f750fbad78dfbbd6cef48ff9e38b95cd7ba96.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 1, 2018, 7:04am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/11 "2018-10-01T07:04:53Z")

</div>

I can not see any reason for this unless one of the nodes is misconfigured, there is some hardware issue or you are using one of the features I mentioned. What does disk I/O look like on the different nodes? Is there anything in the Elasticsearch logs?

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 7:10am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/12 "2018-10-01T07:10:02Z")

</div>

here's the disk stats...

black = es2  
purple = es1  
blue = es3

 ![disk](https://us1.discourse-cdn.com/elastic/original/3X/a/d/add5da435974ebb492df6e32ee34e72ccf85d24d.png)

I'll check the logs..

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 1, 2018, 7:22am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/13 "2018-10-01T07:22:27Z")

</div>

Also check if you are having any problems with the disks.

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 8:21am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/14 "2018-10-01T08:21:16Z")

</div>

I can't see anything in the log file

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 8:56am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/15 "2018-10-01T08:56:55Z")

</div>

I found a similar issue here, could it be the version of Java? i'm using the Open version

openjdk version "1.8.0\_181"  
OpenJDK Runtime Environment (build 1.8.0\_181-b13)  
OpenJDK 64-Bit Server VM (build 25.181-b13, mixed mode)

> [@High CPU, disk load only on one node in cluster](https://discuss.elastic.co/t/high-cpu-disk-load-only-on-one-node-in-cluster/108926):
>
> Hello there. Setup that i have: Cluster of 3 nodes with ES 5.3.0. All nodes are VMs in KVM hypervisor, with no more highloaded VMs on local hypervisors. VMs: Ubuntu 16.04 linux-4.4.0 kernel (different minor builds), JVM: OpenJDK 64-Bit 1.8.0\_131+ (different minor builds). 16 vCPU, 32Gb RAM, 0 swap, 1Tb disk space. Hardware: Per hypervisor 2 Intel Xeon E5 v3, Samsung SSDs, other doesn't matter AFAIK. Cluster setup: 3 nodes, JVM heap 16Gb, x-pack installed, JMX enabled, ES config: cluste…

---

<div class="post-metadata">

**Author:** ![stuayre](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@stuayre](https://discuss.elastic.co/u/stuayre)\
**Post date:** [October 1, 2018, 9:29pm UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/16 "2018-10-01T21:29:44Z")

</div>

I switched to the Oracle Java SDK on all 3 servers and restarted Elastic Search but it made no difference...

```
ip heap.percent ram.percent cpu load_1m load_5m load_15m node.role master name
xxx.xxx.xxx.223 12 99 10 2.91 3.28 2.19 mdi - ES3
xxx.xxx.xxx.48 11 99 16 39.08 35.10 20.84 mdi - ES2
xxx.xxx.xxx.77 13 99 15 9.01 9.31 5.75 mdi * ES1

```

So I shut down ES2 to see what would happen...

```
ip heap.percent ram.percent cpu load_1m load_5m load_15m node.role master name
xxx.xxx.xxx.223 12 99 28 9.80 8.25 5.25 mdi - ES3
xxx.xxx.xxx.77 13 99 62 39.81 31.41 17.69 mdi * ES1

```

All the load jumped to ES1!

so I shut down ES1....

```
ip heap.percent ram.percent cpu load_1m load_5m load_15m node.role master name
xxx.xxx.xxx.223 11 99 28 10.14 11.29 9.00 mdi * ES3

```

Now everything is running fine and stable just on a one node cluster (ES3)!

why does adding another node increase the load on that extra node?

any ideas?

---

<div class="post-metadata">

**Author:** ![siddaram\_kj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/siddaram_kj/32/46048_2.png) [@siddaram\_kj](https://discuss.elastic.co/u/siddaram_kj)\
**Post date:** [October 5, 2018, 9:43am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/17 "2018-10-05T09:43:01Z")

</div>

Can you check in which node the shards of the index to which you are indexing the data are allocated?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 2, 2018, 9:43am UTC](https://discuss.elastic.co/t/all-load-is-being-concentrated-on-one-node/150488/18 "2018-11-02T09:43:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
