# Load not evenly distributed

**URL:** <https://discuss.elastic.co/t/load-not-evenly-distributed/101343>\
**Category:** Elasticsearch\
**Created:** [September 21, 2017, 1:38pm UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343 "2017-09-21T13:38:34Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 21, 2017, 1:38pm UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/1 "2017-09-21T13:38:34Z")

</div>

Hi, I have a 3+ node setup, with all nodes having all roles.  
1 node gets up to 90% cpu and frequent garbage collection 2nd node is a bit less but reasonable and then nodes 3+ are doing nearly nothing. If I stop and start nodes a different node will get the high load.

Are there any suggestions on how I can even the load over 3+ nodes?

Its running on linux and on their own machines v5.5.2.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 21, 2017, 10:02pm UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/2 "2017-09-21T22:02:41Z")

</div>

How are you interacting with the cluster, Kibana, something else?

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 6:22am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/3 "2017-09-22T06:22:44Z")

</div>

With x-pack and kibana yes

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 22, 2017, 6:24am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/4 "2017-09-22T06:24:47Z")

</div>

Right, but what about writing to the cluster?

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 6:55am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/5 "2017-09-22T06:55:40Z")

</div>

Its used for exceptionless.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 22, 2017, 6:57am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/6 "2017-09-22T06:57:59Z")

</div>

Does it do load balancing or does it just talk to a single node?

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 6:59am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/7 "2017-09-22T06:59:29Z")

</div>

It does loadbalancing (I tried with the ip of all nodes in config as well as a round robin ip).

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 22, 2017, 7:00am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/8 "2017-09-22T07:00:00Z")

</div>

What does Monitoring tell you about differences in load/indexing/query loads?

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 7:05am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/9 "2017-09-22T07:05:09Z")

</div>

Atm its on 2 nodes because the 3rd node as just about no load and what I can see is:  
node 1 cpu: 26% node 2 cpu 80% mem for node 1 gets garbage collection to almost 0 where node 2 is 70%+ after garbage collection.  
Request rate (indexing just over 2k for both and search rate just over 1k for both)  
not sure how to check query load?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 22, 2017, 7:27am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/10 "2017-09-22T07:27:55Z")

</div>

What is the output of the [cat nodes API](https://www.elastic.co/guide/en/elasticsearch/reference/5.6/cat-nodes.html)?

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 7:38am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/11 "2017-09-22T07:38:25Z")

</div>

heap.percent ram.percent cpu load\_1m load\_5m load\_15m node.role master  
43-----------------97---------------43--1.27-------1.35---------1.28 mdi -  
45-----------------98---------------81--3.65-------3.64--------- 3.80 mdi \*

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 22, 2017, 8:06am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/12 "2017-09-22T08:06:45Z")

</div>

This may be related: [https://github.com/elastic/elasticsearch/issues/24642](https://github.com/elastic/elasticsearch/issues/24642)  
The fix is in the impending 6.0 release.

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 8:13am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/14 "2017-09-22T08:13:54Z")

</div>

Thanks, is there anything I can set to fix it on 5.5.2?

Will removing kibana (x-pack on the nodes) solve the problem (if thats even possible to do without breaking elasticsearch) or using new nodes without x-pack/kibana?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 22, 2017, 9:03am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/15 "2017-09-22T09:03:39Z")

</div>

Two ugly choices:

- Set replicas to zero to rebalance with all primaries (and no redundancy!)
- Install a proxy between Kibana and elasticsearch to strip out preference=sessionId parameters

Kibana uses sessionId based routing to ensure each user revisits the same nodes and has warm caches for their queries so is desirable but the 5.x primary vs replica selection routing logic is not ideal when using this feature.

Note however, that with many users their collective loads should be spread evenly across data nodes but a single user will load the cluster unevenly.

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 9:35am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/16 "2017-09-22T09:35:34Z")

</div>

Thank you!

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 9:46am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/17 "2017-09-22T09:46:06Z")

</div>

I dont see how it can be kibana because exceptionless connects directly to the nodes so the traffic should'nt be affected by kibana.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 22, 2017, 10:02am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/18 "2017-09-22T10:02:32Z")

</div>

You said you had 3 nodes in the cluster, but I only see two here? Have you set `minimum_master_nodes` correctly to avoid split brain scenarios?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 22, 2017, 10:06am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/19 "2017-09-22T10:06:52Z")

</div>

> [@jannesvh](#):
>
> I dont see how it can be kibana because exceptionless connects directly to the nodes

SessionID based query routing is a feature supported by elasticsearch and used by Kibana. It may also be used by exceptionless.

---

<div class="post-metadata">

**Author:** ![jannesvh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jannesvh/32/22371_2.png) [@jannesvh](https://discuss.elastic.co/u/jannesvh)\
**Post date:** [September 22, 2017, 10:14am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/20 "2017-09-22T10:14:44Z")

</div>

Yes the minimum\_master\_nodes is setup correctly. It used to be 3 nodes but I stoped the 1 because it was literately doing nothing (5% cpu and 4 garbage collection a day) and I'm aware its not an ideal setup atm, minimum\_master\_nodes is set to 2 so split brain should not be a problem.

The number of shards is set to the number of nodes so as I understand it the primary index should be split between all nodes.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 23, 2017, 7:52am UTC](https://discuss.elastic.co/t/load-not-evenly-distributed/101343/21 "2017-09-23T07:52:04Z")

</div>

> [@jannesvh](#):
>
> I'm aware its not an ideal setup atm, minimum\_master\_nodes is set to 2 so split brain should not be a problem

No, but if you lose one master you lose the cluster.

[Next page](https://discuss.elastic.co/t/load-not-evenly-distributed/101343.md?page=2)
