# The es-cluster warm-node rebalance tasks keeps going and some node shards decrease for long time

**URL:** <https://discuss.elastic.co/t/the-es-cluster-warm-node-rebalance-tasks-keeps-going-and-some-node-shards-decrease-for-long-time/306074>\
**Category:** Elasticsearch\
**Created:** [June 1, 2022, 2:33am UTC](https://discuss.elastic.co/t/the-es-cluster-warm-node-rebalance-tasks-keeps-going-and-some-node-shards-decrease-for-long-time/306074 "2022-06-01T02:33:12Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![yitiao\_feiyu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yitiao_feiyu/32/106414_2.png) [@yitiao\_feiyu](https://discuss.elastic.co/u/yitiao_feiyu)\
**Post date:** [June 1, 2022, 2:33am UTC](https://discuss.elastic.co/t/the-es-cluster-warm-node-rebalance-tasks-keeps-going-and-some-node-shards-decrease-for-long-time/306074/1 "2022-06-01T02:33:12Z")

</div>

version ：7.17  
The es-cluster warm-node rebalance tasks keeps going and not stop, warm node 's index comes from hot-node by ILM policy control, I didn't update warm index directly. I hava snapshot task, but I think it's not the cause of rebalane keeps goling.  
I tried reduce or enlarge the cluster\_concurrent\_rebalance and wait but the reblanace is still didn't stop.  
Thre wired thing is the warm-20 and warm-15's disk space and shard keep decrease for 8 hours , snapshot bellow:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/7/b7cbf119ff3ba078d11c41d1110815cd2dfd78f2.jpeg)

I open the allocator trace log, found that warm-20 as move source node , the index weight is more than 30(it's not resonable), but warm-20 as move target node is negative or less then 5, I calculate myself the weight 30+ mabe wrong. the warm-20's shards is very less compire other's warm node. I think bad weight value is the cause of rebalancing issue. so what is the reason behind this? see below logging snapshot.  
`node_id "WvAX8WYTQiGIYdFdosEBjw" is the name of "warm-20"`

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/a/3/a34b142dce8be192adbad5076f20e038fcd19301.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/4/748b25f28c16816abe1d26bf2f84545642669aab.png)

> > code ： /elasticsearch-7.17.3-sources.jar!/org/elasticsearch/cluster/routing/allocation/allocator/BalancedShardsAllocator.java:552  
> > private void balanceByWeights() { ...."Balancing from node [{}] weight: [{}] to node [{}] weight: [{}] delta: [{}]",

**my cluster setting**  
GET /\_cluster/settings?include\_default  
"persistent" : {  
"cluster.routing.allocation.allow\_rebalance" : "indices\_all\_active",  
"cluster.routing.allocation.balance.threshold" : "3",  
"cluster.routing.allocation.cluster\_concurrent\_rebalance" : "2",  
"cluster.routing.allocation.enable" : "all"  
},  
"transient" : {  
"cluster.routing.allocation.balance.threshold" : "2.0",  
"cluster.routing.allocation.cluster\_concurrent\_rebalance" : "24",  
"cluster.routing.allocation.disk.watermark.flood\_stage" : "99%",  
"cluster.routing.allocation.disk.watermark.high" : "94%",  
"cluster.routing.allocation.disk.watermark.low" : "94%",  
....  
**current allocation info**  
GET \_cat/allocation?v=true&&s=shards&h=node,shards,disk.\*  
node shards disk.indices disk.used disk.avail disk.total disk.percent  
logging-hot-11 150 1.1tb 1.2tb 492gb 1.7tb 72  
logging-hot-12 151 1.2tb 1.3tb 337.6gb 1.7tb 80  
logging-hot-8 151 1.2tb 1.3tb 400.8gb 1.7tb 77  
logging-hot-19 151 1.1tb 1.1tb 533.1gb 1.7tb 69  
logging-hot-15 151 1tb 1.1tb 583.6gb 1.7tb 66  
logging-hot-18 151 973.3gb 1tb 691.6gb 1.7tb 60  
logging-hot-1 152 1.2tb 1.3tb 381.1gb 1.7tb 78  
logging-hot-9 152 1009.1gb 1tb 657.4gb 1.7tb 62  
logging-hot-5 152 1.1tb 1.1tb 533.6gb 1.7tb 69  
logging-hot-17 152 1019.2gb 1tb 646.2gb 1.7tb 63  
logging-hot-2 152 1.1tb 1.1tb 534gb 1.7tb 69  
logging-hot-0 152 1.1tb 1.2tb 513.6gb 1.7tb 70  
logging-hot-7 152 984gb 1tb 680.6gb 1.7tb 61  
logging-hot-14 153 1.1tb 1.1tb 533.6gb 1.7tb 69  
logging-hot-3 153 1.1tb 1.2tb 529.7gb 1.7tb 69  
logging-hot-16 153 1tb 1.1tb 619.1gb 1.7tb 64  
logging-hot-6 154 1.1tb 1.2tb 503.5gb 1.7tb 71  
logging-hot-10 154 1.1tb 1.1tb 531.8gb 1.7tb 69  
logging-hot-4 154 1tb 1.1tb 575.2gb 1.7tb 67  
logging-hot-13 154 1tb 1.1tb 608.9gb 1.7tb 65  
logging-warm-20 165 1.7tb 2.1tb 4.9tb 7tb 29  
logging-warm-15 205 3tb 3.3tb 3.7tb 7tb 47  
logging-warm-8 292 5.5tb 5.9tb 1.1tb 7tb 83  
logging-warm-2 294 5.2tb 5.6tb 1.4tb 7tb 79  
logging-warm-12 305 5.6tb 5.9tb 1.1tb 7tb 84  
logging-warm-4 306 5.7tb 6tb 1003.8gb 7tb 86  
logging-warm-22 307 5.6tb 6tb 1tb 7tb 85  
logging-warm-9 310 5.4tb 5.8tb 1.2tb 7tb 82  
logging-warm-7 311 5.8tb 6.2tb 879.2gb 7tb 87  
logging-warm-23 311 5.8tb 6.2tb 882.8gb 7tb 87  
logging-warm-14 312 5.7tb 6.1tb 992.5gb 7tb 86  
logging-warm-10 315 5.3tb 5.7tb 1.2tb 7tb 81  
logging-warm-18 315 5.4tb 5.8tb 1.2tb 7tb 82  
logging-warm-11 315 5.4tb 5.8tb 1.2tb 7tb 82  
logging-warm-6 316 5.2tb 5.5tb 1.4tb 7tb 79  
logging-warm-16 317 5.1tb 5.4tb 1.5tb 7tb 77  
logging-warm-17 317 5.2tb 5.6tb 1.4tb 7tb 79  
logging-warm-3 317 5.6tb 5.9tb 1.1tb 7tb 84  
logging-warm-19 318 4.5tb 4.8tb 2.1tb 7tb 69  
logging-warm-0 318 4.7tb 5.1tb 1.9tb 7tb 72  
logging-warm-21 318 5.4tb 5.7tb 1.2tb 7tb 82  
logging-warm-5 319 5.2tb 5.6tb 1.4tb 7tb 80  
logging-warm-13 320 5.5tb 5.9tb 1.1tb 7tb 83  
logging-warm-1 321 5.6tb 5.9tb 1tb 7tb 84  
**warm-disk-usage trend**

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/0/a/0a75fe1857c754c49337439aa1bcfd1d5ee8aec9.jpeg)

---

<div class="post-metadata">

**Author:** ![yitiao\_feiyu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yitiao_feiyu/32/106414_2.png) [@yitiao\_feiyu](https://discuss.elastic.co/u/yitiao_feiyu)\
**Post date:** [June 1, 2022, 2:48am UTC](https://discuss.elastic.co/t/the-es-cluster-warm-node-rebalance-tasks-keeps-going-and-some-node-shards-decrease-for-long-time/306074/2 "2022-06-01T02:48:45Z")

</div>

I calculate warm-20 index weight myself, it may not right. it's as a refrence should be negative not 30+ as snapshot show.

```auto
args: node_id=None, node=logging-warm-20, index=phpback_2022-05-12
----------------------
公式: weightindex= indexBalance * (shards_of_node_index_count - avg_shards_per_node_index = 0.55 * (2- 1.043) = 0.526
公式: weightShard= shardBalance * (node_shards_count- avg_shards_per_node) = 0.45 * (168- 223.652 = -25.043)
    统计 node: logging-warm-20, index: phpback_2022-05-12, node_index_weight: -24.517 

```

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/7/779d8faa2dca76dd2722b6baeec255d1a8c8be47.png)

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 1, 2022, 10:28am UTC](https://discuss.elastic.co/t/the-es-cluster-warm-node-rebalance-tasks-keeps-going-and-some-node-shards-decrease-for-long-time/306074/3 "2022-06-01T10:28:20Z")

</div>

Strangely I have just been investigating another case that looks very similar and found some strange effects when too much concurrent balancing is allowed. I opened [#87279](https://github.com/elastic/elasticsearch/issues/87279) with some more details but the short answer is "remove `cluster.routing.allocation.cluster_concurrent_rebalance` from your config".

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 29, 2022, 10:28am UTC](https://discuss.elastic.co/t/the-es-cluster-warm-node-rebalance-tasks-keeps-going-and-some-node-shards-decrease-for-long-time/306074/4 "2022-06-29T10:28:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
