# Cluster shards unbalanced and keep moving shards around after upgrade to 8.8.1

**URL:** <https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573>\
**Category:** Elasticsearch\
**Created:** [June 21, 2023, 12:05pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573 "2023-06-21T12:05:56Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2023, 12:05pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/1 "2023-06-21T12:05:56Z")

</div>

Hello,

Yesterday we upgraded our cluster from 8.5.1 to 8.8.1 and now the shards are unbalacend between the nodes and the cluster keeps moving shards around to try to balance it.

I have a hot/warm architecture with 4 hot nodes and 12 warm nodes with just ~ 50 TB of data in total, even after 12 hours it is still moving shards around.

This is how some of the warm nodes looks like:

 ![Screenshot from 2023-06-21 09-00-41](https://us1.discourse-cdn.com/elastic/original/3X/3/9/3966b21fa65d8d62856d48e0a02821562d1c1454.png)

All warm nodes have the same specs.

One big change that we got on this upgrade is that on 8.6 the size of the shards started to being considered on the way elastic balance shards, could this be an issue related to this?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2023, 12:21pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/2 "2023-06-21T12:21:44Z")

</div>

Just found this [github issue](https://github.com/elastic/elasticsearch/issues/87279) and I think that it is related to my issue.

I have custom settings on 3 of the 4 settings mentioned in the workaround part of the issue:

```auto
cluster.routing.allocation.cluster_concurrent_rebalance
cluster.routing.allocation.node_concurrent_incoming_recoveries
cluster.routing.allocation.node_concurrent_outgoing_recoveries

```

Will set them to the default values to see if it helps.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 21, 2023, 12:37pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/3 "2023-06-21T12:37:26Z")

</div>

Yes please use the defaults for these settings.

Also it looks like you have shards of quite different sizes, so when you upgrade we'd expect quite a bit of shard movement to try and balance the disk usage better. Rebalancing does not happen as fast as possible by design (it's a background process which we don't want to affect the regular operation of the cluster) so it could take quite some time to rebalance 50TiB of data.

`GET _internal/desired_balance` is an internal API but it should be possible for you to watch the balancing progress here.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2023, 12:51pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/4 "2023-06-21T12:51:23Z")

</div>

Hello @DavidTurner,

With the change made in 8.6 that now takes the disk usage as a factor when rebalacing shard, should I expect an uneven number of shards on the nodes of the same tier in the case of shards of different sizes?

On 8.5.1 we had an almost even number of shards on each node, but the disk usage was pretty diffrent, some nodes had for example 1 TB of free space and other only 180 GB of free space, which was an issue for us, so after this upgrade should I expect a more even use of disk space but an uneven number of shards on each node?

> [@DavidTurner](#):
>
> `GET _internal/desired_balance` is an internal API but it should be possible for you to watch the balancing progress here.

What should I look on this API as there is no public documentation about it? I'm not sure how to interpret the response.

For example, this is part of the result of the API, without the nodes and shards information and in my case `data_content` tier is on the `data_warm` nodes as well:

```auto
{
  "stats": {
    "computation_converged_index": 1010,
    "computation_active": false,
    "computation_submitted": 1011,
    "computation_executed": 1011,
    "computation_converged": 1004,
    "computation_iterations": 6905,
    "computed_shard_movements": 17927,
    "computation_time_in_millis": 203993,
    "reconciliation_time_in_millis": 4806
  },
  "cluster_balance_stats": {
    "tiers": {
      "data_warm": {
        "shard_count": {
          "total": 2306,
          "min": 169,
          "max": 212,
          "average": 192.16666666666666,
          "std_dev": 15.03237247483664
        },
        "forecast_write_load": {
          "total": 0,
          "min": 0,
          "max": 0,
          "average": 0,
          "std_dev": 0
        },
        "forecast_disk_usage": {
          "total": 46521804262408,
          "min": 3643675442782,
          "max": 3977820690712,
          "average": 3876817021867.3335,
          "std_dev": 98262085739.02898
        },
        "actual_disk_usage": {
          "total": 46521804262408,
          "min": 3643675442782,
          "max": 3977820690712,
          "average": 3876817021867.3335,
          "std_dev": 98262085739.02898
        }
      },
      "data_hot": {
        "shard_count": {
          "total": 1166,
          "min": 286,
          "max": 297,
          "average": 291.5,
          "std_dev": 5.024937810560445
        },
        "forecast_write_load": {
          "total": 0,
          "min": 0,
          "max": 0,
          "average": 0,
          "std_dev": 0
        },
        "forecast_disk_usage": {
          "total": 12949387024675,
          "min": 3117120702163,
          "max": 3368843902592,
          "average": 3237346756168.75,
          "std_dev": 110375675929.2402
        },
        "actual_disk_usage": {
          "total": 12949387024675,
          "min": 3117120702163,
          "max": 3368843902592,
          "average": 3237346756168.75,
          "std_dev": 110375675929.2402
        }
      },
      "data_content": {
        "shard_count": {
          "total": 2306,
          "min": 169,
          "max": 212,
          "average": 192.16666666666666,
          "std_dev": 15.03237247483664
        },
        "forecast_write_load": {
          "total": 0,
          "min": 0,
          "max": 0,
          "average": 0,
          "std_dev": 0
        },
        "forecast_disk_usage": {
          "total": 46521804262408,
          "min": 3643675442782,
          "max": 3977820690712,
          "average": 3876817021867.3335,
          "std_dev": 98262085739.02898
        },
        "actual_disk_usage": {
          "total": 46521804262408,
          "min": 3643675442782,
          "max": 3977820690712,
          "average": 3876817021867.3335,
          "std_dev": 98262085739.02898
        }
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2023, 1:07pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/5 "2023-06-21T13:07:29Z")

</div>

I think that now that the disk usage is take into account, the number of shards in each tier will not always be evenly distributed when the shard size is different, which is now kinda obvious and expected.

For example, these are my 4 hot nodes:

 ![Screenshot from 2023-06-21 10-01-12](https://us1.discourse-cdn.com/elastic/original/3X/5/9/59045b0fee33d21322b3a81f184b2139ee2d3256.png)

They are in this state for a couple of time now and there is no shard movement between them, so it seems that this is the balanced state of my hot tier now, which is nice since the disk usage is pretty similar and the specs are all the same for them, 4 TB of disk on each.

So, answering my own question, it is now obvious that an uneven number of shards on each node of a specific tier is expected as the disk usage is now taken into consideration.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 21, 2023, 1:10pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/6 "2023-06-21T13:10:00Z")

</div>

The `routing_table` section has an entry per shard, with a `node_is_desired` flag. The ones with `false` here are still to be moved.

> [@leandrojmp](#):
>
> So, answering my own question, it is now obvious that an uneven number of shards on each node of a specific tier is expected as the disk usage is now taken into consideration.

Yes that's right. Elasticsearch still cares about shard count, but it now aims to strike a balance between even shard count and even disk usage and in doing so will not normally achieve perfection in either variable.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 21, 2023, 1:17pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/7 "2023-06-21T13:17:10Z")

</div>

> [@DavidTurner](#):
>
> The ones with `false` here are still to be moved.

Oh thanks, will write a quick script to get this response and print the shards to be moved in a more friendly way.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 27, 2023, 1:04pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/8 "2023-06-27T13:04:29Z")

</div>

> [@DavidTurner](#):
>
> Elasticsearch still cares about shard count, but it now aims to strike a balance between even shard count and even disk usage and in doing so will not normally achieve perfection in either variable.

I'm using the default settings for the `cluster.routing.allocation.*` but my cluster now is constantly moving shards around, it seems that it never reaches a balanced state on my warm nodes.

It seems that a moved shard to a node is triggering another rebalance, so since I upgraded to 8.8.1 it keeps moving shards around.

Is there anything I can change to make it stop moving shards around? Is this normal now? It increases the network traffic between the nodes.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 27, 2023, 1:40pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/9 "2023-06-27T13:40:11Z")

</div>

Did you manage to analyse the output from `GET _internal/desired_balance`? How many shards with `node_is_desired: false` do you have? And is this remaining constant over time?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 27, 2023, 2:17pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/10 "2023-06-27T14:17:14Z")

</div>

> [@DavidTurner](#):
>
> How many shards with `node_is_desired: false` do you have? And is this remaining constant over time?

I'm not monitoring this API, so I cannot say if the value is constant over time.

Just checked here on a couple of intervals, it had _822_ shards with `false`, then _816_ and then _821_ and after a couple more of minutes it increased to _831_.

I could build a quick script to check this API and get the number of shards with `node_is_desired` as false and run it every 5 minutes, but I will wait a couple more days to see if the issue still persists.

One thing that I need to add is that we have a hot/warm architecture and we are using daily indices with ILM configured, so every day shards will be moved from the hot nodes to the warm nodes and indices will be deleted, this of course will trigger a rebalance.

When the balance of shards didn't take the disk usage in consideration, every node would get a equal number of shards and the balance process would finish pretty quickly, but now it seems that the warm nodes spend all day rebalancing shards and the process doesn't finish before the next ILM trigger, which will then force another rebalance.

I will watch this for a couple more days, but is there any way to go back to the old behavior? Constantly moving shards is not desired because the network load.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 27, 2023, 2:30pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/11 "2023-06-27T14:30:36Z")

</div>

Hmm it seems surprising to have 800+ of your ~2300 shards in the wrong place. Could you try `DELETE /_internal/desired_balance` and see if that gets things to settle down more quickly?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 27, 2023, 2:45pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/12 "2023-06-27T14:45:13Z")

</div>

> [@DavidTurner](#):
>
> DELETE /\_internal/desired\_balance

This returns this:

> Request failed to get to the server (status code: 200)

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 27, 2023, 2:46pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/13 "2023-06-27T14:46:39Z")

</div>

I don't recognise that message - it's not coming from Elasticsearch.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 27, 2023, 2:50pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/14 "2023-06-27T14:50:42Z")

</div>

> [@DavidTurner](#):
>
> I don't recognise that message - it's not coming from Elasticsearch

I'm running on Kibana Dev Tools.

Our Kibana is behind a GCP load balancing, maybe is from there, let me try to run it directly on a Elasticsearch node.

EDIT:

Running from an Elasticsearch node it seems to have worked, but there is no return for the curl command for that endpoint, is that right?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 27, 2023, 2:55pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/15 "2023-06-27T14:55:18Z")

</div>

> there is no return for the curl command for that endpoint, is that right?

That's right. Does `GET _internal/desired_balance` look any better now?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 27, 2023, 2:57pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/16 "2023-06-27T14:57:23Z")

</div>

> [@DavidTurner](#):
>
> Does `GET _internal/desired_balance` look any better now?

Yeah, just got the response to a `json` file and got this:

```auto
$ cat desired-balance.json | grep "node_is_desired" | grep -v relocating | grep false | wc -l
53

```

It is showing 53 now, but a second check it increased to 57 and then to 60.

I will wait a couple of time to see if it starts increasing.

After 2 hours the number increased to 97.

Not sure if this solved anything, after a couple more hours there are now 280 shards with `node_is_desired` as `false`.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 28, 2023, 2:13pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/17 "2023-06-28T14:13:23Z")

</div>

Hello @DavidTurner,

There was another issue, a couple of the nodes were hitting the watermarks and others were close to it, so probably any shard movement could trigger another watermark that would then trigger another shard movement.

I removed some old data to make sure that no node in the warm tier would hit any watermark, and after sometime there are no more shard movements.

```auto
$ cat desired-balance.json | grep "node_is_desired" | grep -v relocating | grep false | wc -l
0

```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 28, 2023, 3:27pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/18 "2023-06-28T15:27:07Z")

</div>

Hmm, that is a little puzzling. The new allocator should let you run nodes closer to full because of its disk balancing.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [June 28, 2023, 3:56pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/19 "2023-06-28T15:56:20Z")

</div>

> [@DavidTurner](#):
>
> The new allocator should let you run nodes closer to full because of its disk balancing.

I'm not sure if I'm right, but after I removed some old data to free more space on the nodes, the shards stopped moving around,

Before that the number of shards with `node_is_desired` as `false` was increasing, last time I checked was around _600_, and from the 12 warm nodes, 2 have hit the second watermark already, the high watermark.

After I freed somespace on the warm tier, things went back to normal.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 26, 2023, 3:56pm UTC](https://discuss.elastic.co/t/cluster-shards-unbalanced-and-keep-moving-shards-around-after-upgrade-to-8-8-1/336573/20 "2023-07-26T15:56:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
