# Shards balance and homogeneity of their sizes

**URL:** <https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661>\
**Category:** Elasticsearch\
**Created:** [August 17, 2021, 10:09am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661 "2021-08-17T10:09:57Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![sebastienf](https://avatars.discourse-cdn.com/v4/letter/s/f07891/32.png) [@sebastienf](https://discuss.elastic.co/u/sebastienf)\
**Post date:** [August 17, 2021, 10:09am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/1 "2021-08-17T10:09:57Z")

</div>

I understood the need of a well balanced ES cluster, with indices made of enough shards (not too much but not too few), but I was wondering if the constancy of the size of shards was important ?  
Indeed, according the purpose of the index, I got yearly, monthly, daily ones leading to very different sizes of shard : is this a problem at the end ?

---

<div class="post-metadata">

**Author:** ![Wolfram\_Haussig](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wolfram_haussig/32/70528_2.png) [@Wolfram\_Haussig](https://discuss.elastic.co/u/Wolfram_Haussig)\
**Post date:** [August 19, 2021, 7:40am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/2 "2021-08-19T07:40:16Z")

</div>

Hello Sebastien,

We already had a problem with this so I can safely say that - yes, this can be a problem! Elasticsearch balances the shards on all nodes based on the shard count which makes sense as the best practices say 20 shards per 1GB of RAM.

We have a cluster 3 equally large nodes(cpu,RAM, storage) but our shard sizes were very different: some were 25GB per shard and some were only 1GB per shard(I know...)

In the end, one node created an alert because the storage was nearly full(\>80%) while the other 2 nodes had still \>40% free.

The best way to solve this is to make all shards more or less the same size. Otherwise - and that was our temporary workaround then - you can take a look at the [index shard allocation](https://www.elastic.co/guide/en/elasticsearch/reference/current/shard-allocation-filtering.html). We used the allocation settings to pin the different large indices to specific nodes within the cluster to better distribute the load.

Best regards  
Wolfram

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 19, 2021, 8:06am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/3 "2021-08-19T08:06:38Z")

</div>

> [@Wolfram\_Haussig](#):
>
> In the end, one node created an alert because the storage was nearly full(\>80%) while the other 2 nodes had still \>40% free.

The problem here seems to be that you set up an overly sensitive alert. Elasticsearch doesn't mind having imbalanced storage like this, really you should only alert if you're approaching the low watermark (85% by default) on all nodes or if one node is _persistently_ over the high watermark (90% by default). Outside of those conditions, no action is really needed so an alert is kind of inappropriate. Moving shards around is expensive (it blows out the filesystem cache for instance) so it's usually preferable to leave things alone.

See [these docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-cluster.html#disk-based-shard-allocation) for more details, in particular:

> **NOTE** : It is normal for nodes to temporarily exceed the high watermark from time to time.

and

> **TIP** : It is normal for the nodes in your cluster to be using very different amounts of disk space. ..

---

<div class="post-metadata">

**Author:** ![Wolfram\_Haussig](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wolfram_haussig/32/70528_2.png) [@Wolfram\_Haussig](https://discuss.elastic.co/u/Wolfram_Haussig)\
**Post date:** [August 19, 2021, 8:13am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/4 "2021-08-19T08:13:22Z")

</div>

> [@DavidTurner](#):
>
> The problem here seems to be that you set up an overly sensitive alert

I am not aware that we changed the alert settings so I guess this is the alert out of the box...

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 19, 2021, 9:52am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/5 "2021-08-19T09:52:21Z")

</div>

> [@Wolfram\_Haussig](#):
>
> I guess this is the alert out of the box...

Where did the box in question come from? I don't think Elasticsearch itself ships with any such alert. If it does, that's definitely a bug.

---

<div class="post-metadata">

**Author:** ![Wolfram\_Haussig](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wolfram_haussig/32/70528_2.png) [@Wolfram\_Haussig](https://discuss.elastic.co/u/Wolfram_Haussig)\
**Post date:** [August 19, 2021, 10:08am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/6 "2021-08-19T10:08:28Z")

</div>

> [@DavidTurner](#):
>
> Where did the box in question come from? I don't think Elasticsearch itself ships with any such alert. If it does, that's definitely a bug.

The installation comes directly from Elastic, the alert is even well [documented](https://www.elastic.co/guide/en/kibana/current/kibana-alerts.html#kibana-alerts-disk-usage-threshold):

> This rule checks for Elasticsearch nodes that are nearly at disk capacity. By default, the condition is set at 80% or more averaged over the last 5 minutes. The default rule checks on a schedule time of 1 minute with a re-notify interval of 1 day.

As long as alerts provided by Elastic are firing I will handle them as an urgent problem...

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 19, 2021, 11:06am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/7 "2021-08-19T11:06:18Z")

</div>

TIL, I did not know this was one of our built-in alerts. I have reported this as a bug:

> <https://github.com/elastic/kibana/issues/109224>
>
> \*\*Kibana version:\*\* 7.14.0 (likely others)
> 
> \*\*Elasticsearch version:\*\* 7.14.0 …(likely others)
> 
> \*\*Describe the bug:\*\*
> 
> The default \[disk usage threshold alert\](https://www.elastic.co/guide/en/kibana/current/kibana-alerts.html#kibana-alerts-disk-usage-threshold) will fire when a single node reaches 80% capacity even though no action is needed at this point. Elasticsearch itself doesn't react at all until disk usage reaches the low watermark (85% capacity by default) and only really starts putting any effort in when it reaches the high watermark (90% capacity by default); intervention from a user is only needed once Elasticsearch runs out of options for moving shards around. See \[these Elasticsearch docs\](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-cluster.html#disk-based-shard-allocation) for more info, noting in particular:
> 
> \> \*\*NOTE\*\*: It is normal for nodes to temporarily exceed the high watermark from time to time.
> 
> \*\*Steps to reproduce:\*\*
> 1. Install Elasticsearch & Kibana with default settings
> 2. Fill one of the disks up to over 80% capacity
> 3. Note that the disk usage threshold alert fires even though Elasticsearch logs no warnings about disk usage.
> 
> \*\*Expected behavior:\*\*
> 
> The alert should only fire when action is needed from the user, which in this context means a node is persistently over its high watermark, or the cluster in total is approaching its low watermark capacity. We shouldn't be firing an alert for a situation that Elasticsearch considers to be normal.
> 
> \*\*Any additional context:\*\*
> 
> Raised in https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661

---

<div class="post-metadata">

**Author:** ![sebastienf](https://avatars.discourse-cdn.com/v4/letter/s/f07891/32.png) [@sebastienf](https://discuss.elastic.co/u/sebastienf)\
**Post date:** [August 20, 2021, 10:29am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/8 "2021-08-20T10:29:42Z")

</div>

Thanks @Wolfram_Haussig and @DavidTurner for your share, it helps me.  
@DavidTurner, does a well-balanced-shards-cluster can be slowed down (read and write) because shards storage are imbalanced ?

Best regards  
Sebastien

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 17, 2021, 10:29am UTC](https://discuss.elastic.co/t/shards-balance-and-homogeneity-of-their-sizes/281661/9 "2021-09-17T10:29:43Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
