# Elastic cluster is getting overloaded by incorrect shard allocation

**URL:** <https://discuss.elastic.co/t/elastic-cluster-is-getting-overloaded-by-incorrect-shard-allocation/323595>\
**Category:** Elasticsearch\
**Created:** [January 20, 2023, 1:29pm UTC](https://discuss.elastic.co/t/elastic-cluster-is-getting-overloaded-by-incorrect-shard-allocation/323595 "2023-01-20T13:29:32Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Petr.Simik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/petr.simik/32/38082_2.png) [@Petr.Simik](https://discuss.elastic.co/u/Petr.Simik)\
**Post date:** [January 20, 2023, 1:29pm UTC](https://discuss.elastic.co/t/elastic-cluster-is-getting-overloaded-by-incorrect-shard-allocation/323595/1 "2023-01-20T13:29:32Z")

</div>

Hi Community,  
May I ask for help:

We have elastic v7.17.0 with 43 nodes. Sizing 2TB SSD, 8cores, 32GB RAM, 16GB Heap.  
Currently about 2400 indices and 6400shards.

some nodes have less shards but have disk full which causes node cpu at 100% and it impacts whole cluster that all ingest pipelines are being rejected with error

```auto
{'error': {'root_cause': [{'type': 'es_rejected_execution_exception', 'reason': 'rejected execution of coordinating operation [coordinating_and_primary_bytes=750630256, replica_bytes=0, all_bytes=750630256, 
coordinating_operation_bytes=1477151, max_coordinating_and_primary_bytes=751619276]'}], 'type': 'es_rejected_execution_exception', 'reason': 'rejected execution of coordinating operation [coordinating_and_primary_bytes=750630256, replica_bytes=0, all_bytes=750630256, coordinating_operation_bytes=1477151, max_coordinating_and_primary_bytes=751619276]'}, 'status': 429}

```

shards are automatically distributed by elastic among 43 nodes, all indices are timeseries indices having templates and ILM doing rollover daily/weekly/monthly depends on size of datasource.

there is a mix of very small indices xxMB size and huge indices xxGB size. We keep recommended defaults 50GB shard max size, but some datasources are very small .

Combination of small and large indices is a problem for elastic it happened that it allocates small indices on one node and huge end up on another node and it results in situation where some nodes have full storage and other are half empty byt Elatic allocates data to nodes with least number of shards.

My workaround is to remove node from cluster and reinsert it back in few hours later.

The problem is the situation occurs repeatedly every few days and it breaks the production.

```auto
GET _cat/allocation?v&s=node

shards disk.indices disk.used disk.avail disk.total disk.percent node
   136 1.6tb 1.6tb 232.2gb 1.9tb 88 tela01prahkz --> problem node smalles num of shard (136) but disk full
   134 1.6tb 1.7tb 201.8gb 1.9tb 89 tela02prahkz --> problem node smalles num of shard (134) but disk full
   179 643.8gb 736gb 1.2tb 1.9tb 37 tela03prahkz
   179 1tb 1.1tb 822.2gb 1.9tb 58 tela04prahkz
   179 738.2gb 836.4gb 1.1tb 1.9tb 42 tela05prahkz
....

```

thank you for any advice

I already reported this issue before but no resolution:

> [@Elasticseach shards allocation](https://discuss.elastic.co/t/elasticseach-shards-allocation/318770):
>
> Hi, I have elastic version 7.17.0 and I have a problem with shard allocation, which causes problems with data loading. The cluster is rejecting requests and limiting ingest due to overloading 2 nodes. The nodes are overloaded due to a small number of shards, but with a large size they are filling up the disk and affecting the cluster operation. Do you have any advice what I have to check/do? I temporarily fixed the problem by excluding nodes from cluster and including them back but I am af…

> [@Shard allocation on single node causes cluster overload](https://discuss.elastic.co/t/shard-allocation-on-single-node-causes-cluster-overload/279794):
>
> Cluster creates index with 5 shards on same node. And it picks the node with least disk space. High performance ETL process writes all the data on to single node instead of to 5 multiple nodes. This causes overloads of this node CPU (100%) and whole cluster becomes in high latency/ unavailable. (writing a single document lasts several seconds) The problem is related to cluster shard allocation. When new index is created via ILM it creates all shards on this node. I tried to temporarily ex…

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 20, 2023, 1:50pm UTC](https://discuss.elastic.co/t/elastic-cluster-is-getting-overloaded-by-incorrect-shard-allocation/323595/2 "2023-01-20T13:50:43Z")

</div>

Try upgrading to 8.6, noting that the [release blog](https://www.elastic.co/blog/whats-new-elasticsearch-kibana-cloud-8-6-0) calls out some improvements to shard balancing in this version:

> [@](#):
>
> # Better balancing of shards
> 
> [...]
> 
> Furthermore, we introduce two additional variables into the [balancing computation](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-cluster.html). In earlier versions, shards were all considered to be equivalent and were balanced among the nodes by shard count only. From 8.6, [shard size will also be taken into account by default](https://github.com/elastic/elasticsearch/pull/91561) to achieve a more uniform disk usage. This avoids larger shards concentrating on particular nodes, which can cause disk watermark hotspots. All Elastic Cloud and all self-managed Enterprise-subscription users will further benefit from the shard balancing strategy [based on observed data stream write load](https://github.com/elastic/elasticsearch/pull/91425). This will reduce indexing hotspots by spreading out the shards of high-traffic data streams to improve the balance of CPU usage across the hot nodes.

---

<div class="post-metadata">

**Author:** ![Petr.Simik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/petr.simik/32/38082_2.png) [@Petr.Simik](https://discuss.elastic.co/u/Petr.Simik)\
**Post date:** [January 20, 2023, 2:12pm UTC](https://discuss.elastic.co/t/elastic-cluster-is-getting-overloaded-by-incorrect-shard-allocation/323595/3 "2023-01-20T14:12:55Z")

</div>

thank you this is definitelly my plan

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 17, 2023, 2:13pm UTC](https://discuss.elastic.co/t/elastic-cluster-is-getting-overloaded-by-incorrect-shard-allocation/323595/4 "2023-02-17T14:13:15Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
