# Too many unassigned replica shards everyday

**URL:** <https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550>\
**Category:** Elasticsearch\
**Created:** [December 20, 2017, 5:44am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550 "2017-12-20T05:44:21Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![sbanagiri](https://avatars.discourse-cdn.com/v4/letter/s/ba9def/32.png) [@sbanagiri](https://discuss.elastic.co/u/sbanagiri)\
**Post date:** [December 20, 2017, 5:44am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/1 "2017-12-20T05:44:21Z")

</div>

Hello,

We have a ELK cluster setup with 35 nodes (3 masters and 32 data nodes).  
New indices get created everyday once.  
Approximately, 280 indices are created everyday.  
Number of shards per index is 24 with replication factor of 1.  
The disk size of each node is 15TB and RAM is 125GB.

The problem is that, everyday after the index creation, the cluster goes to yellow state with too many unassigned indices and takes a lot of time to recover. Following is the state of the cluster after 4.5 hours of new index creation.

> {  
> "cluster\_name" : \<cluster\_name\>,  
> "status" : "yellow",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 35,  
> "number\_of\_data\_nodes" : 32,  
> "active\_primary\_shards" : 48589,  
> "active\_shards" : 89434,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 30,  
> "unassigned\_shards" : 7760,  
> "delayed\_unassigned\_shards" : 0,  
> "number\_of\_pending\_tasks" : 1936,  
> "number\_of\_in\_flight\_fetch" : 0,  
> "task\_max\_waiting\_in\_queue\_millis" : 2036819,  
> "active\_shards\_percent\_as\_number" : 91.98757508434132  
> }

The reasons for shards being unassigned are the following:

> INDEX\_CREATED - on newly created indices  
> NODE\_LEFT - on older indices

The following is the output from \_cat/allocation/explain api

> .  
> .  
> .  
> .  
> {  
> "node\_id" : "XZAy1ITXSjmO\_a-8yulfig",  
> "node\_name" : "node1",  
> "transport\_address" : ":9300",  
> "node\_attributes" : {  
> "ml.max\_open\_jobs" : "10",  
> "ml.enabled" : "true"  
> },  
> "node\_decision" : "throttled",  
> "deciders" : [  
> {  
> "decider" : "throttling",  
> "decision" : "THROTTLE",  
> "explanation" : "reached the limit of incoming shard recoveries [10], cluster setting [cluster.routing.allocation.node\_concurrent\_incoming\_recoveries=10] (can also be set via [cluster.routing.allocation.node\_concurrent\_recoveries])"  
> }  
> ]  
> },  
> {  
> "node\_id" : "dnB6TExcSOaI\_poL3d3t9g",  
> "node\_name" : "node2",  
> "transport\_address" : ":9300",  
> "node\_attributes" : {  
> "ml.max\_open\_jobs" : "10",  
> "ml.enabled" : "true"  
> },  
> "node\_decision" : "throttled",  
> "store" : {  
> "matching\_sync\_id" : true  
> },  
> "deciders" : [  
> {  
> "decider" : "throttling",  
> "decision" : "THROTTLE",  
> "explanation" : "reached the limit of incoming shard recoveries [10], cluster setting [cluster.routing.allocation.node\_concurrent\_incoming\_recoveries=10] (can also be set via [cluster.routing.allocation.node\_concurrent\_recoveries])"  
> }  
> ]  
> },  
> {  
> "node\_id" : "BrIKoXg4TkuJ2Ekn8YKuPQ",  
> "node\_name" : "node3",  
> "transport\_address" : ":9300",  
> "node\_attributes" : {  
> "ml.max\_open\_jobs" : "10",  
> "ml.enabled" : "true"  
> },  
> "node\_decision" : "no",  
> "store" : {  
> "matching\_sync\_id" : true  
> },  
> "deciders" : [  
> {  
> "decider" : "same\_shard",  
> "decision" : "NO",  
> "explanation" : "the shard cannot be allocated to the same node on which a copy of the shard already exists [[logstash-hadoop\_axonitered-hdfsproxy-audit-s3-2017.12.15][18], node[BrIKoXg4TkuJ2Ekn8YKuPQ], [P], s[STARTED], a[id=OzSdAIySQFWKj7PuzUHbjQ]]"  
> }  
> ]  
> }

We don't want the cluster to go into yellow state everyday where thousands of shards go unassigned and takes time to recover. We want the index creation, shard distribution of both primary and replica to be smooth.Would increasing the number of shards per index help. If yes, by how much?

Also, what does NODE\_LEFT mean? I do not see any services getting restarted on nodes.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 20, 2017, 6:45am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/2 "2017-12-20T06:45:46Z")

</div>

> [@sbanagiri](#):
>
> "number\_of\_nodes" : 35,  
> "number\_of\_data\_nodes" : 32,  
> "active\_primary\_shards" : 48589,  
> "active\_shards" : 89434,

That is waaaaaaaayyyyyyyy too many. You need to reduce that quite dramatically.

---

<div class="post-metadata">

**Author:** ![zqc0512](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zqc0512/32/32141_2.png) [@zqc0512](https://discuss.elastic.co/u/zqc0512)\
**Post date:** [December 20, 2017, 7:10am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/3 "2017-12-20T07:10:10Z")

</div>

the shards set about indices how many? can change the `cluster.routing.allocation.node_initial_primaries_recoveries: big number`  
output the **pending\_tasks** see what pending..

---

<div class="post-metadata">

**Author:** ![sbanagiri](https://avatars.discourse-cdn.com/v4/letter/s/ba9def/32.png) [@sbanagiri](https://discuss.elastic.co/u/sbanagiri)\
**Post date:** [December 20, 2017, 10:33am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/4 "2017-12-20T10:33:07Z")

</div>

Following is the subset of output from the pending tasks

> {  
> "insert\_order" : 180619,  
> "priority" : "URGENT",  
> "source" : "shard-started shard id [[logstash-hadoop\_radiumtan-oozie-oozie-ops-s3-2017.12.17][5]], allocation id [2k\_o8\_i9RtO3zF6AaG\_6yA], primary term [0], message [after peer recovery]",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 12994,  
> "time\_in\_queue" : "12.9s"  
> },  
> {  
> "insert\_order" : 179771,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 438956,  
> "time\_in\_queue" : "7.3m"  
> },  
> {  
> "insert\_order" : 179758,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 461314,  
> "time\_in\_queue" : "7.6m"  
> },  
> {  
> "insert\_order" : 179772,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 438956,  
> "time\_in\_queue" : "7.3m"  
> },  
> {  
> "insert\_order" : 179757,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 461485,  
> "time\_in\_queue" : "7.6m"  
> },  
> {  
> "insert\_order" : 179751,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 462079,  
> "time\_in\_queue" : "7.7m"  
> },  
> {  
> "insert\_order" : 179763,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 460810,  
> "time\_in\_queue" : "7.6m"  
> },  
> {  
> "insert\_order" : 179765,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 458949,  
> "time\_in\_queue" : "7.6m"  
> },  
> {  
> "insert\_order" : 179796,  
> "priority" : "HIGH",  
> "source" : "shard-failed",  
> "executing" : false,  
> "time\_in\_queue\_millis" : 437708,  
> "time\_in\_queue" : "7.2m"  
> },

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 20, 2017, 10:37am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/5 "2017-12-20T10:37:03Z")

</div>

You have too many shards.

---

<div class="post-metadata">

**Author:** ![sbanagiri](https://avatars.discourse-cdn.com/v4/letter/s/ba9def/32.png) [@sbanagiri](https://discuss.elastic.co/u/sbanagiri)\
**Post date:** [December 20, 2017, 2:26pm UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/6 "2017-12-20T14:26:32Z")

</div>

Number of shards needs to be reduced per index? @warkolm

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 20, 2017, 4:33pm UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/7 "2017-12-20T16:33:49Z")

</div>

I suspect you need to reduce the number of indices as well as the number of shards. Please read [this blog post about shards and sharding practices](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) for some guidelines.

---

<div class="post-metadata">

**Author:** ![pinky.floyano](https://avatars.discourse-cdn.com/v4/letter/p/5fc32e/32.png) [@pinky.floyano](https://discuss.elastic.co/u/pinky.floyano)\
**Post date:** [December 21, 2017, 4:14pm UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/8 "2017-12-21T16:14:11Z")

</div>

Excuse me for interrupting, I have a similar problem.

[http://localhost:9200/\_cat/shards](http://localhost:9200/_cat/shards)

censos2 0 p STARTED 6608080 199.7mb 192.168.0.104 nodo2  
censos2 0 r UNASSIGNED  
.kibana 0 p STARTED 2 6.1kb 192.168.0.104 nodo2  
.kibana 0 r UNASSIGNED  
.marvel-es-data-1 0 p STARTED 4 4.5kb 192.168.0.104 nodo2  
.marvel-es-data-1 0 r UNASSIGNED  
.marvel-es-1-2017.12.21 0 p STARTED 95 147.5kb 192.168.0.104 nodo2  
.marvel-es-1-2017.12.21 0 r UNASSIGNED  
censos 3 p STARTED 26732520 760.9mb 192.168.0.105 nodo3  
censos 3 r STARTED 26732520 760.9mb 192.168.0.104 nodo2  
censos 4 p STARTED 26678994 717.3mb 192.168.0.105 nodo3  
censos 4 r STARTED 26678994 717.3mb 192.168.0.104 nodo2  
censos 1 p STARTED 26722913 688.5mb 192.168.0.105 nodo3  
censos 1 r STARTED 26722913 688.5mb 192.168.0.104 nodo2  
censos 2 p STARTED 26720830 701.6mb 192.168.0.105 nodo3  
censos 2 r STARTED 26720830 701.6mb 192.168.0.104 nodo2  
censos 0 p STARTED 26689213 695.7mb 192.168.0.105 nodo3  
censos 0 r STARTED 26689213 695.7mb 192.168.0.104 nodo2

[http://localhost:9200/\_cat/shards?h=index,shard,prirep,state,unassigned.reason](http://localhost:9200/_cat/shards?h=index,shard,prirep,state,unassigned.reason)

censos2 0 p STARTED  
censos2 0 r UNASSIGNED CLUSTER\_RECOVERED  
.kibana 0 p STARTED  
.kibana 0 r UNASSIGNED CLUSTER\_RECOVERED  
.marvel-es-data-1 0 p STARTED  
.marvel-es-data-1 0 r UNASSIGNED INDEX\_CREATED  
.marvel-es-1-2017.12.21 0 p STARTED  
.marvel-es-1-2017.12.21 0 r UNASSIGNED INDEX\_CREATED  
censos 3 p STARTED  
censos 3 r STARTED  
censos 4 p STARTED  
censos 4 r STARTED  
censos 1 p STARTED  
censos 1 r STARTED  
censos 2 p STARTED  
censos 2 r STARTED  
censos 0 p STARTED  
censos 0 r STARTED

Considering that this did not happen before, but it started a couple of weeks ago. Sometimes I solve it, first stopping the allocation of fragments and then restarting the cluster, and after that, I activate the fragment assignment again.

But when every day the marvel indexes are automatically created again, the same thing happens, and it is tedious to have to do the same and the same thing every day, and sometimes I have to do it several times a day for this problem, someone knows the final solution so that this does not happen again?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 22, 2017, 3:06am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/9 "2017-12-22T03:06:54Z")

</div>

It'd be better if you created a new thread please 🙂

---

<div class="post-metadata">

**Author:** ![pinky.floyano](https://avatars.discourse-cdn.com/v4/letter/p/5fc32e/32.png) [@pinky.floyano](https://discuss.elastic.co/u/pinky.floyano)\
**Post date:** [December 22, 2017, 12:47pm UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/10 "2017-12-22T12:47:35Z")

</div>

Lo probaré. Gracias

---

<div class="post-metadata">

**Author:** ![sbanagiri](https://avatars.discourse-cdn.com/v4/letter/s/ba9def/32.png) [@sbanagiri](https://discuss.elastic.co/u/sbanagiri)\
**Post date:** [January 9, 2018, 9:33am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/11 "2018-01-09T09:33:51Z")

</div>

Reduced the number of indices and the number of shards.

> {  
> "cluster\_name": "xxx",  
> "status": "green",  
> "timed\_out": false,  
> "number\_of\_nodes": 35,  
> "number\_of\_data\_nodes": 32,  
> "active\_primary\_shards": 4184,  
> "active\_shards": 8414,  
> "relocating\_shards": 0,  
> "initializing\_shards": 0,  
> "unassigned\_shards": 0,  
> "delayed\_unassigned\_shards": 0,  
> "number\_of\_pending\_tasks": 0,  
> "number\_of\_in\_flight\_fetch": 0,  
> "task\_max\_waiting\_in\_queue\_millis": 0,  
> "active\_shards\_percent\_as\_number": 100  
> }

The cluster is very stable since then.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 6, 2018, 9:34am UTC](https://discuss.elastic.co/t/too-many-unassigned-replica-shards-everyday/112550/12 "2018-02-06T09:34:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
