# Elasticserach cluster in red state and a lot of unassigned shards

**URL:** <https://discuss.elastic.co/t/elasticserach-cluster-in-red-state-and-a-lot-of-unassigned-shards/248579>\
**Category:** Elasticsearch\
**Created:** [September 14, 2020, 8:02pm UTC](https://discuss.elastic.co/t/elasticserach-cluster-in-red-state-and-a-lot-of-unassigned-shards/248579 "2020-09-14T20:02:19Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sahishnu\_Maharjan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sahishnu_maharjan/32/44972_2.png) [@Sahishnu\_Maharjan](https://discuss.elastic.co/u/Sahishnu_Maharjan)\
**Post date:** [September 14, 2020, 8:02pm UTC](https://discuss.elastic.co/t/elasticserach-cluster-in-red-state-and-a-lot-of-unassigned-shards/248579/1 "2020-09-14T20:02:19Z")

</div>

Greetings, I was given a broken graylog cluster to fix. Upon examination, elasticsearch cluster (7 node: 1 master 6 data) was red as its fs was 100%. After clearing older data, I started noticing unassigned shards error. I kinda figured out what that meant by looking it in the web and I tried a lot of fix but didn't work. Here are some of the things I tried.

1. After fs space was made available(1tb each total, used 50% now), I restarted all the data nodes - Unassigned shards no. increased initially but then subsided to original figure 1800
2. I ran curl -s -X POST 'http://:9200$es-master/\_cluster/reroute?retry\_failed=true&pretty' - I get a lot of state started output with allocation id. it does assign a few like 10-15 but thats it.
3. I also made sure allocation is enabled on all data nodes.  
C02RT2M5FVH6:~ mahars01$ curl -s "dfwlnpgles-10:9200/\_cluster/settings?pretty"  
{  
"persistent" : { },  
"transient" : {  
"cluster" : {  
"routing" : {  
"allocation" : {  
"enable" : "all"  
}  
}  
}  
}  
}

But none of these seem to help.

Here are some info:

C02RT2M5FVH6:~ mahars01$ curl dfwlnpgles-10:9200/\_cat/health?v  
epoch timestamp cluster status node.total node.data shards pri relo init unassign pending\_tasks max\_task\_wait\_time active\_shards\_percent  
1600112243 14:37:23 elasticsearch2 red 7 7 1311 1303 0 2 1749 0 - 42.8%  
C02RT2M5FVH6:~ mahars01$ curl dfwlnpgles-10:9200/\_cat/nodes  
10.5.4.78 17 99 53 5.17 6.27 6.16 mdi - dfwlnpgles-16  
10.5.4.73 18 98 3 0.28 0.20 0.15 mdi - dfwlnpgles-11  
10.5.4.75 25 99 58 6.54 7.17 7.10 mdi - dfwlnpgles-13  
10.5.4.74 57 99 88 7.69 6.90 5.92 mdi - dfwlnpgles-12  
10.5.4.77 30 99 33 1.82 1.74 1.77 mdi - dfwlnpgles-15  
10.5.4.76 15 98 28 1.95 1.87 1.83 mdi - dfwlnpgles-14  
10.5.4.72 44 41 7 0.17 0.10 0.13 mdi \* dfwlnpgles-10

C02RT2M5FVH6:~ mahars01$ curl -s dfwlnpgles-10:9200/\_cat/indices | wc -l  
632

I check one of the unassigned shards

C02RT2M5FVH6:~ mahars01$ curl '[http://dfwlnpgles-10:9200/\_cluster/allocation/explain?pretty](http://dfwlnpgles-10:9200/_cluster/allocation/explain?pretty)' -d '{  
"index": "weekly\_1372",  
"shard": 3,  
"primary": true  
}'  
{  
"index" : "weekly\_1372",  
"shard" : 3,  
"primary" : true,  
"current\_state" : "unassigned",  
"unassigned\_info" : {  
"reason" : "ALLOCATION\_FAILED",  
"at" : "2020-09-03T18:22:53.137Z", `#NOTE: this date is from the time when fs was 100%`  
"failed\_allocation\_attempts" : 5,  
"details" : "failed to create shard, failure IOException[No space left on device]",  
"last\_allocation\_status" : "no"  
},  
"can\_allocate" : "no",  
"allocate\_explanation" : "cannot allocate because allocation is not permitted to any of the nodes that hold an in-sync shard copy",  
"node\_allocation\_decisions" : [  
{  
"node\_id" : "12SVd1opTqukG9D5q8CAwg",  
"node\_name" : "dfwlnpgles-10",  
"transport\_address" : "10.5.4.72:9300",  
"node\_decision" : "no",  
"store" : {  
"found" : false  
}  
},  
{  
"node\_id" : "76V1negYQOK-cZOMLJAifw",  
"node\_name" : "dfwlnpgles-13",  
"transport\_address" : "10.5.4.75:9300",  
"node\_decision" : "no",  
"store" : {  
"found" : false  
}  
},  
{  
"node\_id" : "7Lh9tonFTWeZ6aPF4tlF3g",  
"node\_name" : "dfwlnpgles-16",  
"transport\_address" : "10.5.4.78:9300",  
"node\_decision" : "no",  
"store" : {  
"in\_sync" : true,  
"allocation\_id" : "MFcr1BX9T8yN2ihCiGsyPQ"  
},  
"deciders" : [  
{  
"decider" : "max\_retry",  
"decision" : "NO",  
"explanation" : "shard has exceeded the maximum number of retries [5] on failed allocation attempts - manually call [/\_cluster/reroute?retry\_failed=true] to retry, [unassigned\_info[[reason=ALLOCATION\_FAILED], at[2020-09-03T18:22:53.137Z], failed\_attempts[5], delayed=false, details[failed to create shard, failure IOException[No space left on device]], allocation\_status[deciders\_no]]]"  
}  
]  
},  
{  
"node\_id" : "Dcks1X5LSxy\_RtrSw\_GoTw",  
"node\_name" : "dfwlnpgles-12",  
"transport\_address" : "10.5.4.74:9300",  
"node\_decision" : "no",  
"store" : {  
"found" : false  
}  
},  
{  
"node\_id" : "Qguk\_\_nnSXmT8e9bny19VQ",  
"node\_name" : "dfwlnpgles-11",  
"transport\_address" : "10.5.4.73:9300",  
"node\_decision" : "no",  
"store" : {  
"found" : false  
}  
},  
{  
"node\_id" : "XuGZf3x3RTqb71eQEbWWdQ",  
"node\_name" : "dfwlnpgles-15",  
"transport\_address" : "10.5.4.77:9300",  
"node\_decision" : "no",  
"store" : {  
"in\_sync" : false,  
"allocation\_id" : "jAKEjQS-QEy\_jCuQS13s0g",  
"store\_exception" : {  
"type" : "file\_not\_found\_exception",  
"reason" : "no segments\* file found in SimpleFSDirectory@/app/data/nodes/0/indices/blsAn4FlQmyzPtkp6y1xcw/3/index lockFactory=org.apache.lucene.store.NativeFSLockFactory@7824675a: files: [write.lock]"  
}  
}  
},  
{  
"node\_id" : "gfDZofDKQwmhswnGAiFKag",  
"node\_name" : "dfwlnpgles-14",  
"transport\_address" : "10.5.4.76:9300",  
"node\_decision" : "no",  
"store" : {  
"found" : false  
}  
}  
]  
}

Anything more I can provide here?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 12, 2020, 8:02pm UTC](https://discuss.elastic.co/t/elasticserach-cluster-in-red-state-and-a-lot-of-unassigned-shards/248579/2 "2020-10-12T20:02:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
