# Elasticsearch cluster is in Red state. How to recover it?

**URL:** <https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444>\
**Category:** Elasticsearch\
**Created:** [November 12, 2019, 5:53am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444 "2019-11-12T05:53:06Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![chaitra\_hegde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chaitra_hegde/32/54151_2.png) [@chaitra\_hegde](https://discuss.elastic.co/u/chaitra_hegde)\
**Post date:** [November 12, 2019, 5:53am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/1 "2019-11-12T05:53:06Z")

</div>

Hi,  
I am using Elasticsearch 6.6.1 in k8s environment. My cluster was in green state before. But now my cluster is in red state due to UNASSIGNED shards. I see many shards are in PRIMARY\_FAILED state.

You can find the details below.  
Response from primary shard:  
`curl -X GET "http://xx.xx.xx.xx:9200/_cluster/allocation/explain?pretty" -H 'Content-Type: application/json' -d'`  
` {`  
`"index": "log-2019-10-08",`  
` "shard": 9,`  
`"primary": true`  
` }`  
` '`  
` {`  
` "index" : "log-2019-10-08",`  
`"shard" : 9,`  
` "primary" : true,`  
`"current_state" : "initializing",`  
` "unassigned_info" : {`  
` "reason" : "NODE_LEFT",`  
` "at" : "2019-10-22T09:26:26.488Z",`  
` "details" : "node_left[UVNJBTB8SuC3OiVnaB4Tfw]",`  
` "last_allocation_status" : "awaiting_info"`  
` },`  
` "current_node" : {`  
` "id" : "UVNJBTB8SuC3OiVnaB4Tfw",`  
` "name" : "elasticsearch-data-4",`  
` "transport_address" : "xx.xx.xx.xx:9300"`  
` },`  
`"explanation" : "the shard is in the process of initializing on node [elasticsearch-data-4], wait until initialization has completed"`  
` }`

The replica response is below.  
`curl -X GET "http://xx.xx.xx.xx:9200/_cluster/allocation/explain?pretty" -H 'Content-Type: application/json' -d'`  
` {`  
`"index": "log-2019-10-08",`  
` "shard": 9,`  
`"primary": false`  
` }`  
` '`  
`{`  
`"index" : "log-2019-10-08",`  
`"shard" : 9,`  
` "primary" : false,`  
`"current_state" : "unassigned",`  
` "unassigned_info" : {`  
` "reason" : "PRIMARY_FAILED",`  
`"at" : "2019-10-21T20:06:43.919Z",`  
`"details" : "primary failed while replica initializing",`  
` "last_allocation_status" : "no_attempt"`  
` },`  
` "can_allocate" : "no",`  
`"allocate_explanation" : "cannot allocate because allocation is not permitted to any of the nodes",`  
` "node_allocation_decisions" : [`  
`{`  
` "node_id" : "Ljab6IuXQTOUNxk8RkcuGg",`  
`"node_name" : "elasticsearch-data-1",`  
`"transport_address" : "xx.xx.xx.xx:9300",`  
`"node_decision" : "no",`  
` "deciders" : [`  
`{`  
`"decider" : "replica_after_primary_active",`  
` "decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
` },`  
`{`  
` "decider" : "throttling",`  
`"decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
`}`  
` ]`  
` },`  
` {`  
` "node_id" : "UVNJBTB8SuC3OiVnaB4Tfw",`  
`"node_name" : "elasticsearch-data-4",`  
` "transport_address" : "xx.xx.xx.xx:9300",`  
`"node_decision" : "no",`  
` "deciders" : [`  
` {`  
` "decider" : "replica_after_primary_active",`  
`"decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
` },`  
` {`  
`"decider" : "same_shard",`  
` "decision" : "NO",`  
` "explanation" : "the shard cannot be allocated to the same node on which a copy of the shard already exists [[log-2019-10-08][9],` `node[UVNJBTB8SuC3OiVnaB4Tfw], [P], recovery_source[existing store recovery; bootstrap_history_uuid=false], s[INITIALIZING], a[id=clqmzyGgSQC4HBKolutV-Q], unassigned_info[[reason=NODE_LEFT], at[2019-10-22T09:26:26.488Z], delayed=false, details[node_left[UVNJBTB8SuC3OiVnaB4Tfw]], allocation_status[fetching_shard_data]]]"`  
`},`  
`{`  
` "decider" : "throttling",`  
` "decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
` }`  
` ]`  
` },`  
` {`  
`"node_id" : "V4qtWbtLRqyDqW9f6T0mog",`  
`"node_name" : "elasticsearch-data-2",`  
` "transport_address" : "xx.xx.xx.xx:9300",`  
` "node_decision" : "no",`  
` "deciders" : [`  
` {`  
` "decider" : "replica_after_primary_active",`  
`"decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
` },`  
`{`  
`"decider" : "throttling",`  
` "decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
` }`  
` ]`  
` },`  
` {`  
` "node_id" : "r6UCUEPzR6aY0Kz8NiauDg",`  
`"node_name" : "elasticsearch-data-0",`  
`"transport_address" : "xx.xx.xx.xx:9300",`  
` "node_decision" : "no",`  
` "deciders" : [`  
`{`  
`"decider" : "replica_after_primary_active",`  
` "decision" : "NO",`  
` "explanation" : "primary shard for this replica is not yet active"`  
` },`  
` {`  
` "decider" : "throttling",`  
` "decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
`}`  
` ]`  
` },`  
`{`  
`"node_id" : "sNmtl-VvQMqS2bcXEycB-g",`  
` "node_name" : "elasticsearch-data-3",`  
` "transport_address" : "xx.xx.xx.xx:9300",`  
` "node_decision" : "no",`  
` "deciders" : [`  
` {`  
`"decider" : "replica_after_primary_active",`  
` "decision" : "NO",`  
`"explanation" : "primary shard for this replica is not yet active"`  
` },`  
`{`  
` "decider" : "throttling",`  
` "decision" : "NO",`  
` "explanation" : "primary shard for this replica is not yet active"`  
`}`  
` ]`  
`}`  
` ]`  
`}`  
How can I bring back my cluster to healthy state?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 12, 2019, 6:26am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/2 "2019-11-12T06:26:08Z")

</div>

What is the full output of the [cluster health API](https://www.elastic.co/guide/en/elasticsearch/reference/7.4/cluster-health.html)?

---

<div class="post-metadata">

**Author:** ![chaitra\_hegde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chaitra_hegde/32/54151_2.png) [@chaitra\_hegde](https://discuss.elastic.co/u/chaitra_hegde)\
**Post date:** [November 12, 2019, 7:16am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/3 "2019-11-12T07:16:16Z")

</div>

Response from cluster health API is below:  
`curl -XGET xx.xx.xx.xx:9200/_cluster/health?pretty`  
`{`  
` "cluster_name" : "my-cluster-1",`  
` "status" : "red",`  
` "timed_out" : false,`  
`"number_of_nodes" : 11,`  
`"number_of_data_nodes" : 6,`  
` "active_primary_shards" : 494,`  
` "active_shards" : 716,`  
` "relocating_shards" : 0,`  
`"initializing_shards" : 449,`  
` "unassigned_shards" : 401,`  
` "delayed_unassigned_shards" : 0,`  
` "number_of_pending_tasks" : 20,`  
` "number_of_in_flight_fetch" : 0,`  
`"task_max_waiting_in_queue_millis" : 6691643,`  
` "active_shards_percent_as_number" : 45.721583652618136`  
`}`

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 12, 2019, 8:01am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/4 "2019-11-12T08:01:46Z")

</div>

How did you end up in this state? Are all original data nodes now part of the cluster?

---

<div class="post-metadata">

**Author:** ![chaitra\_hegde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chaitra_hegde/32/54151_2.png) [@chaitra\_hegde](https://discuss.elastic.co/u/chaitra_hegde)\
**Post date:** [November 12, 2019, 9:07am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/5 "2019-11-12T09:07:16Z")

</div>

Hi,  
Two of the worker node had gone down in few hours time difference. When we recovered the worker nodes the issue started appearing in elasticsearch cluster.  
All the original data nodes are now the part of cluster.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 12, 2019, 11:14am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/6 "2019-11-12T11:14:35Z")

</div>

The primary shard you looked at above is recovering:

> [@chaitra\_hegde](#):
>
> `"primary" : true,`  
> `"current_state" : "initializing"`

This means you should just wait and eventually it will recover.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 12, 2019, 12:23pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/8 "2019-11-12T12:23:09Z")

</div>

> [@chaitra\_hegde](#):
>
> `"initializing_shards" : 449,`

This suggests you have increased a setting such as `cluster.routing.allocation.node_concurrent_recoveries` far too high. Your cluster may be [deadlocked](https://github.com/elastic/elasticsearch/issues/36195). Could you set it (and any other related settings) back to the default and perform a full cluster restart?

---

<div class="post-metadata">

**Author:** ![chaitra\_hegde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chaitra_hegde/32/54151_2.png) [@chaitra\_hegde](https://discuss.elastic.co/u/chaitra_hegde)\
**Post date:** [November 18, 2019, 6:45am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/9 "2019-11-18T06:45:08Z")

</div>

But the primary shards are in intializing state since 4-5days.

---

<div class="post-metadata">

**Author:** ![chaitra\_hegde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chaitra_hegde/32/54151_2.png) [@chaitra\_hegde](https://discuss.elastic.co/u/chaitra_hegde)\
**Post date:** [November 18, 2019, 6:46am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/10 "2019-11-18T06:46:08Z")

</div>

Hi,  
I am using k8s environment. What do you mean by full cluster restart? How can I restart my ES cluster?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 18, 2019, 8:32am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/11 "2019-11-18T08:32:27Z")

</div>

I don't know about Kubernetes specifically, but a full cluster restart is where you shut all of the nodes down and then start them all up again.

---

<div class="post-metadata">

**Author:** ![chaitra\_hegde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chaitra_hegde/32/54151_2.png) [@chaitra\_hegde](https://discuss.elastic.co/u/chaitra_hegde)\
**Post date:** [November 18, 2019, 9:38am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/12 "2019-11-18T09:38:36Z")

</div>

Hi,  
On what all conditions/scenarios can `cluster.routing.allocation.node_concurrent_recoveries` be tuned to other values than default values?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 18, 2019, 10:26am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/13 "2019-11-18T10:26:59Z")

</div>

I would only use this parameter for experiments in a test environment. I would not recommend adjusting it from the default in a production environment.

---

<div class="post-metadata">

**Author:** ![alex\_polisevschi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alex_polisevschi/32/52431_2.png) [@alex\_polisevschi](https://discuss.elastic.co/u/alex_polisevschi)\
**Post date:** [November 19, 2019, 6:13pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/14 "2019-11-19T18:13:24Z")

</div>

Have you tried the reroute API?

POST /\_cluster/reroute?retry\_failed=true

> The cluster will attempt to allocate a shard a maximum of `index.allocation.max_retries` times in a row (defaults to `5` ), before giving up and leaving the shard unallocated. This scenario can be caused by structural problems such as having an analyzer which refers to a stopwords file which doesn’t exist on all nodes.
> 
> Once the problem has been corrected, allocation can be manually retried by calling the [`reroute`](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/cluster-reroute.html) API with the `?retry_failed` URI query parameter, which will attempt a single retry round for these shards.

> **[Cluster Reroute | Elasticsearch Guide \[6.6\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/cluster-reroute.html#_retrying_failed_allocations)**

---

<div class="post-metadata">

**Author:** ![chaitra\_hegde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chaitra_hegde/32/54151_2.png) [@chaitra\_hegde](https://discuss.elastic.co/u/chaitra_hegde)\
**Post date:** [November 27, 2019, 6:58am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/15 "2019-11-27T06:58:36Z")

</div>

Hi,  
I have set back the `cluster.routing.allocation.node_concurrent_recoveries` to default value and performed a full cluster restart. And also i have performed reroute API with `?retry_failed`.  
Now I am able to reduce the number of shards which are in red state. So after performing this i have 66 unassigned shards.  
Now I have many indices with some shards are in red state. Since i have recovered some of the shards of that index, I do not want to delete the full index in red state to bring back my cluster to healthy state.  
So how can I delete particular red shard of an index without deleting the whole index?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 25, 2019, 6:58am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-in-red-state-how-to-recover-it/207444/16 "2019-12-25T06:58:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
