# Unable to recover my cluster

**URL:** <https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000>\
**Category:** Elasticsearch\
**Created:** [May 9, 2023, 8:06pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000 "2023-05-09T20:06:52Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 9, 2023, 8:06pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/1 "2023-05-09T20:06:52Z")

</div>

We moved our data to new version of elasticsearch. The new cluster have 3 master, 3 hot and 3 warm nodes. Everything was working fine till this morning and all of a cluster health went red. After looking at it further, I realized due to low storage on warm nodes, shard allocation failed. I increased the volumes and restarted the warm nodes. With \_cluster reroute api, I tried recovering the cluster. But nothing happened and now  
with \_cat/shards?v=true&h=index,shard,prirep,state,node,unassigned.reason&s=state  
I am getting almost 83 shards unassigned with reason NODE\_LEFT.

Please help.

Thanks in Advance

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 10, 2023, 3:11am UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/2 "2023-05-10T03:11:15Z")

</div>

Welcome to our community! 😃

What version are you on? What do your master logs show? What does an [`?explain`](https://www.elastic.co/guide/en/elasticsearch/reference/8.7/cluster-reroute.html) show against one of the shards? What does [`_cat/recovery?v`](https://www.elastic.co/guide/en/elasticsearch/reference/8.7/cat-recovery.html) show?

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 1:17pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/3 "2023-05-10T13:17:11Z")

</div>

Hi Mark,  
Thanks for your response.  
I don't see anything specific to the shard issue on master log  
for \_cluster/allocation/explain

```auto
{
note: "No shard was specified in the explain API request, so this response explains a randomly chosen unassigned shard. There may be other unassigned shards in this cluster which cannot be assigned for different reasons. It may not be possible to assign this shard until one of the other shards is assigned correctly. To explain the allocation of other shards (whether assigned or unassigned) you must specify the target shard in the request to this API.",
index: "hashgraphaccounttransfer-000017",
shard: 3,
primary: true,
current_state: "unassigned",
unassigned_info: {
reason: "NODE_LEFT",
at: "2023-05-09T12:56:28.001Z",
details: "node_left [_u38-owTRr2UwR0kTUH2rw]",
last_allocation_status: "throttled"
},
can_allocate: "throttled",
allocate_explanation: "allocation temporarily throttled",
node_allocation_decisions: [
{
node_id: "_u38-owTRr2UwR0kTUH2rw",
node_name: "warm-node-2",
transport_address: "172.32.1.125:9300",
node_attributes: {
data: "warm",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "throttled",
store: {
in_sync: true,
allocation_id: "r_QbqvGXSkOaLcKouFXdKg"
},
deciders: [
{
decider: "throttling",
decision: "THROTTLE",
explanation: "reached the limit of ongoing initial primary recoveries [6], cluster setting [cluster.routing.allocation.node_initial_primaries_recoveries=6]"
}
]
},
{
node_id: "4nDKSda9Q3SsGH4NqYHgSA",
node_name: "hot-node-3",
transport_address: "172.32.2.17:9300",
node_attributes: {
data: "hot",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
},
{
node_id: "7JUClIzEQaaqxEd6HpBfUg",
node_name: "warm-node-3a",
transport_address: "172.32.2.68:9300",
node_attributes: {
data: "warm",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
},
{
node_id: "FsiWEnj-QBO0Tp_5DsFS5Q",
node_name: "warm-node-1",
transport_address: "172.32.0.208:9300",
node_attributes: {
data: "warm",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
},
{
node_id: "IatGaiCUQ3SEKg9BSXjYtQ",
node_name: "warm-node-2a",
transport_address: "172.32.1.79:9300",
node_attributes: {
data: "warm",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
},
{
node_id: "SjsTQhboQEC0USoMCP5khQ",
node_name: "warm-node-3",
transport_address: "172.32.2.227:9300",
node_attributes: {
data: "warm",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
},
{
node_id: "TTeJyCyrSIC281GTQUEh7g",
node_name: "hot-node-1",
transport_address: "172.32.0.246:9300",
node_attributes: {
data: "hot",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
},
{
node_id: "aVPNgVNwTCiNGT7eQdNevA",
node_name: "hot-node-2",
transport_address: "172.32.1.248:9300",
node_attributes: {
data: "hot",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
},
{
node_id: "e-LTEHGyR3OUcqEsqC2Yew",
node_name: "warm-node-1a",
transport_address: "172.32.0.176:9300",
node_attributes: {
data: "warm",
xpack.installed: "true",
transform.node: "false"
},
node_decision: "no",
store: {
found: false
}
}
]
}

```

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 1:18pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/4 "2023-05-10T13:18:38Z")

</div>

/\_cat/recovery?v

```auto
index shard time type stage source_host source_node target_host target_node repository snapshot files files_recovered files_percent files_total bytes bytes_recovered bytes_percent bytes_total translog_ops translog_ops_recovered translog_ops_percent
hashgraphtxnsummary-000007 0 81ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000007 0 90ms peer done 172.32.2.17 hot-node-3 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 16 16 100.0%
hashgraphtxnsummary-000007 1 70ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000007 1 93ms peer done 172.32.2.17 hot-node-3 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 12 12 100.0%
hashgraphtxnsummary-000007 2 89ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000007 2 93ms peer done 172.32.2.17 hot-node-3 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 13 13 100.0%
hashgraphtxnsummary-000007 3 89ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000007 3 98ms peer done 172.32.2.17 hot-node-3 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 15 15 100.0%
hashgraphtxnsummary-000007 4 88ms peer done 172.32.2.17 hot-node-3 172.32.0.246 hot-node-1 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 16 16 100.0%
hashgraphtxnsummary-000007 4 79ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphhcstxnsummary-000002 0 44ms empty_store done n/a n/a 172.32.0.246 hot-node-1 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphhcstxnsummary-000002 0 49ms peer done 172.32.0.246 hot-node-1 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 0 0 100.0%
hashgraphhcstxnsummary-000002 1 77ms peer done 172.32.2.17 hot-node-3 172.32.0.246 hot-node-1 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 0 0 100.0%
hashgraphhcstxnsummary-000002 1 75ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphhcstxnsummary-000002 2 72ms peer done 172.32.1.248 hot-node-2 172.32.2.17 hot-node-3 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 0 0 100.0%
hashgraphhcstxnsummary-000002 2 34ms empty_store done n/a n/a 172.32.1.248 hot-node-2 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphhcstxnsummary-000002 3 53ms empty_store done n/a n/a 172.32.0.246 hot-node-1 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphhcstxnsummary-000002 3 62ms peer done 172.32.0.246 hot-node-1 172.32.2.17 hot-node-3 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 0 0 100.0%
hashgraphhcstxnsummary-000002 4 61ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphhcstxnsummary-000002 4 51ms peer done 172.32.2.17 hot-node-3 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 0 0 100.0%
hashgraphtxnsummary-000006 0 50ms empty_store done n/a n/a 172.32.0.246 hot-node-1 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000006 0 434ms peer done 172.32.0.246 hot-node-1 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 2476 2476 100.0%
hashgraphtxnsummary-000006 1 62ms empty_store done n/a n/a 172.32.2.17 hot-node-3 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000006 1 1.5s peer done 172.32.2.17 hot-node-3 172.32.1.248 hot-node-2 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 6017 6017 100.0%
hashgraphtxnsummary-000006 2 29.3s peer done 172.32.1.248 hot-node-2 172.32.0.246 hot-node-1 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 121953 121953 100.0%
hashgraphtxnsummary-000006 2 39ms empty_store done n/a n/a 172.32.1.248 hot-node-2 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000006 3 47ms empty_store done n/a n/a 172.32.0.246 hot-node-1 n/a n/a 0 0 0.0% 0 0 0 0.0% 0 0 0 100.0%
hashgraphtxnsummary-000006 3 474ms peer done 172.32.0.246 hot-node-1 172.32.2.17 hot-node-3 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 2404 2404 100.0%
hashgraphtxnsummary-000006 4 1.6s peer done 172.32.2.17 hot-node-3 172.32.0.246 hot-node-1 n/a n/a 1 1 100.0% 1 226 226 100.0% 226 6075 6075 100.0%

```

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 1:21pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/5 "2023-05-10T13:21:08Z")

</div>

We are on version 7.17.9

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [May 10, 2023, 2:12pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/6 "2023-05-10T14:12:04Z")

</div>

post some output of

GET \_cluster/settings  
GET \_cluster/health

mainly check routing allocation and how many shard is still unassigned and relocating

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 2:47pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/7 "2023-05-10T14:47:50Z")

</div>

GET \_cluster/health

```auto
{
cluster_name: "mainnet-es-cluster",
status: "red",
timed_out: false,
number_of_nodes: 12,
number_of_data_nodes: 9,
active_primary_shards: 113,
active_shards: 206,
relocating_shards: 5,
initializing_shards: 19,
unassigned_shards: 63,
delayed_unassigned_shards: 0,
number_of_pending_tasks: 0,
number_of_in_flight_fetch: 0,
task_max_waiting_in_queue_millis: 0,
active_shards_percent_as_number: 71.52777777777779
}

```

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 2:48pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/8 "2023-05-10T14:48:47Z")

</div>

\_cluster/settings

```auto
persistent: {
cluster: {
routing: {
allocation: {
disk: {
watermark: {
low: "90%",
high: "93%"
}
}
}
}
}
},
transient: {
cluster: {
routing: {
allocation: {
node_initial_primaries_recoveries: "6"
}
}
}
}
}

```

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 2:51pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/9 "2023-05-10T14:51:01Z")

</div>

Total of below 2 is always 82

initializing\_shards: 19,  
unassigned\_shards: 63,

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [May 10, 2023, 4:07pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/10 "2023-05-10T16:07:08Z")

</div>

what about  
GET \_cat/shards

this should show you many shard are reinitializing. if that is the case you have to wait

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 7:50pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/11 "2023-05-10T19:50:26Z")

</div>

This is the # of shards reinitializing but it doesn't work either  
initializing\_shards: 19

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [May 10, 2023, 8:25pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/12 "2023-05-10T20:25:34Z")

</div>

> [@Ashu\_Mahajan](#):
>
> ```auto
> explanation: "reached the limit of ongoing initial primary recoveries [6], cluster setting [cluster.routing.allocation.node_initial_primaries_recoveries=6]"
> 
> ```

this one might causing problem as limit is 6.  
may be set to 20 and see if it makes any difference

```auto
PUT _cluster/settings
{
  "persistent": {
    "cluster.routing.allocation.node_initial_primaries_recoveries": 20
  }
}

```

---

<div class="post-metadata">

**Author:** ![Ashu\_Mahajan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashu_mahajan/32/117721_2.png) [@Ashu\_Mahajan](https://discuss.elastic.co/u/Ashu_Mahajan)\
**Post date:** [May 10, 2023, 10:58pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/13 "2023-05-10T22:58:56Z")

</div>

I actually tried changing that it didn't help

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 7, 2023, 10:59pm UTC](https://discuss.elastic.co/t/unable-to-recover-my-cluster/333000/14 "2023-06-07T22:59:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
