# Elasticsearch All Shards failed on cluster with multiple nodes on azure VM

**URL:** <https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809>\
**Category:** Elasticsearch\
**Created:** [July 19, 2020, 5:31pm UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809 "2020-07-19T17:31:48Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![pranav\_patwardhan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pranav_patwardhan/32/72409_2.png) [@pranav\_patwardhan](https://discuss.elastic.co/u/pranav_patwardhan)\
**Post date:** [July 19, 2020, 5:31pm UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/1 "2020-07-19T17:31:49Z")

</div>

I have Elasticsearch cluster with 3 master nodes, 2 data nodes, and cluster node deployed on different Virtual Machine on Azure. It was working fine but suddenly it failed and now `[search_phase_execution_exception] all shards failed` error coming when we are trying to search any data.  
There are a total of 200+ indexes with all as **red** status.  
Following is the health of the entire cluster

> "status": "red",  
> "timed\_out": false,  
> "number\_of\_nodes": 6,  
> "number\_of\_data\_nodes": 2,  
> "active\_primary\_shards": 5,  
> "active\_shards": 10,  
> "relocating\_shards": 0,  
> "initializing\_shards": 0,  
> "unassigned\_shards": 524,  
> "delayed\_unassigned\_shards": 0,  
> "number\_of\_pending\_tasks": 0,  
> "number\_of\_in\_flight\_fetch": 0,  
> "task\_max\_waiting\_in\_queue\_millis": 0,  
> "active\_shards\_percent\_as\_number": 1.8726591760299627

What could be the possible solution for this?  
Help would be appreciated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 19, 2020, 7:27pm UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/2 "2020-07-19T19:27:21Z")

</div>

Which version are you using? Is there anything in the Elasticsearch logs that provides any clues?

---

<div class="post-metadata">

**Author:** ![Steve\_Mushero](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steve_mushero/32/22441_2.png) [@Steve\_Mushero](https://discuss.elastic.co/u/Steve_Mushero)\
**Post date:** [July 20, 2020, 7:57am UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/3 "2020-07-20T07:57:22Z")

</div>

And what is cluster architecture as 6 nodes but only 2 data is pretty odd - did you lose most of your data nodes somehow?

Must be some history here - any nodes VM or process restart, as something had to happen. This sit on VMWare or some shared SAN storage or something that could have failed? You can do an explain to get first unassigned and why:

GET /\_cluster/allocation/explain

Which might help - any allocation awareness and maybe lost nodes or properties that prevents allocation?

---

<div class="post-metadata">

**Author:** ![pranav\_patwardhan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pranav_patwardhan/32/72409_2.png) [@pranav\_patwardhan](https://discuss.elastic.co/u/pranav_patwardhan)\
**Post date:** [July 21, 2020, 11:33am UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/4 "2020-07-21T11:33:46Z")

</div>

6.5.3

---

<div class="post-metadata">

**Author:** ![pranav\_patwardhan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pranav_patwardhan/32/72409_2.png) [@pranav\_patwardhan](https://discuss.elastic.co/u/pranav_patwardhan)\
**Post date:** [July 21, 2020, 2:02pm UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/5 "2020-07-21T14:02:49Z")

</div>

Thanks Steve, I am trying now with allocation/explain

---

<div class="post-metadata">

**Author:** ![pranav\_patwardhan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pranav_patwardhan/32/72409_2.png) [@pranav\_patwardhan](https://discuss.elastic.co/u/pranav_patwardhan)\
**Post date:** [July 21, 2020, 2:53pm UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/6 "2020-07-21T14:53:26Z")

</div>

Hello @Steve_Mushero,

````auto
        "primary": true,
        "current_state": "unassigned",
        "unassigned_info": {
            "reason": "CLUSTER_RECOVERED",
            "at": "2020-07-19T16:11:27.591Z",
            "last_allocation_status": "no_valid_shard_copy"
        },
        "can_allocate": "no_valid_shard_copy",
        "allocate_explanation": "cannot allocate because all found copies of the shard are either stale or corrupt",``` 

I am getting this response, is there any way to assign this shard again?
````

---

<div class="post-metadata">

**Author:** ![Steve\_Mushero](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steve_mushero/32/22441_2.png) [@Steve\_Mushero](https://discuss.elastic.co/u/Steve_Mushero)\
**Post date:** [July 22, 2020, 2:39am UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/7 "2020-07-22T02:39:49Z")

</div>

So what is history here and why only two data nodes? This error implies it had your indexes but lost them, likely because they are stale, i.e. there was a primary on another node that was lost before replicas were updated or something - you might check a few shards/indexes (there is an option to explain for this, see docs), but maybe the same.

I don't recall if you can promote a stale shard; I think there is an API for it but can fail, but better to find the bad nodes or understand what happened here.

And if you have snapshots, best to recover from them.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 19, 2020, 2:40am UTC](https://discuss.elastic.co/t/elasticsearch-all-shards-failed-on-cluster-with-multiple-nodes-on-azure-vm/241809/8 "2020-08-19T02:40:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
