# Shard failure after restart of node - ES 1.7.5

**URL:** <https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057>\
**Category:** Elasticsearch\
**Created:** [August 26, 2016, 4:20pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057 "2016-08-26T16:20:02Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![kmrdiscuss](https://avatars.discourse-cdn.com/v4/letter/k/db5fbb/32.png) [@kmrdiscuss](https://discuss.elastic.co/u/kmrdiscuss)\
**Post date:** [August 26, 2016, 4:20pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/1 "2016-08-26T16:20:02Z")

</div>

Hi. We have a three node cluster, running ES 1.7.5, with the following config.  
index.number\_of\_shards: 5  
index.number\_of\_replicas: 2  
node.master: true  
node.data: true

After writing our data we see each node has five shards [0..4] as expected.  
We perform a query and obtain the correct number of documents, etc...  
We then stop nodes 1 and 2 and successfully re-query the cluster.  
All well and good.

We then stop node 3 and then start node 3. We receive the error.  
"SearchPhaseExecutionException[Failed to execute phase [query], all shards failed]"  
We receive the same error no matter which node is started.

We interpret our configuration to mean we should be able to successfully query our cluster with only one node running.  
Any ideas?

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [August 26, 2016, 5:05pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/2 "2016-08-26T17:05:58Z")

</div>

Which version of Elasticsearch are you running?

---

<div class="post-metadata">

**Author:** ![kmrdiscuss](https://avatars.discourse-cdn.com/v4/letter/k/db5fbb/32.png) [@kmrdiscuss](https://discuss.elastic.co/u/kmrdiscuss)\
**Post date:** [August 26, 2016, 5:22pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/3 "2016-08-26T17:22:31Z")

</div>

1.7.5

---

<div class="post-metadata">

**Author:** ![kmrdiscuss](https://avatars.discourse-cdn.com/v4/letter/k/db5fbb/32.png) [@kmrdiscuss](https://discuss.elastic.co/u/kmrdiscuss)\
**Post date:** [August 26, 2016, 6:24pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/4 "2016-08-26T18:24:18Z")

</div>

I see the same behavior using 2.3.3

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [August 26, 2016, 7:02pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/5 "2016-08-26T19:02:22Z")

</div>

This is an issue where ES wants to have a "quorum" of shard copies available before it recovers on a cluster restart. You can set `index.recovery.initial_shards` to `1` so that it only waits for one shard copy to be available before the primary recovers.

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [August 26, 2016, 7:06pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/6 "2016-08-26T19:06:31Z")

</div>

As a side note, you could have also set the number of replicas to something less (like 0 or 1) and it would've also recovered your primary. ES is essentially waiting for enough nodes for `index.recovery.initial_shards` to be able to be met. If the default is quorum, which means for 3 shard copies, you would need 2 nodes to hold a quorum of those copies, then ES won't recover until those 2 nodes are up. If you set the number of replicas to 1 and keep the initial\_shards setting to quorum, then you would have met the quorum by just starting one node.

---

<div class="post-metadata">

**Author:** ![kmrdiscuss](https://avatars.discourse-cdn.com/v4/letter/k/db5fbb/32.png) [@kmrdiscuss](https://discuss.elastic.co/u/kmrdiscuss)\
**Post date:** [August 26, 2016, 7:41pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/7 "2016-08-26T19:41:22Z")

</div>

@abeyad - Thanks for your insight it solved our issue. Setting index.number\_of\_replicas: 1 did not work it needed to be set to 0. Setting index.recovery.initial\_shards to 1 worked.

Thanks again.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:24pm UTC](https://discuss.elastic.co/t/shard-failure-after-restart-of-node-es-1-7-5/59057/8 "2017-07-05T22:24:45Z")

</div>


