# Theory question: How to recover data from 3 node cluster after 2 nodes failed

**URL:** <https://discuss.elastic.co/t/theory-question-how-to-recover-data-from-3-node-cluster-after-2-nodes-failed/310262>\
**Category:** Elasticsearch\
**Created:** [July 21, 2022, 9:13am UTC](https://discuss.elastic.co/t/theory-question-how-to-recover-data-from-3-node-cluster-after-2-nodes-failed/310262 "2022-07-21T09:13:54Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![defalt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/defalt/32/71379_2.png) [@defalt](https://discuss.elastic.co/u/defalt)\
**Post date:** [July 21, 2022, 9:13am UTC](https://discuss.elastic.co/t/theory-question-how-to-recover-data-from-3-node-cluster-after-2-nodes-failed/310262/1 "2022-07-21T09:13:54Z")

</div>

Long time no see, but hello again.

We had a **3 Node cluster** running very well till it stopped running very well. The SSD's of **2 Nodes** failed in a timespan of 4 hours. This was in the middle of the night so no actions could be taken. We now had 2 corrupted nodes and one node still up. We though everything was fine because we had 3 nodes and **every node** had a **replica of every index**.

The problem was that after a restart of this node it could not elect itself as a master and could not start back up again. No matter what we tried we were not able to get it running. In the end we resorted to using [unsafe cluster bootstraping](https://www.elastic.co/guide/en/elasticsearch/reference/8.2/node-tool.html#node-tool-unsafe-bootstrap) to recover at least some of the data.

But this cant be the "best" way. So what is the actual best practice when 2 of 3 nodes fail? What is the best way to get the data back, to get the cluster back up again? Or is everyone just hoping that this never happens?

Thanks for your ideas 🙂

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 21, 2022, 10:33pm UTC](https://discuss.elastic.co/t/theory-question-how-to-recover-data-from-3-node-cluster-after-2-nodes-failed/310262/2 "2022-07-21T22:33:53Z")

</div>

What version are you on?

---

<div class="post-metadata">

**Author:** ![defalt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/defalt/32/71379_2.png) [@defalt](https://discuss.elastic.co/u/defalt)\
**Post date:** [July 22, 2022, 6:28am UTC](https://discuss.elastic.co/t/theory-question-how-to-recover-data-from-3-node-cluster-after-2-nodes-failed/310262/3 "2022-07-22T06:28:41Z")

</div>

7.17.3

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 22, 2022, 8:07am UTC](https://discuss.elastic.co/t/theory-question-how-to-recover-data-from-3-node-cluster-after-2-nodes-failed/310262/4 "2022-07-22T08:07:54Z")

</div>

> [@defalt](#):
>
> The SSD's of **2 Nodes** failed in a timespan of 4 hours.

Was there any correlation between the failures? E.g. were they running in physical proximity so they might have suffered similar damage due to heat or vibration or some other environmental effect? Were they the same model of drive? The same manufacturing batch perhaps?

> [@defalt](#):
>
> Or is everyone just hoping that this never happens?

There's no watertight protection against multiple failures so it really depends how paranoid you want to be. Physically separating the nodes helps. Decorrelating their hardware helps. RAID helps, especially on the master nodes - dedicated masters don't need much storage but they do need it to be reliable. But ultimately it's all about probabilities and you can still be extraordinarily unlucky. In that case, [the manual recommends restoring from a recent snapshot](https://www.elastic.co/guide/en/elasticsearch/reference/current/discovery-troubleshooting.html#discovery-no-master):

> If you can’t start enough nodes to form a quorum, start a new cluster and restore data from a recent snapshot.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 19, 2022, 8:08am UTC](https://discuss.elastic.co/t/theory-question-how-to-recover-data-from-3-node-cluster-after-2-nodes-failed/310262/5 "2022-08-19T08:08:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
