# If Primary and Replica shards both fail how to recover?

**URL:** <https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636>\
**Category:** Elasticsearch\
**Created:** [February 15, 2019, 7:03pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636 "2019-02-15T19:03:36Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![jwlee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jwlee/32/61090_2.png) [@jwlee](https://discuss.elastic.co/u/jwlee)\
**Post date:** [February 15, 2019, 7:03pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636/1 "2019-02-15T19:03:36Z")

</div>

I tried to find a relevant document but I could not get the answer.

Let's say there are primary shard and replica shard. I believe if the primary shard goes down, the replica shard promoted to the primary shard and recreate the replica. On the other hand, if replica shard goes down, it will simply recreate the replica based off from the primary.

The question that I have is:

1. What happens when primary and replica shards go down? Do it impact other shards as well?
2. Is there a way that can fully recover when both primary and replica shards go down? If there is a recovery action, what happens when there are new inserts or updates to the documents that belong to these shards during the recovery? I believe we can schedule a daily snapshot of the database but this means it will lose any new data that are coming to these shards after the snapshot was taken.

In practical even though the chances of both primary and replica shard failure are rare I believe we should still take more than one replica and distribute to different data center if possible. However, in case of the worst scenario, I am really curious how we can recover from all the primary and replica shards failure.

Any inputs are really appreciated. Thanks.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 15, 2019, 7:18pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636/2 "2019-02-15T19:18:28Z")

</div>

It depends on what you mean by "go down". If a primary shard fails then you are correct that the master promotes one of the active in-sync replicas to become the new primary. If there are currently no active in-sync replicas then it waits until one appears, and your cluster health reports as `RED`. So if "go down" means "... and come back up again" then everything's fine. However if all of the in-sync copies of your data _permanently_ disappear then there's not a lot that can be done: you have by definition lost data.

> [@jwlee](#):
>
> I believe we should still take more than one replica and distribute to different data center if possible.

It's not recommended to split a single Elasticsearch cluster across data centres. It expects the node-to-node connections to have low latency.

---

<div class="post-metadata">

**Author:** ![jwlee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jwlee/32/61090_2.png) [@jwlee](https://discuss.elastic.co/u/jwlee)\
**Post date:** [February 15, 2019, 8:56pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636/3 "2019-02-15T20:56:24Z")

</div>

Thanks for the inputs.

> [@DavidTurner](#):
>
> If there are currently no active in-sync replicas then it waits until one appears, and your cluster health reports as `RED` . So if "go down" means "... and come back up again" then everything's fine.

Not sure if everything is fine. Wouldn't new requests be lost when it waits for active in-sync replica to appear?

What happens when there are new insert or update requests to the shards during the cluster health is RED?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 15, 2019, 10:25pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636/4 "2019-02-15T22:25:49Z")

</div>

> [@jwlee](#):
>
> Wouldn't new requests be lost when it waits for active in-sync replica to appear?

Not lost, no, such requests would be rejected.

---

<div class="post-metadata">

**Author:** ![jwlee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jwlee/32/61090_2.png) [@jwlee](https://discuss.elastic.co/u/jwlee)\
**Post date:** [February 20, 2019, 7:24pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636/5 "2019-02-20T19:24:11Z")

</div>

So if primary shard fails and there is no active in-sync replicas then the data has been lost. Correct?

If so, is there a way that we can mitigate this risk proactively? Such as taking daily db snapshot as a backup?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 20, 2019, 9:45pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636/6 "2019-02-20T21:45:01Z")

</div>

> [@jwlee](#):
>
> is there a way that we can mitigate this risk proactively?

The best thing to do is to make sure that you _do_ have active in-sync replicas, by keeping your cluster health at `GREEN`. That way your data is safe unless you suffer from multiple failures involving independent devices all at the same time.

> [@jwlee](#):
>
> Such as taking daily db snapshot as a backup?

Sure, you can do this too. Indeed many people take snapshots much more frequently than daily which works because they are incremental - a snapshot every 30 minutes is not unusual.

Restoring from a snapshot still means you lose any data written after the snapshot was taken, but maybe this is ok in your case?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 20, 2019, 9:50pm UTC](https://discuss.elastic.co/t/if-primary-and-replica-shards-both-fail-how-to-recover/168636/7 "2019-03-20T21:50:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
