# Unassigned primary + replica shard, minimise data loss

**URL:** https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393
**Category:** Elasticsearch
**Created:** [October 25, 2022, 10:03am UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393 "2022-10-25T10:03:20Z")
**Posts on this page:** 13
**Page:** 1

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [October 25, 2022, 10:03am UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/1 "2022-10-25T10:03:20Z")

</div>

Hello, my cluster is currently in red state due to one of both the primary and replica shard of an index becoming unassigned. This happened after a number of large tasks were executed simultaneously by accident on the same index alias. The index in question that was affected was the current write index of that index alias. The index contains valuable data and I would like to try to minimise any data loss if possible, ideally with none.

Calling GET \_cluster/allocation/explain on the the primary returns `unassigned_info`-\>`reason`:

"ALLOCATION\_FAILED"

and `allocate_explanation`:

"cannot allocate because all found copies of the shard are either stale or corrupt"

finally within `unassigned_info`-\>`details`:

""failed shard on node [\<node\_id\>]: shard failure, reason [merge failed], failure NotSerializableExceptionWrapper[merge\_exception: org.apache.lucene.index.CorruptIndexException: checksum failed (hardware problem?)..."

Calling the same on the replica returns:

`unassigned_info`-\>`reason`: "ALLOCATION\_FAILED"

`allocate_explanation`:"cannot allocate because allocation is not permitted to any of the nodes"

`unassigned_info`-\>`details`:

"failed shard on node [\<node\_id\>]: failed to perform indices:data/write/bulk[s] on replica [\<index\_name\>][\<shard\_num\>], node[\<node\_id\>], [R], s[STARTED], a[id=\<\>], failure IndexShardClosedException[CurrentState[CLOSED] Primary closed.]"

I attempted a dry run of manually reallocating the replica using the reroute API, and received a status 400 with: "[allocate\_replica] trying to allocate a replica shard [\<index\_name\>][\<shard\_num\>], while corresponding primary shard is still unassigned"

What is the best course of action here? I gather I need to assign the primary shard before I can do anything with the replica. I am concerned given the CorruptIndex exception that the primary shard (and potentially the replica too..) has suffered data losses, so my thinking was that recovering from the replica was my best bet? Is my understanding incorrect here / am I going to have to be content with data losses?

Many thanks

---

<div class="post-metadata">

### Author: ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)
#### Post date: [October 26, 2022, 1:16am UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/2 "2022-10-26T01:16:36Z")

</div>

What version are you running?

> [@tg295](#):
>
> org.apache.lucene.index.CorruptIndexException: checksum failed (hardware problem?)

That's not good and may indicate data loss. Do you have snapshots?

---

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [October 26, 2022, 8:31am UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/3 "2022-10-26T08:31:32Z")

</div>

I'm running 7.13.2. And no I do not have any snapshots...

---

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [October 28, 2022, 11:05am UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/4 "2022-10-28T11:05:15Z")

</div>

does anyone have any recommendations for how to proceed?

---

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [October 28, 2022, 11:10am UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/5 "2022-10-28T11:10:03Z")

</div>

My current thinking was to follow an approach similar to: [When everything else fails. We are using Elasticsearch on a Google… | by Remco Verhoef | Medium](https://medium.com/@remco_verhoef/when-everything-else-fails-cc8a27c679d0)

Use the CheckIndex tool to see whether the shards are actually corrupt, and then proceed from there?

---

<div class="post-metadata">

### Author: ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)
#### Post date: [October 28, 2022, 2:04pm UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/6 "2022-10-28T14:04:10Z")

</div>

> [@tg295](#):
>
> Use the CheckIndex tool to see whether the shards are actually corrupt, and then proceed from there?

You may use the `elasticsearch-shard` cli tool to see if you can recover something, here is the [documentation](https://www.elastic.co/guide/en/elasticsearch/reference/7.13/shard-tool.html).

Be aware of this warning in the documentation.

> You will lose the corrupted data when you run `elasticsearch-shard` . This tool should only be used as a last resort if there is no way to recover from another copy of the shard or restore a snapshot.

From what you shared there is not much else you can do since it appears that your index is corrupted and you may already have some level o data loss.

---

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [October 28, 2022, 3:52pm UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/7 "2022-10-28T15:52:45Z")

</div>

Thank you! Wish me luck...

---

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [November 2, 2022, 2:38pm UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/8 "2022-11-02T14:38:43Z")

</div>

The documentation explicitly says to "Stop Elasticsearch before running elasticsearch-shard."

Does this apply to the particular node or my entire cluster?

---

<div class="post-metadata">

### Author: ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)
#### Post date: [November 2, 2022, 2:47pm UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/9 "2022-11-02T14:47:41Z")

</div>

> [@tg295](#):
>
> Does this apply to the particular node or my entire cluster?

Never used this command, but I would assume that it is the Elasticsearch node that has the shard you want to try to fix.

---

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [November 4, 2022, 10:16am UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/10 "2022-11-04T10:16:23Z")

</div>

Due to the size of our cluster, stopping Elasticsearch / reallocating an entire node is quite an operation - do you know of any method that might allow us to address the issue without doing so? I assume not but worth an ask.. Many thanks for all your help by the way 🙂

---

<div class="post-metadata">

### Author: ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)
#### Post date: [November 4, 2022, 1:23pm UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/11 "2022-11-04T13:23:24Z")

</div>

> [@tg295](#):
>
> do you know of any method that might allow us to address the issue without doing so?

The `elasticsearch-shard` command is already a last resort for cases like yours, and there is no guarantee that it will work, but to be able to try it you will need to shutdown this node.

---

<div class="post-metadata">

### Author: ![tg295](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tg295/32/84798_2.png) [@tg295](https://discuss.elastic.co/u/tg295)
#### Post date: [November 4, 2022, 2:27pm UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/12 "2022-11-04T14:27:53Z")

</div>

Ok, thanks for the info and your swift reply.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [December 2, 2022, 2:28pm UTC](https://discuss.elastic.co/t/unassigned-primary-replica-shard-minimise-data-loss/317393/13 "2022-12-02T14:28:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
