# Failed shard recovery after hard shutdown

**URL:** https://discuss.elastic.co/t/failed-shard-recovery-after-hard-shutdown/160998
**Category:** Elasticsearch
**Created:** [December 15, 2018, 12:18pm UTC](https://discuss.elastic.co/t/failed-shard-recovery-after-hard-shutdown/160998 "2018-12-15T12:18:36Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![kamal](https://avatars.discourse-cdn.com/v4/letter/k/a698b9/32.png) [@kamal](https://discuss.elastic.co/u/kamal)
#### Post date: [December 15, 2018, 12:18pm UTC](https://discuss.elastic.co/t/failed-shard-recovery-after-hard-shutdown/160998/1 "2018-12-15T12:18:36Z")

</div>

Hi  
I was doing a huge indexing job(about 400 billion records), but suddenly one of the nodes went down because the power failed, after fixing the power, one shard is missing, and the error is:

- shard failure, reason [failed to recover from translog], failure EngineException, nested: EOFException[read past EOF. pos [4590678] length: [4] end: [4590678].
- cannot allocate because allocation is not permitted to any of nodes that hold an in-sync shard copy.

As it was a huge indexing, (and still is running very slow after the problem), there is no replica.  
What should I do?

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [December 15, 2018, 3:49pm UTC](https://discuss.elastic.co/t/failed-shard-recovery-after-hard-shutdown/160998/2 "2018-12-15T15:49:36Z")

</div>

The translog was not properly written to disk before the power outage and is now corrupt. Did you set `index.translog.durability: async`? If not, my guess is that your storage hardware does not properly support the `fsync()` call, claiming to have persisted some writes before actually having done so.

The shard in question is broken, and the only truly reliable way forwards is to start again. You can wipe out the corrupt translog using [the `elasticsearch-translog` tool](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-modules-translog.html#corrupt-translog-truncation) (or [`elasticsearch-shard` if in 6.5 or later](https://www.elastic.co/guide/en/elasticsearch/reference/6.5/shard-tool.html)) which will lose any writes that were not also written to Lucene. There's no way to tell which writes will be lost, unless you can somehow compare the data in Elasticsearch to your source data and fix it up.

---

<div class="post-metadata">

### Author: ![kamal](https://avatars.discourse-cdn.com/v4/letter/k/a698b9/32.png) [@kamal](https://discuss.elastic.co/u/kamal)
#### Post date: [December 18, 2018, 7:24am UTC](https://discuss.elastic.co/t/failed-shard-recovery-after-hard-shutdown/160998/4 "2018-12-18T07:24:51Z")

</div>

I didn't set `index.translog.durability: async`, and the filesystem is ext4 and disks are raid 10.  
So what should I do for this so it will not happen again?

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [December 18, 2018, 8:07am UTC](https://discuss.elastic.co/t/failed-shard-recovery-after-hard-shutdown/160998/5 "2018-12-18T08:07:04Z")

</div>

> [@kamal](#):
>
> So what should I do for this so it will not happen again?

As I said, my guess is that your storage hardware does not properly support the `fsync()` call. This is often due to a misconfiguration: write caching is sometimes enabled for performance reasons but this breaks `fsync()` unless all such caches are battery-backed. A simple way to check for this kind of problem is [described in this article](https://brad.livejournal.com/2116715.html).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [January 15, 2019, 8:07am UTC](https://discuss.elastic.co/t/failed-shard-recovery-after-hard-shutdown/160998/6 "2019-01-15T08:07:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
