# Index relocation failing at 100% bytes

**URL:** <https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934>\
**Category:** Elasticsearch\
**Created:** [January 8, 2025, 12:28pm UTC](https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934 "2025-01-08T12:28:43Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [January 8, 2025, 12:28pm UTC](https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934/1 "2025-01-08T12:28:43Z")

</div>

After upgrade from 7.17.23 to 7.17.26 I started seeing patterns like this:

- ILM starts relocating from warm tier to cold
- in `/_cat/allocation` transfer goes to 100% bp
- nothing in log of source node
- error in log of target node
- repeat

```auto

[2025-01-08T13:07:16,075][WARN][o.e.i.c.IndicesClusterStateService] [Cold2] [set1_228][1] marking and sending shard failed due to [failed recovery]
org.elasticsearch.indices.recovery.RecoveryFailedException: [set1_228][1]: Recovery failed from {Warm1}{id-removed}{id-removed}{192.168.0.1}{192.168.0.1:9300}{hiw}{xpack.installed=true, transform.node=false} into {Cold2}{id-removed}{id-removed}{192.168.0.12}{192.168.0.12:9300}{cmv}{xpack.installed=true, transform.node=false} (failed to retry recovery)
        at org.elasticsearch.indices.recovery.RecoveriesCollection.resetRecovery(RecoveriesCollection.java:137) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.PeerRecoveryTargetService.retryRecovery(PeerRecoveryTargetService.java:199) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.PeerRecoveryTargetService.retryRecovery(PeerRecoveryTargetService.java:195) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.PeerRecoveryTargetService$RecoveryResponseHandler.handleException(PeerRecoveryTargetService.java:767) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.transport.TransportService$ContextRestoreResponseHandler.handleException(TransportService.java:1481) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.transport.TransportService$ContextRestoreResponseHandler.handleException(TransportService.java:1481) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.transport.InboundHandler.lambda$handleException$3(InboundHandler.java:380) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingRunnable.run(ThreadContext.java:718) [elasticsearch-7.17.26.jar:7.17.26]
        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1144) [?:?]
        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:642) [?:?]
        at java.lang.Thread.run(Thread.java:1570) [?:?]
Caused by: java.lang.IllegalStateException: cannot reset recovery as previous attempt made it past finalization step
        at org.elasticsearch.indices.recovery.RecoveryTarget.resetRecovery(RecoveryTarget.java:241) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.RecoveriesCollection.resetRecovery(RecoveriesCollection.java:114) ~[elasticsearch-7.17.26.jar:7.17.26]
        ... 10 more
[2025-01-08T13:13:57,339][WARN][o.e.i.c.IndicesClusterStateService] [Cold2] [set2_228][1] marking and sending shard failed due to [failed recovery]
org.elasticsearch.indices.recovery.RecoveryFailedException: [set2_228][1]: Recovery failed from {Warm1}{id-removed}{id-removed}{192.168.0.1}{192.168.0.1:9300}{hiw}{xpack.installed=true, transform.node=false} into {Cold2}{id-removed}{id-removed}{192.168.0.12}{192.168.0.12:9300}{cmv}{xpack.installed=true, transform.node=false} (failed to retry recovery)
        at org.elasticsearch.indices.recovery.RecoveriesCollection.resetRecovery(RecoveriesCollection.java:137) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.PeerRecoveryTargetService.retryRecovery(PeerRecoveryTargetService.java:199) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.PeerRecoveryTargetService.retryRecovery(PeerRecoveryTargetService.java:195) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.PeerRecoveryTargetService$RecoveryResponseHandler.handleException(PeerRecoveryTargetService.java:767) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.transport.TransportService$ContextRestoreResponseHandler.handleException(TransportService.java:1481) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.transport.TransportService$ContextRestoreResponseHandler.handleException(TransportService.java:1481) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.transport.InboundHandler.lambda$handleException$3(InboundHandler.java:380) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingRunnable.run(ThreadContext.java:718) [elasticsearch-7.17.26.jar:7.17.26]
        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1144) [?:?]
        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:642) [?:?]
        at java.lang.Thread.run(Thread.java:1570) [?:?]
Caused by: java.lang.IllegalStateException: cannot reset recovery as previous attempt made it past finalization step
        at org.elasticsearch.indices.recovery.RecoveryTarget.resetRecovery(RecoveryTarget.java:241) [elasticsearch-7.17.26.jar:7.17.26]
        at org.elasticsearch.indices.recovery.RecoveriesCollection.resetRecovery(RecoveriesCollection.java:114) ~[elasticsearch-7.17.26.jar:7.17.26]
        ... 10 more

```

Internet search for "cannot reset recovery as previous attempt made it past finalization step" gives no results for my profile.

Updates

- Relocating another index from Cold back to Warm works fine
- Relocating another index from different Warm node to Cold works fine
- Relocating affected index from Warm node to different warm node fails
- Relocating affected shard to Cold still fails today
- Relocating different shard of same index to affected Warm node works
- Relocating this different shard to other Warm node works
- Relocating original affected shard to same Warn node fails

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [January 8, 2025, 12:41pm UTC](https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934/2 "2025-01-08T12:41:17Z")

</div>

Looks like it is currently about 4 shards stuck like this. This is more that outgoing recovery limit, so there is always 2-3 shards transfering and one waiting. It looks like because of this each transfer starts from zero?

I override all but one back to warm tier and now it looks like cold node can keep transfered data between attemps. This speeds up loop to like 10 attempts per second. Error is same.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [January 9, 2025, 11:12am UTC](https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934/3 "2025-01-09T11:12:34Z")

</div>

So it looks like it is particular shards that are affected by this.  
Other shards of the same index and other indexes move fine.

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [January 10, 2025, 3:03pm UTC](https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934/4 "2025-01-10T15:03:06Z")

</div>

Decided to reboot node hosting affected shards.

No errors in logs but replicas did not want to synchronize and stuck in INITIALIZING state for more than normal. I was worried but after about 10 minutes it finally synchronized.

- Trying to move originally affected shard again now

---

<div class="post-metadata">

**Author:** ![nisow95612](https://avatars.discourse-cdn.com/v4/letter/n/3d9bf3/32.png) [@nisow95612](https://discuss.elastic.co/u/nisow95612)\
**Post date:** [January 13, 2025, 9:59am UTC](https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934/5 "2025-01-13T09:59:21Z")

</div>

Good news. After that restart indexes move with no problems.

---

<div class="post-metadata">

**Author:** ![jessgarson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jessgarson/32/129841_2.png) [@jessgarson](https://discuss.elastic.co/u/jessgarson)\
**Post date:** [January 13, 2025, 11:09pm UTC](https://discuss.elastic.co/t/index-relocation-failing-at-100-bytes/372934/6 "2025-01-13T23:09:54Z")

</div>

Thanks so much for sharing these troubleshooting details, @nisow95612.
