# Recovery failed, no activity after 30m, networking seems fine

**URL:** <https://discuss.elastic.co/t/recovery-failed-no-activity-after-30m-networking-seems-fine/106540>\
**Category:** Elasticsearch\
**Created:** [November 6, 2017, 2:27pm UTC](https://discuss.elastic.co/t/recovery-failed-no-activity-after-30m-networking-seems-fine/106540 "2017-11-06T14:27:15Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ndtreviv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ndtreviv/32/22494_2.png) [@ndtreviv](https://discuss.elastic.co/u/ndtreviv)\
**Post date:** [November 6, 2017, 2:27pm UTC](https://discuss.elastic.co/t/recovery-failed-no-activity-after-30m-networking-seems-fine/106540/1 "2017-11-06T14:27:15Z")

</div>

Last night I started restoring a snapshot of 12 small indices to a 6-node cluster (v.2.3.1) running on m4.2xlarge AWS instance types, each with 2TB disk.

The shards won't allocate to any node.

Every node has plenty of available disk (at last 1TB each).  
All nodes are master-eligible, data nodes that also accept client comms. They all have marvel and kibana running on them.

The errors I'm seeing in the logs look like this:

```auto
[2017-11-06 04:20:55,982][WARN][indices.cluster] [es-live-6] [[metrics-persistent-2017-01][4]] marking and sending shard failed due to [failed recovery]
RecoveryFailedException[[metrics-persistent-2017-01][4]: Recovery failed from {es-live-5}{nJSjNL5GRo-N-cI0YH5RbQ}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true} into {es-live-6}{PGXmWTURQ1eQZS_5E1k2Kw}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true} (no activity after [30m])]; nested: ElasticsearchTimeoutException[no activity after [30m]];
	at org.elasticsearch.indices.recovery.RecoveriesCollection$RecoveryMonitor.doRun(RecoveriesCollection.java:236)
	at org.elasticsearch.common.util.concurrent.AbstractRunnable.run(AbstractRunnable.java:37)
	at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
	at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
	at java.lang.Thread.run(Thread.java:748)
Caused by: ElasticsearchTimeoutException[no activity after [30m]]

```

I can successfully cURL the addresses and ports (obvs port 9300 just returns with `This is not a HTTP port`, as expected), so assuming that networking is fine given that fact.

All nodes are showing very high `ram.percent` (\>90 on all of them), but average `heap.percent` (around the 50s).  
netadata shows the system RAM usage as not being particularly high.

Hot threads look like this:

```auto
::: {es-live-7}{OPbLZrE8RHyC0-Eva3WaGQ}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true}
   Hot threads at 2017-11-06T14:13:01.633Z, interval=500ms, busiestThreads=3, ignoreIdleThreads=true:
   
    0.0% (99.3micros out of 500ms) cpu usage by thread 'elasticsearch[es-live-7][transport_client_timer][T#1]{Hashed wheel timer #1}'
     10/10 snapshots sharing following 5 elements
       java.lang.Thread.sleep(Native Method)
       org.jboss.netty.util.HashedWheelTimer$Worker.waitForNextTick(HashedWheelTimer.java:445)
       org.jboss.netty.util.HashedWheelTimer$Worker.run(HashedWheelTimer.java:364)
       org.jboss.netty.util.ThreadRenamingRunnable.run(ThreadRenamingRunnable.java:108)
       java.lang.Thread.run(Thread.java:748)

::: {es-live-9}{r79b4XstTAm0UKfzdxHaOw}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true}
   Hot threads at 2017-11-06T14:13:01.634Z, interval=500ms, busiestThreads=3, ignoreIdleThreads=true:
   
    0.0% (76.4micros out of 500ms) cpu usage by thread 'elasticsearch[es-live-9][transport_client_timer][T#1]{Hashed wheel timer #1}'
     10/10 snapshots sharing following 5 elements
       java.lang.Thread.sleep(Native Method)
       org.jboss.netty.util.HashedWheelTimer$Worker.waitForNextTick(HashedWheelTimer.java:445)
       org.jboss.netty.util.HashedWheelTimer$Worker.run(HashedWheelTimer.java:364)
       org.jboss.netty.util.ThreadRenamingRunnable.run(ThreadRenamingRunnable.java:108)
       java.lang.Thread.run(Thread.java:748)

::: {es-live-8}{0lcMCUmdRwqr6NaAau5Ehg}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true}
   Hot threads at 2017-11-06T14:13:01.633Z, interval=500ms, busiestThreads=3, ignoreIdleThreads=true:
   
    0.0% (71.7micros out of 500ms) cpu usage by thread 'elasticsearch[es-live-8][transport_client_timer][T#1]{Hashed wheel timer #1}'
     10/10 snapshots sharing following 5 elements
       java.lang.Thread.sleep(Native Method)
       org.jboss.netty.util.HashedWheelTimer$Worker.waitForNextTick(HashedWheelTimer.java:445)
       org.jboss.netty.util.HashedWheelTimer$Worker.run(HashedWheelTimer.java:364)
       org.jboss.netty.util.ThreadRenamingRunnable.run(ThreadRenamingRunnable.java:108)
       java.lang.Thread.run(Thread.java:748)

::: {es-live-5}{nJSjNL5GRo-N-cI0YH5RbQ}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true}
   Hot threads at 2017-11-06T14:13:01.631Z, interval=500ms, busiestThreads=3, ignoreIdleThreads=true:
   
    0.0% (94.1micros out of 500ms) cpu usage by thread 'elasticsearch[es-live-5][transport_client_timer][T#1]{Hashed wheel timer #1}'
     10/10 snapshots sharing following 5 elements
       java.lang.Thread.sleep(Native Method)
       org.jboss.netty.util.HashedWheelTimer$Worker.waitForNextTick(HashedWheelTimer.java:445)
       org.jboss.netty.util.HashedWheelTimer$Worker.run(HashedWheelTimer.java:364)
       org.jboss.netty.util.ThreadRenamingRunnable.run(ThreadRenamingRunnable.java:108)
       java.lang.Thread.run(Thread.java:748)

::: {es-live-6}{PGXmWTURQ1eQZS_5E1k2Kw}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true}
   Hot threads at 2017-11-06T14:13:01.633Z, interval=500ms, busiestThreads=3, ignoreIdleThreads=true:

::: {es-live-10}{29trlJX5QoiCub0xEon4xw}{XXX.XXX.XXX.XXX}{XXX.XXX.XXX.XXX:9300}{master=true}
   Hot threads at 2017-11-06T14:13:01.633Z, interval=500ms, busiestThreads=3, ignoreIdleThreads=true:

```

I've even tried running a script to manually route the shards to a node. It allocates the shards, they spend an inordinate amount of time in the `INITIALIZING` state, and then eventually go back to `UNASSIGNED`.

Some of these indices are only a few KB.

What is going on here? Why won't the shards allocate?

Thanks for any help!

---

<div class="post-metadata">

**Author:** ![ndtreviv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ndtreviv/32/22494_2.png) [@ndtreviv](https://discuss.elastic.co/u/ndtreviv)\
**Post date:** [November 6, 2017, 2:35pm UTC](https://discuss.elastic.co/t/recovery-failed-no-activity-after-30m-networking-seems-fine/106540/2 "2017-11-06T14:35:22Z")

</div>

I'm also seeing this sort of message a lot in the logs on at least 3 of the nodes:

```auto
[2017-11-06 14:34:18,928][DEBUG][index.shard] [es-live-10] [metrics-persistent-2017-11][2] updateBufferSize: engine is closed; skipping

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 4, 2017, 2:35pm UTC](https://discuss.elastic.co/t/recovery-failed-no-activity-after-30m-networking-seems-fine/106540/3 "2017-12-04T14:35:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
