# Trouble restarting after crash

**URL:** <https://discuss.elastic.co/t/trouble-restarting-after-crash/7585>\
**Category:** Elasticsearch\
**Created:** [May 7, 2012, 6:22pm UTC](https://discuss.elastic.co/t/trouble-restarting-after-crash/7585 "2012-05-07T18:22:54Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chuck\_McKenzie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chuck_mckenzie/32/19660_2.png) [@Chuck\_McKenzie](https://discuss.elastic.co/u/Chuck_McKenzie)\
**Post date:** [May 7, 2012, 6:22pm UTC](https://discuss.elastic.co/t/trouble-restarting-after-crash/7585/1 "2012-05-07T18:22:54Z")

</div>

I've inherited a several large 18.7 elasticsearch clusters and I'm  
having some trouble getting one to restart after a crash. (We ran out  
of open filehandles.) I've since upped the limit, and I'll be  
cleaning up old indices after it comes back up, so that shouldn't  
happen again, but I can't get the cluster to finish starting.

Here's the problem I'm seeing:

[2012-05-07 13:11:04,926][WARN][indices.cluster]  
[node\_name] [shard\_name][8] failed to start shard  
org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
[shard\_name][8] shard allocated for local recovery (post api), should  
exists, but doesn't  
at  
org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
99)  
at org.elasticsearch.index.gateway.IndexShardGatewayService  
$1.run(IndexShardGatewayService.java:179)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.runTask(ThreadPoolExecutor.java:886)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.run(ThreadPoolExecutor.java:908)  
at java.lang.Thread.run(Thread.java:662)

I've followed earlier instructions on this mailing list that say to  
XDELETE the affected index, but that doesn't seem to be working - it's  
been sitting for an hour as follows:

{

```
cluster_name: es_cluster1
status: red
timed_out: false
number_of_nodes: 4
number_of_data_nodes: 4
active_primary_shards: 5260
active_shards: 9719
relocating_shards: 0
initializing_shards: 6
unassigned_shards: 41

```

}

Any idea how I can get rid of the two tiny test indices that are  
having problems, without deleting several TB of data from the other  
indices?

---

<div class="post-metadata">

**Author:** ![Rafal\_Kuc\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafal_kuc_3/32/799_2.png) [@Rafal\_Kuc\_3](https://discuss.elastic.co/u/Rafal_Kuc_3)\
**Post date:** [May 7, 2012, 6:28pm UTC](https://discuss.elastic.co/t/trouble-restarting-after-crash/7585/2 "2012-05-07T18:28:16Z")

</div>

Hello!

We had similar issue to yours - did you try running XDELETE on more  
then one nodes ? We had to run XDELETE on two nodes in the cluster to  
actually have problematic indices deleted.

--  
Regards,  
Rafał Kuć  
Sematext :: [http://sematext.com/](http://sematext.com/) :: Solr - Lucene - Nutch - Elasticsearch

> I've inherited a several large 18.7 elasticsearch clusters and I'm  
> having some trouble getting one to restart after a crash. (We ran out  
> of open filehandles.) I've since upped the limit, and I'll be  
> cleaning up old indices after it comes back up, so that shouldn't  
> happen again, but I can't get the cluster to finish starting.

> Here's the problem I'm seeing:

> [2012-05-07 13:11:04,926][WARN][indices.cluster]  
> [node\_name] [shard\_name][8] failed to start shard  
> org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> [shard\_name][8] shard allocated for local recovery (post api), should  
> exists, but doesn't  
> at  
> org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> 99)  
> at org.elasticsearch.index.gateway.IndexShardGatewayService  
> $1.run(IndexShardGatewayService.java:179)  
> at java.util.concurrent.ThreadPoolExecutor  
> $Worker.runTask(ThreadPoolExecutor.java:886)  
> at java.util.concurrent.ThreadPoolExecutor  
> $Worker.run(ThreadPoolExecutor.java:908)  
> at java.lang.Thread.run(Thread.java:662)

> I've followed earlier instructions on this mailing list that say to  
> XDELETE the affected index, but that doesn't seem to be working - it's  
> been sitting for an hour as follows:

> {

> ```
> cluster_name: es_cluster1
> status: red
> timed_out: false
> number_of_nodes: 4
> number_of_data_nodes: 4
> active_primary_shards: 5260
> active_shards: 9719
> relocating_shards: 0
> initializing_shards: 6
> unassigned_shards: 41
> 
> ```

> }

> Any idea how I can get rid of the two tiny test indices that are  
> having problems, without deleting several TB of data from the other  
> indices?

---

<div class="post-metadata">

**Author:** ![Chuck\_McKenzie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chuck_mckenzie/32/19660_2.png) [@Chuck\_McKenzie](https://discuss.elastic.co/u/Chuck_McKenzie)\
**Post date:** [May 7, 2012, 6:39pm UTC](https://discuss.elastic.co/t/trouble-restarting-after-crash/7585/3 "2012-05-07T18:39:21Z")

</div>

They're running against localhost on each of the 4 nodes. Doesn't  
seem to help.

On May 7, 1:28 pm, Rafał Kuć [r....@solr.pl](mailto:r....@solr.pl) wrote:

> Hello!
> 
> We had similar issue to yours - did you try running XDELETE on more  
> then one nodes ? We had to run XDELETE on two nodes in the cluster to  
> actually have problematic indices deleted.
> 
> --  
> Regards,  
> Rafa³ Kuæ  
> Sematext ::[http://sematext.com/::](http://sematext.com/::) Solr - Lucene - Nutch - Elasticsearch
> 
> > I've inherited a several large 18.7 elasticsearch clusters and I'm  
> > having some trouble getting one to restart after a crash. (We ran out  
> > of open filehandles.) I've since upped the limit, and I'll be  
> > cleaning up old indices after it comes back up, so that shouldn't  
> > happen again, but I can't get the cluster to finish starting.  
> > Here's the problem I'm seeing:  
> > [2012-05-07 13:11:04,926][WARN][indices.cluster]  
> > [node\_name] [shard\_name][8] failed to start shard  
> > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > [shard\_name][8] shard allocated for local recovery (post api), should  
> > exists, but doesn't  
> > at  
> > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> > 99)  
> > at org.elasticsearch.index.gateway.IndexShardGatewayService  
> > $1.run(IndexShardGatewayService.java:179)  
> > at java.util.concurrent.ThreadPoolExecutor  
> > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > at java.util.concurrent.ThreadPoolExecutor  
> > $Worker.run(ThreadPoolExecutor.java:908)  
> > at java.lang.Thread.run(Thread.java:662)  
> > I've followed earlier instructions on this mailing list that say to  
> > XDELETE the affected index, but that doesn't seem to be working - it's  
> > been sitting for an hour as follows:  
> > {  
> > cluster\_name: es\_cluster1  
> > status: red  
> > timed\_out: false  
> > number\_of\_nodes: 4  
> > number\_of\_data\_nodes: 4  
> > active\_primary\_shards: 5260  
> > active\_shards: 9719  
> > relocating\_shards: 0  
> > initializing\_shards: 6  
> > unassigned\_shards: 41  
> > }  
> > Any idea how I can get rid of the two tiny test indices that are  
> > having problems, without deleting several TB of data from the other  
> > indices?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 9, 2012, 9:08am UTC](https://discuss.elastic.co/t/trouble-restarting-after-crash/7585/4 "2012-05-09T09:08:53Z")

</div>

DELETE the index will help to remove this message, this problem should be  
fixed in 0.19 with the new local gateway structure and several bug fixes  
(the fact that a shard can't recover).

On Mon, May 7, 2012 at 9:39 PM, Chuck McKenzie [redchuck@gmail.com](mailto:redchuck@gmail.com) wrote:

> They're running against localhost on each of the 4 nodes. Doesn't  
> seem to help.
> 
> On May 7, 1:28 pm, Rafał Kuć [r....@solr.pl](mailto:r....@solr.pl) wrote:
> 
> > Hello!
> > 
> > We had similar issue to yours - did you try running XDELETE on more  
> > then one nodes ? We had to run XDELETE on two nodes in the cluster to  
> > actually have problematic indices deleted.
> > 
> > --  
> > Regards,  
> > Rafa³ Kuæ  
> > Sematext ::[http://sematext.com/::](http://sematext.com/::) Solr - Lucene - Nutch - Elasticsearch
> > 
> > > I've inherited a several large 18.7 elasticsearch clusters and I'm  
> > > having some trouble getting one to restart after a crash. (We ran out  
> > > of open filehandles.) I've since upped the limit, and I'll be  
> > > cleaning up old indices after it comes back up, so that shouldn't  
> > > happen again, but I can't get the cluster to finish starting.  
> > > Here's the problem I'm seeing:  
> > > [2012-05-07 13:11:04,926][WARN][indices.cluster]  
> > > [node\_name] [shard\_name][8] failed to start shard  
> > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > [shard\_name][8] shard allocated for local recovery (post api), should  
> > > exists, but doesn't  
> > > at
> 
> org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:
> 
> > > 1. 
> > > 
> > > ```
> > > at org.elasticsearch.index.gateway.IndexShardGatewayService
> > > 
> > > ```
> > > 
> > > $1.run(IndexShardGatewayService.java:179)  
> > > at java.util.concurrent.ThreadPoolExecutor  
> > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > at java.util.concurrent.ThreadPoolExecutor  
> > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > at java.lang.Thread.run(Thread.java:662)  
> > > I've followed earlier instructions on this mailing list that say to  
> > > XDELETE the affected index, but that doesn't seem to be working - it's  
> > > been sitting for an hour as follows:  
> > > {  
> > > cluster\_name: es\_cluster1  
> > > status: red  
> > > timed\_out: false  
> > > number\_of\_nodes: 4  
> > > number\_of\_data\_nodes: 4  
> > > active\_primary\_shards: 5260  
> > > active\_shards: 9719  
> > > relocating\_shards: 0  
> > > initializing\_shards: 6  
> > > unassigned\_shards: 41  
> > > }  
> > > Any idea how I can get rid of the two tiny test indices that are  
> > > having problems, without deleting several TB of data from the other  
> > > indices?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:29am UTC](https://discuss.elastic.co/t/trouble-restarting-after-crash/7585/5 "2017-07-06T03:29:43Z")

</div>


