# Failed to start shard

**URL:** <https://discuss.elastic.co/t/failed-to-start-shard/5656>\
**Category:** Elasticsearch\
**Created:** [October 21, 2011, 9:53am UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656 "2011-10-21T09:53:27Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![gnulinux](https://avatars.discourse-cdn.com/v4/letter/g/6de8d8/32.png) [@gnulinux](https://discuss.elastic.co/u/gnulinux)\
**Post date:** [October 21, 2011, 9:53am UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/1 "2011-10-21T09:53:27Z")

</div>

Hi

I am evaluating ElasticSearch (0.17.8) for a spatial search platform.  
I was able to setup a two-node cluster and everything was working  
fine. But after rebooting both the nodes, I am getting the following  
error on both.

[2011-10-19 06:02:45,243][WARN][indices.cluster] [linux  
Ubu2] [books][1] failed to start shard  
org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
[books][1] shard allocated for local recovery (post api), should  
exists, but doesn't  
at  
org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
99)  
at org.elasticsearch.index.gateway.IndexShardGatewayService  
$1.run(IndexShardGatewayService.java:179)  
at  
java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
1110)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.run(ThreadPoolExecutor.java:603)  
at java.lang.Thread.run(Thread.java:679)

Config Files:

Node01 (Master)

cluster:  
name: gnulinux

node.name: "linux Ubu2"  
node.master: true  
node.data: true  
node.rack: rack01

network:  
bindHost: 192.168.2.10  
publishHost: 192.168.2.10

index.engine.robin.refreshInterval: -1  
index.gateway.snapshot\_interval: -1  
index.gateway.type: local  
index.number\_of\_shards: 5  
index.number\_of\_replicas: 1

gateway.recover\_after\_nodes: 2  
gateway.recover\_after\_time: 5m  
gateway.expected\_nodes: 2  
cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
cluster.routing.allocation.node\_concurrent\_recoveries: 2  
indices.recovery.concurrent\_streams: 5

index:  
store:  
fs:  
memory:  
enabled: true  
discovery:  
jgroups:  
config: tcp  
bind\_port: 9700  
bind\_address: 192.168.2.10  
tcpping:  
initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

Node02:

cluster:  
name: gnulinux

node.name: "linux Ubu1"  
node.master: false  
node.data: true  
node.rack: rack01

network:  
bindHost: 192.168.2.11  
publishHost: 192.168.2.11

index.engine.robin.refreshInterval: -1  
index.gateway.snapshot\_interval: -1  
index.gateway.type: local  
index.number\_of\_shards: 5  
index.number\_of\_replicas: 1

gateway.recover\_after\_nodes: 2  
gateway.recover\_after\_time: 5m  
gateway.expected\_nodes: 2  
cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
cluster.routing.allocation.node\_concurrent\_recoveries: 2  
indices.recovery.concurrent\_streams: 5

index:  
store:  
fs:  
memory:  
enabled: true  
discovery:  
jgroups:  
config: tcp  
bind\_port: 9700  
bind\_address: 192.168.2.11  
tcpping:  
initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 21, 2011, 6:27pm UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/2 "2011-10-21T18:27:01Z")

</div>

Are you sure the two nodes find each other? The configuration you have  
configure jgroups for discovery, which was removed in version 0.6 ....

On Fri, Oct 21, 2011 at 11:53 AM, gnulinux [vijivijayakumar@gmail.com](mailto:vijivijayakumar@gmail.com)wrote:

> Hi
> 
> I am evaluating Elasticsearch (0.17.8) for a spatial search platform.  
> I was able to setup a two-node cluster and everything was working  
> fine. But after rebooting both the nodes, I am getting the following  
> error on both.
> 
> [2011-10-19 06:02:45,243][WARN][indices.cluster] [linux  
> Ubu2] [books][1] failed to start shard  
> org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> [books][1] shard allocated for local recovery (post api), should  
> exists, but doesn't  
> at
> 
> org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> 99)  
> at org.elasticsearch.index.gateway.IndexShardGatewayService  
> $1.run(IndexShardGatewayService.java:179)  
> at  
> java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
> 1110)  
> at java.util.concurrent.ThreadPoolExecutor  
> $Worker.run(ThreadPoolExecutor.java:603)  
> at java.lang.Thread.run(Thread.java:679)
> 
> Config Files:
> 
> Node01 (Master)
> 
> cluster:  
> name: gnulinux
> 
> node.name: "linux Ubu2"  
> node.master: true  
> node.data: true  
> node.rack: rack01
> 
> network:  
> bindHost: 192.168.2.10  
> publishHost: 192.168.2.10
> 
> index.engine.robin.refreshInterval: -1  
> index.gateway.snapshot\_interval: -1  
> index.gateway.type: local  
> index.number\_of\_shards: 5  
> index.number\_of\_replicas: 1
> 
> gateway.recover\_after\_nodes: 2  
> gateway.recover\_after\_time: 5m  
> gateway.expected\_nodes: 2  
> cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> indices.recovery.concurrent\_streams: 5
> 
> index:  
> store:  
> fs:  
> memory:  
> enabled: true  
> discovery:  
> jgroups:  
> config: tcp  
> bind\_port: 9700  
> bind\_address: 192.168.2.10  
> tcpping:  
> initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]
> 
> Node02:
> 
> cluster:  
> name: gnulinux
> 
> node.name: "linux Ubu1"  
> node.master: false  
> node.data: true  
> node.rack: rack01
> 
> network:  
> bindHost: 192.168.2.11  
> publishHost: 192.168.2.11
> 
> index.engine.robin.refreshInterval: -1  
> index.gateway.snapshot\_interval: -1  
> index.gateway.type: local  
> index.number\_of\_shards: 5  
> index.number\_of\_replicas: 1
> 
> gateway.recover\_after\_nodes: 2  
> gateway.recover\_after\_time: 5m  
> gateway.expected\_nodes: 2  
> cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> indices.recovery.concurrent\_streams: 5
> 
> index:  
> store:  
> fs:  
> memory:  
> enabled: true  
> discovery:  
> jgroups:  
> config: tcp  
> bind\_port: 9700  
> bind\_address: 192.168.2.11  
> tcpping:  
> initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

---

<div class="post-metadata">

**Author:** ![Viji\_Nair](https://avatars.discourse-cdn.com/v4/letter/v/7ab992/32.png) [@Viji\_Nair](https://discuss.elastic.co/u/Viji_Nair)\
**Post date:** [October 21, 2011, 8:09pm UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/3 "2011-10-21T20:09:57Z")

</div>

Hi,

Yes, I missed it. But I don't know how, the cluster API was reporting  
"green" and showing the total number of nodes as "2"

Now, I changed to zen discovery and deleted all the existing indexes. The  
node discovery happens properly, but after adding index (this time followed  
the twitter example) and putting some data, a subsequent reboot is giving  
the same issue. Please find the steps I have followed.

1. Installed Java

# java -version

java version "1.6.0\_29"  
Java(TM) SE Runtime Environment (build 1.6.0\_29-b11)  
Java HotSpot(TM) 64-Bit Server VM (build 20.4-b02, mixed mode)

1. This is an ubuntu 64 bit machine (11.04)

# uname -a

Linux ubu-ser 2.6.38-8-generic #42-Ubuntu SMP Mon Apr 11 03:31:24 UTC 2011  
x86\_64 x86\_64 x86\_64 GNU/Linux

1. Installed Elastic Search 0.17.9 and Service Wrapper.

2. Configured a two node cluster

- 

Node01 Configuration\*

# cat /root/elasticsearch-0.17.8/config/elasticsearch.yml

cluster:  
name: gnulinux

node.name: "Ubu2"  
node.master: true  
node.data: true  
node.rack: rack01

network:  
bindHost: 192.168.2.10  
publishHost: 192.168.2.10

index.engine.robin.refreshInterval: -1  
index.gateway.snapshot\_interval: -1  
index.gateway.type: local  
index.number\_of\_shards: 5  
index.number\_of\_replicas: 1

gateway.recover\_after\_nodes: 2  
gateway.recover\_after\_time: 5m  
gateway.expected\_nodes: 2  
cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
cluster.routing.allocation.node\_concurrent\_recoveries: 2  
indices.recovery.concurrent\_streams: 5

index:  
store:  
fs:  
memory:  
enabled: true

discovery:  
zen:  
ping\_timeout: 30s  
ping:  
multicast:  
enabled: false  
unicast:  
enabled: true  
hosts: 192.168.2.10, 192.168.2.11  
fd:  
ping\_retries: 10  
ping\_interval: 5s  
ping\_timeout: 30s

_Node02 Configuration_

#cat /root/elasticsearch-0.17.8/config/elasticsearch.yml  
cluster:  
name: gnulinux

node.name: "Ubu1"  
node.master: false  
node.data: true  
node.rack: rack01

network:  
bindHost: 192.168.2.11  
publishHost: 192.168.2.11

index.engine.robin.refreshInterval: -1  
index.gateway.snapshot\_interval: -1  
index.gateway.type: local  
index.number\_of\_shards: 5  
index.number\_of\_replicas: 1

gateway.recover\_after\_nodes: 2  
gateway.recover\_after\_time: 5m  
gateway.expected\_nodes: 2  
cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
cluster.routing.allocation.node\_concurrent\_recoveries: 2  
indices.recovery.concurrent\_streams: 5

index:  
store:  
fs:  
memory:  
enabled: true

discovery:  
zen:  
ping\_timeout: 30s  
ping:  
multicast:  
enabled: false  
unicast:  
enabled: true  
hosts: 192.168.2.10, 192.168.2.11  
fd:  
ping\_retries: 10  
ping\_interval: 5s  
ping\_timeout: 30s

1. Started both the nodes and checked the status, both were up. Verified the  
log file as well. Everything was fine till this step.

# curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'

{  
"cluster\_name" : "gnulinux",  
"status" : "green",  
"timed\_out" : false,  
"number\_of\_nodes" : 2,  
"number\_of\_data\_nodes" : 2,  
"active\_primary\_shards" : 0,  
"active\_shards" : 0,  
"relocating\_shards" : 0,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 0  
}

# curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'

{  
"cluster\_name" : "gnulinux",  
"status" : "green",  
"timed\_out" : false,  
"number\_of\_nodes" : 2,  
"number\_of\_data\_nodes" : 2,  
"active\_primary\_shards" : 0,  
"active\_shards" : 0,  
"relocating\_shards" : 0,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 0  
}

1. Added some data and restarted the nodes, the cluster status is red and  
log file is giving the same error.

# curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'

{  
"cluster\_name" : "gnulinux",  
"status" : "red",  
"timed\_out" : false,  
"number\_of\_nodes" : 2,  
"number\_of\_data\_nodes" : 2,  
"active\_primary\_shards" : 0,  
"active\_shards" : 0,  
"relocating\_shards" : 0,  
"initializing\_shards" : 5,  
"unassigned\_shards" : 5  
}

# curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'

{  
"cluster\_name" : "gnulinux",  
"status" : "red",  
"timed\_out" : false,  
"number\_of\_nodes" : 2,  
"number\_of\_data\_nodes" : 2,  
"active\_primary\_shards" : 0,  
"active\_shards" : 0,  
"relocating\_shards" : 0,  
"initializing\_shards" : 5,  
"unassigned\_shards" : 5

[2011-10-22 01:18:46,703][WARN][cluster.action.shard] [Ubu1] sending  
failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
s[INITIALIZING], reason [Failed to start shard, message  
[IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
recovery (post api), should exists, but doesn't]]]  
[2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1] sending  
failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
s[INITIALIZING], reason [Failed to start shard, message  
[IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
recovery (post api), should exists, but doesn't]]]  
[2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1] sending  
failed shard for [twitter][0], node[fgELpN11R6m2XbIKdLHYgg], [P],  
s[INITIALIZING], reason [Failed to start shard, message  
[IndexShardGatewayRecoveryException[[twitter][0] shard allocated for local  
recovery (post api), should exists, but doesn't]]]  
[2011-10-22 01:18:46,725][WARN][indices.cluster] [Ubu1]  
[twitter][1] failed to start shard  
org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
[twitter][1] shard allocated for local recovery (post api), should exists,  
but doesn't  
at  
org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
at  
org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
at  
java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
at  
java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
at java.lang.Thread.run(Thread.java:662)  
[2011-10-22 01:18:46,736][WARN][indices.cluster] [Ubu1]  
[twitter][4] failed to start shard  
org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
[twitter][4] shard allocated for local recovery (post api), should exists,  
but doesn't  
at  
org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
at  
org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
at  
java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
at  
java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
at java.lang.Thread.run(Thread.java:662)  
[2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1] sending  
failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
s[INITIALIZING], reason [Failed to start shard, message  
[IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
recovery (post api), should exists, but doesn't]]]  
[2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1] sending  
failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
s[INITIALIZING], reason [Failed to start shard, message  
[IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
recovery (post api), should exists, but doesn't]]]

Thanks  
Viji

On Fri, Oct 21, 2011 at 11:57 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Are you sure the two nodes find each other? The configuration you have  
> configure jgroups for discovery, which was removed in version 0.6 ....
> 
> On Fri, Oct 21, 2011 at 11:53 AM, gnulinux [vijivijayakumar@gmail.com](mailto:vijivijayakumar@gmail.com)wrote:
> 
> > Hi
> > 
> > I am evaluating Elasticsearch (0.17.8) for a spatial search platform.  
> > I was able to setup a two-node cluster and everything was working  
> > fine. But after rebooting both the nodes, I am getting the following  
> > error on both.
> > 
> > [2011-10-19 06:02:45,243][WARN][indices.cluster] [linux  
> > Ubu2] [books][1] failed to start shard  
> > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > [books][1] shard allocated for local recovery (post api), should  
> > exists, but doesn't  
> > at
> > 
> > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> > 99)  
> > at org.elasticsearch.index.gateway.IndexShardGatewayService  
> > $1.run(IndexShardGatewayService.java:179)  
> > at  
> > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
> > 1110)  
> > at java.util.concurrent.ThreadPoolExecutor  
> > $Worker.run(ThreadPoolExecutor.java:603)  
> > at java.lang.Thread.run(Thread.java:679)
> > 
> > Config Files:
> > 
> > Node01 (Master)
> > 
> > cluster:  
> > name: gnulinux
> > 
> > node.name: "linux Ubu2"  
> > node.master: true  
> > node.data: true  
> > node.rack: rack01
> > 
> > network:  
> > bindHost: 192.168.2.10  
> > publishHost: 192.168.2.10
> > 
> > index.engine.robin.refreshInterval: -1  
> > index.gateway.snapshot\_interval: -1  
> > index.gateway.type: local  
> > index.number\_of\_shards: 5  
> > index.number\_of\_replicas: 1
> > 
> > gateway.recover\_after\_nodes: 2  
> > gateway.recover\_after\_time: 5m  
> > gateway.expected\_nodes: 2  
> > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > indices.recovery.concurrent\_streams: 5
> > 
> > index:  
> > store:  
> > fs:  
> > memory:  
> > enabled: true  
> > discovery:  
> > jgroups:  
> > config: tcp  
> > bind\_port: 9700  
> > bind\_address: 192.168.2.10  
> > tcpping:  
> > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]
> > 
> > Node02:
> > 
> > cluster:  
> > name: gnulinux
> > 
> > node.name: "linux Ubu1"  
> > node.master: false  
> > node.data: true  
> > node.rack: rack01
> > 
> > network:  
> > bindHost: 192.168.2.11  
> > publishHost: 192.168.2.11
> > 
> > index.engine.robin.refreshInterval: -1  
> > index.gateway.snapshot\_interval: -1  
> > index.gateway.type: local  
> > index.number\_of\_shards: 5  
> > index.number\_of\_replicas: 1
> > 
> > gateway.recover\_after\_nodes: 2  
> > gateway.recover\_after\_time: 5m  
> > gateway.expected\_nodes: 2  
> > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > indices.recovery.concurrent\_streams: 5
> > 
> > index:  
> > store:  
> > fs:  
> > memory:  
> > enabled: true  
> > discovery:  
> > jgroups:  
> > config: tcp  
> > bind\_port: 9700  
> > bind\_address: 192.168.2.11  
> > tcpping:  
> > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 21, 2011, 11:53pm UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/4 "2011-10-21T23:53:04Z")

</div>

Are you sure that you don't delete the index content between restarts?

On Fri, Oct 21, 2011 at 10:09 PM, Viji Nair [viji@linux.com](mailto:viji@linux.com) wrote:

> Hi,
> 
> Yes, I missed it. But I don't know how, the cluster API was reporting  
> "green" and showing the total number of nodes as "2"
> 
> Now, I changed to zen discovery and deleted all the existing indexes. The  
> node discovery happens properly, but after adding index (this time followed  
> the twitter example) and putting some data, a subsequent reboot is giving  
> the same issue. Please find the steps I have followed.
> 
> 1. Installed Java
> 
> # java -version
> 
> java version "1.6.0\_29"  
> Java(TM) SE Runtime Environment (build 1.6.0\_29-b11)  
> Java HotSpot(TM) 64-Bit Server VM (build 20.4-b02, mixed mode)
> 
> 1. This is an ubuntu 64 bit machine (11.04)
> 
> # uname -a
> 
> Linux ubu-ser 2.6.38-8-generic #42-Ubuntu SMP Mon Apr 11 03:31:24 UTC 2011  
> x86\_64 x86\_64 x86\_64 GNU/Linux
> 
> 1. Installed Elastic Search 0.17.9 and Service Wrapper.
> 
> 2. Configured a two node cluster
> 
> - 
> 
> Node01 Configuration\*
> 
> # cat /root/elasticsearch-0.17.8/config/elasticsearch.yml
> 
> cluster:  
> name: gnulinux
> 
> node.name: "Ubu2"
> 
> node.master: true  
> node.data: true  
> node.rack: rack01
> 
> network:  
> bindHost: 192.168.2.10  
> publishHost: 192.168.2.10
> 
> index.engine.robin.refreshInterval: -1  
> index.gateway.snapshot\_interval: -1  
> index.gateway.type: local  
> index.number\_of\_shards: 5  
> index.number\_of\_replicas: 1
> 
> gateway.recover\_after\_nodes: 2  
> gateway.recover\_after\_time: 5m  
> gateway.expected\_nodes: 2  
> cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> indices.recovery.concurrent\_streams: 5
> 
> index:  
> store:  
> fs:  
> memory:  
> enabled: true
> 
> discovery:  
> zen:  
> ping\_timeout: 30s  
> ping:  
> multicast:  
> enabled: false  
> unicast:  
> enabled: true  
> hosts: 192.168.2.10, 192.168.2.11  
> fd:  
> ping\_retries: 10  
> ping\_interval: 5s  
> ping\_timeout: 30s
> 
> _Node02 Configuration_
> 
> #cat /root/elasticsearch-0.17.8/config/elasticsearch.yml  
> cluster:  
> name: gnulinux
> 
> node.name: "Ubu1"
> 
> node.master: false  
> node.data: true  
> node.rack: rack01
> 
> network:  
> bindHost: 192.168.2.11  
> publishHost: 192.168.2.11
> 
> index.engine.robin.refreshInterval: -1  
> index.gateway.snapshot\_interval: -1  
> index.gateway.type: local  
> index.number\_of\_shards: 5  
> index.number\_of\_replicas: 1
> 
> gateway.recover\_after\_nodes: 2  
> gateway.recover\_after\_time: 5m  
> gateway.expected\_nodes: 2  
> cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> indices.recovery.concurrent\_streams: 5
> 
> index:  
> store:  
> fs:  
> memory:  
> enabled: true
> 
> discovery:  
> zen:  
> ping\_timeout: 30s  
> ping:  
> multicast:  
> enabled: false  
> unicast:  
> enabled: true  
> hosts: 192.168.2.10, 192.168.2.11  
> fd:  
> ping\_retries: 10  
> ping\_interval: 5s  
> ping\_timeout: 30s
> 
> 1. Started both the nodes and checked the status, both were up. Verified  
> the log file as well. Everything was fine till this step.
> 
> # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> 
> {  
> "cluster\_name" : "gnulinux",  
> "status" : "green",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 2,  
> "number\_of\_data\_nodes" : 2,  
> "active\_primary\_shards" : 0,  
> "active\_shards" : 0,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 0,  
> "unassigned\_shards" : 0  
> }
> 
> # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> 
> {  
> "cluster\_name" : "gnulinux",  
> "status" : "green",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 2,  
> "number\_of\_data\_nodes" : 2,  
> "active\_primary\_shards" : 0,  
> "active\_shards" : 0,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 0,  
> "unassigned\_shards" : 0  
> }
> 
> 1. Added some data and restarted the nodes, the cluster status is red and  
> log file is giving the same error.
> 
> # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> 
> {  
> "cluster\_name" : "gnulinux",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 2,  
> "number\_of\_data\_nodes" : 2,  
> "active\_primary\_shards" : 0,  
> "active\_shards" : 0,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 5,  
> "unassigned\_shards" : 5  
> }
> 
> # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> 
> {  
> "cluster\_name" : "gnulinux",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 2,  
> "number\_of\_data\_nodes" : 2,  
> "active\_primary\_shards" : 0,  
> "active\_shards" : 0,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 5,  
> "unassigned\_shards" : 5
> 
> [2011-10-22 01:18:46,703][WARN][cluster.action.shard] [Ubu1] sending  
> failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> recovery (post api), should exists, but doesn't]]]  
> [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1] sending  
> failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> recovery (post api), should exists, but doesn't]]]  
> [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1] sending  
> failed shard for [twitter][0], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[twitter][0] shard allocated for local  
> recovery (post api), should exists, but doesn't]]]  
> [2011-10-22 01:18:46,725][WARN][indices.cluster] [Ubu1]  
> [twitter][1] failed to start shard  
> org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> [twitter][1] shard allocated for local recovery (post api), should exists,  
> but doesn't
> 
> ```
> at
> 
> ```
> 
> org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> at  
> org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> at  
> java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> at  
> java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> at java.lang.Thread.run(Thread.java:662)  
> [2011-10-22 01:18:46,736][WARN][indices.cluster] [Ubu1]  
> [twitter][4] failed to start shard  
> org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> [twitter][4] shard allocated for local recovery (post api), should exists,  
> but doesn't
> 
> ```
> at
> 
> ```
> 
> org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> at  
> org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> at  
> java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> at  
> java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> at java.lang.Thread.run(Thread.java:662)  
> [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1] sending  
> failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> recovery (post api), should exists, but doesn't]]]  
> [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1] sending  
> failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> recovery (post api), should exists, but doesn't]]]
> 
> Thanks  
> Viji
> 
> On Fri, Oct 21, 2011 at 11:57 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Are you sure the two nodes find each other? The configuration you have  
> > configure jgroups for discovery, which was removed in version 0.6 ....
> > 
> > On Fri, Oct 21, 2011 at 11:53 AM, gnulinux [vijivijayakumar@gmail.com](mailto:vijivijayakumar@gmail.com)wrote:
> > 
> > > Hi
> > > 
> > > I am evaluating Elasticsearch (0.17.8) for a spatial search platform.  
> > > I was able to setup a two-node cluster and everything was working  
> > > fine. But after rebooting both the nodes, I am getting the following  
> > > error on both.
> > > 
> > > [2011-10-19 06:02:45,243][WARN][indices.cluster] [linux  
> > > Ubu2] [books][1] failed to start shard  
> > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > [books][1] shard allocated for local recovery (post api), should  
> > > exists, but doesn't  
> > > at
> > > 
> > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> > > 99)  
> > > at org.elasticsearch.index.gateway.IndexShardGatewayService  
> > > $1.run(IndexShardGatewayService.java:179)  
> > > at
> > > 
> > > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
> > > 1110)  
> > > at java.util.concurrent.ThreadPoolExecutor  
> > > $Worker.run(ThreadPoolExecutor.java:603)  
> > > at java.lang.Thread.run(Thread.java:679)
> > > 
> > > Config Files:
> > > 
> > > Node01 (Master)
> > > 
> > > cluster:  
> > > name: gnulinux
> > > 
> > > node.name: "linux Ubu2"  
> > > node.master: true  
> > > node.data: true  
> > > node.rack: rack01
> > > 
> > > network:  
> > > bindHost: 192.168.2.10  
> > > publishHost: 192.168.2.10
> > > 
> > > index.engine.robin.refreshInterval: -1  
> > > index.gateway.snapshot\_interval: -1  
> > > index.gateway.type: local  
> > > index.number\_of\_shards: 5  
> > > index.number\_of\_replicas: 1
> > > 
> > > gateway.recover\_after\_nodes: 2  
> > > gateway.recover\_after\_time: 5m  
> > > gateway.expected\_nodes: 2  
> > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > indices.recovery.concurrent\_streams: 5
> > > 
> > > index:  
> > > store:  
> > > fs:  
> > > memory:  
> > > enabled: true  
> > > discovery:  
> > > jgroups:  
> > > config: tcp  
> > > bind\_port: 9700  
> > > bind\_address: 192.168.2.10  
> > > tcpping:  
> > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]
> > > 
> > > Node02:
> > > 
> > > cluster:  
> > > name: gnulinux
> > > 
> > > node.name: "linux Ubu1"  
> > > node.master: false  
> > > node.data: true  
> > > node.rack: rack01
> > > 
> > > network:  
> > > bindHost: 192.168.2.11  
> > > publishHost: 192.168.2.11
> > > 
> > > index.engine.robin.refreshInterval: -1  
> > > index.gateway.snapshot\_interval: -1  
> > > index.gateway.type: local  
> > > index.number\_of\_shards: 5  
> > > index.number\_of\_replicas: 1
> > > 
> > > gateway.recover\_after\_nodes: 2  
> > > gateway.recover\_after\_time: 5m  
> > > gateway.expected\_nodes: 2  
> > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > indices.recovery.concurrent\_streams: 5
> > > 
> > > index:  
> > > store:  
> > > fs:  
> > > memory:  
> > > enabled: true  
> > > discovery:  
> > > jgroups:  
> > > config: tcp  
> > > bind\_port: 9700  
> > > bind\_address: 192.168.2.11  
> > > tcpping:  
> > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

---

<div class="post-metadata">

**Author:** ![Viji\_Nair](https://avatars.discourse-cdn.com/v4/letter/v/7ab992/32.png) [@Viji\_Nair](https://discuss.elastic.co/u/Viji_Nair)\
**Post date:** [October 22, 2011, 3:26am UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/5 "2011-10-22T03:26:47Z")

</div>

Yes, I am sure. Deleted the old index, reconfigured freshly as explained,  
added data, tested , and restarted the nodes. No deletion in-between.

On Sat, Oct 22, 2011 at 5:23 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Are you sure that you don't delete the index content between restarts?
> 
> On Fri, Oct 21, 2011 at 10:09 PM, Viji Nair [viji@linux.com](mailto:viji@linux.com) wrote:
> 
> > Hi,
> > 
> > Yes, I missed it. But I don't know how, the cluster API was reporting  
> > "green" and showing the total number of nodes as "2"
> > 
> > Now, I changed to zen discovery and deleted all the existing indexes. The  
> > node discovery happens properly, but after adding index (this time followed  
> > the twitter example) and putting some data, a subsequent reboot is giving  
> > the same issue. Please find the steps I have followed.
> > 
> > 1. Installed Java
> > 
> > # java -version
> > 
> > java version "1.6.0\_29"  
> > Java(TM) SE Runtime Environment (build 1.6.0\_29-b11)  
> > Java HotSpot(TM) 64-Bit Server VM (build 20.4-b02, mixed mode)
> > 
> > 1. This is an ubuntu 64 bit machine (11.04)
> > 
> > # uname -a
> > 
> > Linux ubu-ser 2.6.38-8-generic #42-Ubuntu SMP Mon Apr 11 03:31:24 UTC 2011  
> > x86\_64 x86\_64 x86\_64 GNU/Linux
> > 
> > 1. Installed Elastic Search 0.17.9 and Service Wrapper.
> > 
> > 2. Configured a two node cluster
> > 
> > - 
> > 
> > Node01 Configuration\*
> > 
> > # cat /root/elasticsearch-0.17.8/config/elasticsearch.yml
> > 
> > cluster:  
> > name: gnulinux
> > 
> > node.name: "Ubu2"
> > 
> > node.master: true  
> > node.data: true  
> > node.rack: rack01
> > 
> > network:  
> > bindHost: 192.168.2.10  
> > publishHost: 192.168.2.10
> > 
> > index.engine.robin.refreshInterval: -1  
> > index.gateway.snapshot\_interval: -1  
> > index.gateway.type: local  
> > index.number\_of\_shards: 5  
> > index.number\_of\_replicas: 1
> > 
> > gateway.recover\_after\_nodes: 2  
> > gateway.recover\_after\_time: 5m  
> > gateway.expected\_nodes: 2  
> > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > indices.recovery.concurrent\_streams: 5
> > 
> > index:  
> > store:  
> > fs:  
> > memory:  
> > enabled: true
> > 
> > discovery:  
> > zen:  
> > ping\_timeout: 30s  
> > ping:  
> > multicast:  
> > enabled: false  
> > unicast:  
> > enabled: true  
> > hosts: 192.168.2.10, 192.168.2.11  
> > fd:  
> > ping\_retries: 10  
> > ping\_interval: 5s  
> > ping\_timeout: 30s
> > 
> > _Node02 Configuration_
> > 
> > #cat /root/elasticsearch-0.17.8/config/elasticsearch.yml  
> > cluster:  
> > name: gnulinux
> > 
> > node.name: "Ubu1"
> > 
> > node.master: false  
> > node.data: true  
> > node.rack: rack01
> > 
> > network:  
> > bindHost: 192.168.2.11  
> > publishHost: 192.168.2.11
> > 
> > index.engine.robin.refreshInterval: -1  
> > index.gateway.snapshot\_interval: -1  
> > index.gateway.type: local  
> > index.number\_of\_shards: 5  
> > index.number\_of\_replicas: 1
> > 
> > gateway.recover\_after\_nodes: 2  
> > gateway.recover\_after\_time: 5m  
> > gateway.expected\_nodes: 2  
> > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > indices.recovery.concurrent\_streams: 5
> > 
> > index:  
> > store:  
> > fs:  
> > memory:  
> > enabled: true
> > 
> > discovery:  
> > zen:  
> > ping\_timeout: 30s  
> > ping:  
> > multicast:  
> > enabled: false  
> > unicast:  
> > enabled: true  
> > hosts: 192.168.2.10, 192.168.2.11  
> > fd:  
> > ping\_retries: 10  
> > ping\_interval: 5s  
> > ping\_timeout: 30s
> > 
> > 1. Started both the nodes and checked the status, both were up. Verified  
> > the log file as well. Everything was fine till this step.
> > 
> > # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> > 
> > {  
> > "cluster\_name" : "gnulinux",  
> > "status" : "green",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 2,  
> > "number\_of\_data\_nodes" : 2,  
> > "active\_primary\_shards" : 0,  
> > "active\_shards" : 0,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 0,  
> > "unassigned\_shards" : 0  
> > }
> > 
> > # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> > 
> > {  
> > "cluster\_name" : "gnulinux",  
> > "status" : "green",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 2,  
> > "number\_of\_data\_nodes" : 2,  
> > "active\_primary\_shards" : 0,  
> > "active\_shards" : 0,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 0,  
> > "unassigned\_shards" : 0  
> > }
> > 
> > 1. Added some data and restarted the nodes, the cluster status is red and  
> > log file is giving the same error.
> > 
> > # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> > 
> > {  
> > "cluster\_name" : "gnulinux",  
> > "status" : "red",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 2,  
> > "number\_of\_data\_nodes" : 2,  
> > "active\_primary\_shards" : 0,  
> > "active\_shards" : 0,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 5,  
> > "unassigned\_shards" : 5  
> > }
> > 
> > # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> > 
> > {  
> > "cluster\_name" : "gnulinux",  
> > "status" : "red",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 2,  
> > "number\_of\_data\_nodes" : 2,  
> > "active\_primary\_shards" : 0,  
> > "active\_shards" : 0,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 5,  
> > "unassigned\_shards" : 5
> > 
> > [2011-10-22 01:18:46,703][WARN][cluster.action.shard] [Ubu1] sending  
> > failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > s[INITIALIZING], reason [Failed to start shard, message  
> > [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> > recovery (post api), should exists, but doesn't]]]  
> > [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1] sending  
> > failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > s[INITIALIZING], reason [Failed to start shard, message  
> > [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> > recovery (post api), should exists, but doesn't]]]  
> > [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1] sending  
> > failed shard for [twitter][0], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > s[INITIALIZING], reason [Failed to start shard, message  
> > [IndexShardGatewayRecoveryException[[twitter][0] shard allocated for local  
> > recovery (post api), should exists, but doesn't]]]  
> > [2011-10-22 01:18:46,725][WARN][indices.cluster] [Ubu1]  
> > [twitter][1] failed to start shard  
> > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > [twitter][1] shard allocated for local recovery (post api), should exists,  
> > but doesn't
> > 
> > ```
> > at
> > 
> > ```
> > 
> > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> > at  
> > org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> > at  
> > java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> > at  
> > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> > at java.lang.Thread.run(Thread.java:662)  
> > [2011-10-22 01:18:46,736][WARN][indices.cluster] [Ubu1]  
> > [twitter][4] failed to start shard  
> > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > [twitter][4] shard allocated for local recovery (post api), should exists,  
> > but doesn't
> > 
> > ```
> > at
> > 
> > ```
> > 
> > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> > at  
> > org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> > at  
> > java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> > at  
> > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> > at java.lang.Thread.run(Thread.java:662)  
> > [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1] sending  
> > failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > s[INITIALIZING], reason [Failed to start shard, message  
> > [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> > recovery (post api), should exists, but doesn't]]]  
> > [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1] sending  
> > failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > s[INITIALIZING], reason [Failed to start shard, message  
> > [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> > recovery (post api), should exists, but doesn't]]]
> > 
> > Thanks  
> > Viji
> > 
> > On Fri, Oct 21, 2011 at 11:57 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > > Are you sure the two nodes find each other? The configuration you have  
> > > configure jgroups for discovery, which was removed in version 0.6 ....
> > > 
> > > On Fri, Oct 21, 2011 at 11:53 AM, gnulinux [vijivijayakumar@gmail.com](mailto:vijivijayakumar@gmail.com)wrote:
> > > 
> > > > Hi
> > > > 
> > > > I am evaluating Elasticsearch (0.17.8) for a spatial search platform.  
> > > > I was able to setup a two-node cluster and everything was working  
> > > > fine. But after rebooting both the nodes, I am getting the following  
> > > > error on both.
> > > > 
> > > > [2011-10-19 06:02:45,243][WARN][indices.cluster] [linux  
> > > > Ubu2] [books][1] failed to start shard  
> > > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > > [books][1] shard allocated for local recovery (post api), should  
> > > > exists, but doesn't  
> > > > at
> > > > 
> > > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> > > > 99)  
> > > > at org.elasticsearch.index.gateway.IndexShardGatewayService  
> > > > $1.run(IndexShardGatewayService.java:179)  
> > > > at
> > > > 
> > > > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
> > > > 1110)  
> > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > $Worker.run(ThreadPoolExecutor.java:603)  
> > > > at java.lang.Thread.run(Thread.java:679)
> > > > 
> > > > Config Files:
> > > > 
> > > > Node01 (Master)
> > > > 
> > > > cluster:  
> > > > name: gnulinux
> > > > 
> > > > node.name: "linux Ubu2"  
> > > > node.master: true  
> > > > node.data: true  
> > > > node.rack: rack01
> > > > 
> > > > network:  
> > > > bindHost: 192.168.2.10  
> > > > publishHost: 192.168.2.10
> > > > 
> > > > index.engine.robin.refreshInterval: -1  
> > > > index.gateway.snapshot\_interval: -1  
> > > > index.gateway.type: local  
> > > > index.number\_of\_shards: 5  
> > > > index.number\_of\_replicas: 1
> > > > 
> > > > gateway.recover\_after\_nodes: 2  
> > > > gateway.recover\_after\_time: 5m  
> > > > gateway.expected\_nodes: 2  
> > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > indices.recovery.concurrent\_streams: 5
> > > > 
> > > > index:  
> > > > store:  
> > > > fs:  
> > > > memory:  
> > > > enabled: true  
> > > > discovery:  
> > > > jgroups:  
> > > > config: tcp  
> > > > bind\_port: 9700  
> > > > bind\_address: 192.168.2.10  
> > > > tcpping:  
> > > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]
> > > > 
> > > > Node02:
> > > > 
> > > > cluster:  
> > > > name: gnulinux
> > > > 
> > > > node.name: "linux Ubu1"  
> > > > node.master: false  
> > > > node.data: true  
> > > > node.rack: rack01
> > > > 
> > > > network:  
> > > > bindHost: 192.168.2.11  
> > > > publishHost: 192.168.2.11
> > > > 
> > > > index.engine.robin.refreshInterval: -1  
> > > > index.gateway.snapshot\_interval: -1  
> > > > index.gateway.type: local  
> > > > index.number\_of\_shards: 5  
> > > > index.number\_of\_replicas: 1
> > > > 
> > > > gateway.recover\_after\_nodes: 2  
> > > > gateway.recover\_after\_time: 5m  
> > > > gateway.expected\_nodes: 2  
> > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > indices.recovery.concurrent\_streams: 5
> > > > 
> > > > index:  
> > > > store:  
> > > > fs:  
> > > > memory:  
> > > > enabled: true  
> > > > discovery:  
> > > > jgroups:  
> > > > config: tcp  
> > > > bind\_port: 9700  
> > > > bind\_address: 192.168.2.11  
> > > > tcpping:  
> > > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 23, 2011, 12:15am UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/6 "2011-10-23T00:15:34Z")

</div>

The error comes when a shard is allocated on a node where it expects the  
index to exists, but its not there. Maybe you can somehow try and recreate  
it locally (you can easily start 2 nodes locally on your machine), see if it  
happens then. If so, gist the steps you use and I can check it.

On Sat, Oct 22, 2011 at 5:26 AM, Viji Nair [viji@linux.com](mailto:viji@linux.com) wrote:

> Yes, I am sure. Deleted the old index, reconfigured freshly as explained,  
> added data, tested , and restarted the nodes. No deletion in-between.
> 
> On Sat, Oct 22, 2011 at 5:23 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Are you sure that you don't delete the index content between restarts?
> > 
> > On Fri, Oct 21, 2011 at 10:09 PM, Viji Nair [viji@linux.com](mailto:viji@linux.com) wrote:
> > 
> > > Hi,
> > > 
> > > Yes, I missed it. But I don't know how, the cluster API was reporting  
> > > "green" and showing the total number of nodes as "2"
> > > 
> > > Now, I changed to zen discovery and deleted all the existing indexes. The  
> > > node discovery happens properly, but after adding index (this time followed  
> > > the twitter example) and putting some data, a subsequent reboot is giving  
> > > the same issue. Please find the steps I have followed.
> > > 
> > > 1. Installed Java
> > > 
> > > # java -version
> > > 
> > > java version "1.6.0\_29"  
> > > Java(TM) SE Runtime Environment (build 1.6.0\_29-b11)  
> > > Java HotSpot(TM) 64-Bit Server VM (build 20.4-b02, mixed mode)
> > > 
> > > 1. This is an ubuntu 64 bit machine (11.04)
> > > 
> > > # uname -a
> > > 
> > > Linux ubu-ser 2.6.38-8-generic #42-Ubuntu SMP Mon Apr 11 03:31:24 UTC  
> > > 2011 x86\_64 x86\_64 x86\_64 GNU/Linux
> > > 
> > > 1. Installed Elastic Search 0.17.9 and Service Wrapper.
> > > 
> > > 2. Configured a two node cluster
> > > 
> > > - 
> > > 
> > > Node01 Configuration\*
> > > 
> > > # cat /root/elasticsearch-0.17.8/config/elasticsearch.yml
> > > 
> > > cluster:  
> > > name: gnulinux
> > > 
> > > node.name: "Ubu2"
> > > 
> > > node.master: true  
> > > node.data: true  
> > > node.rack: rack01
> > > 
> > > network:  
> > > bindHost: 192.168.2.10  
> > > publishHost: 192.168.2.10
> > > 
> > > index.engine.robin.refreshInterval: -1  
> > > index.gateway.snapshot\_interval: -1  
> > > index.gateway.type: local  
> > > index.number\_of\_shards: 5  
> > > index.number\_of\_replicas: 1
> > > 
> > > gateway.recover\_after\_nodes: 2  
> > > gateway.recover\_after\_time: 5m  
> > > gateway.expected\_nodes: 2  
> > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > indices.recovery.concurrent\_streams: 5
> > > 
> > > index:  
> > > store:  
> > > fs:  
> > > memory:  
> > > enabled: true
> > > 
> > > discovery:  
> > > zen:  
> > > ping\_timeout: 30s  
> > > ping:  
> > > multicast:  
> > > enabled: false  
> > > unicast:  
> > > enabled: true  
> > > hosts: 192.168.2.10, 192.168.2.11  
> > > fd:  
> > > ping\_retries: 10  
> > > ping\_interval: 5s  
> > > ping\_timeout: 30s
> > > 
> > > _Node02 Configuration_
> > > 
> > > #cat /root/elasticsearch-0.17.8/config/elasticsearch.yml  
> > > cluster:  
> > > name: gnulinux
> > > 
> > > node.name: "Ubu1"
> > > 
> > > node.master: false  
> > > node.data: true  
> > > node.rack: rack01
> > > 
> > > network:  
> > > bindHost: 192.168.2.11  
> > > publishHost: 192.168.2.11
> > > 
> > > index.engine.robin.refreshInterval: -1  
> > > index.gateway.snapshot\_interval: -1  
> > > index.gateway.type: local  
> > > index.number\_of\_shards: 5  
> > > index.number\_of\_replicas: 1
> > > 
> > > gateway.recover\_after\_nodes: 2  
> > > gateway.recover\_after\_time: 5m  
> > > gateway.expected\_nodes: 2  
> > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > indices.recovery.concurrent\_streams: 5
> > > 
> > > index:  
> > > store:  
> > > fs:  
> > > memory:  
> > > enabled: true
> > > 
> > > discovery:  
> > > zen:  
> > > ping\_timeout: 30s  
> > > ping:  
> > > multicast:  
> > > enabled: false  
> > > unicast:  
> > > enabled: true  
> > > hosts: 192.168.2.10, 192.168.2.11  
> > > fd:  
> > > ping\_retries: 10  
> > > ping\_interval: 5s  
> > > ping\_timeout: 30s
> > > 
> > > 1. Started both the nodes and checked the status, both were up. Verified  
> > > the log file as well. Everything was fine till this step.
> > > 
> > > # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> > > 
> > > {  
> > > "cluster\_name" : "gnulinux",  
> > > "status" : "green",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 2,  
> > > "number\_of\_data\_nodes" : 2,  
> > > "active\_primary\_shards" : 0,  
> > > "active\_shards" : 0,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 0,  
> > > "unassigned\_shards" : 0  
> > > }
> > > 
> > > # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> > > 
> > > {  
> > > "cluster\_name" : "gnulinux",  
> > > "status" : "green",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 2,  
> > > "number\_of\_data\_nodes" : 2,  
> > > "active\_primary\_shards" : 0,  
> > > "active\_shards" : 0,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 0,  
> > > "unassigned\_shards" : 0  
> > > }
> > > 
> > > 1. Added some data and restarted the nodes, the cluster status is red and  
> > > log file is giving the same error.
> > > 
> > > # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> > > 
> > > {  
> > > "cluster\_name" : "gnulinux",  
> > > "status" : "red",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 2,  
> > > "number\_of\_data\_nodes" : 2,  
> > > "active\_primary\_shards" : 0,  
> > > "active\_shards" : 0,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 5,  
> > > "unassigned\_shards" : 5  
> > > }
> > > 
> > > # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> > > 
> > > {  
> > > "cluster\_name" : "gnulinux",  
> > > "status" : "red",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 2,  
> > > "number\_of\_data\_nodes" : 2,  
> > > "active\_primary\_shards" : 0,  
> > > "active\_shards" : 0,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 5,  
> > > "unassigned\_shards" : 5
> > > 
> > > [2011-10-22 01:18:46,703][WARN][cluster.action.shard] [Ubu1]  
> > > sending failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > s[INITIALIZING], reason [Failed to start shard, message  
> > > [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> > > recovery (post api), should exists, but doesn't]]]  
> > > [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1]  
> > > sending failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > s[INITIALIZING], reason [Failed to start shard, message  
> > > [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> > > recovery (post api), should exists, but doesn't]]]  
> > > [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1]  
> > > sending failed shard for [twitter][0], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > s[INITIALIZING], reason [Failed to start shard, message  
> > > [IndexShardGatewayRecoveryException[[twitter][0] shard allocated for local  
> > > recovery (post api), should exists, but doesn't]]]  
> > > [2011-10-22 01:18:46,725][WARN][indices.cluster] [Ubu1]  
> > > [twitter][1] failed to start shard  
> > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > [twitter][1] shard allocated for local recovery (post api), should exists,  
> > > but doesn't
> > > 
> > > ```
> > > at
> > > 
> > > ```
> > > 
> > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> > > at  
> > > org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> > > at  
> > > java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> > > at  
> > > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> > > at java.lang.Thread.run(Thread.java:662)  
> > > [2011-10-22 01:18:46,736][WARN][indices.cluster] [Ubu1]  
> > > [twitter][4] failed to start shard  
> > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > [twitter][4] shard allocated for local recovery (post api), should exists,  
> > > but doesn't
> > > 
> > > ```
> > > at
> > > 
> > > ```
> > > 
> > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> > > at  
> > > org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> > > at  
> > > java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> > > at  
> > > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> > > at java.lang.Thread.run(Thread.java:662)  
> > > [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1]  
> > > sending failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > s[INITIALIZING], reason [Failed to start shard, message  
> > > [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> > > recovery (post api), should exists, but doesn't]]]  
> > > [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1]  
> > > sending failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > s[INITIALIZING], reason [Failed to start shard, message  
> > > [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> > > recovery (post api), should exists, but doesn't]]]
> > > 
> > > Thanks  
> > > Viji
> > > 
> > > On Fri, Oct 21, 2011 at 11:57 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > 
> > > > Are you sure the two nodes find each other? The configuration you have  
> > > > configure jgroups for discovery, which was removed in version 0.6 ....
> > > > 
> > > > On Fri, Oct 21, 2011 at 11:53 AM, gnulinux [vijivijayakumar@gmail.com](mailto:vijivijayakumar@gmail.com)wrote:
> > > > 
> > > > > Hi
> > > > > 
> > > > > I am evaluating Elasticsearch (0.17.8) for a spatial search platform.  
> > > > > I was able to setup a two-node cluster and everything was working  
> > > > > fine. But after rebooting both the nodes, I am getting the following  
> > > > > error on both.
> > > > > 
> > > > > [2011-10-19 06:02:45,243][WARN][indices.cluster] [linux  
> > > > > Ubu2] [books][1] failed to start shard  
> > > > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > > > [books][1] shard allocated for local recovery (post api), should  
> > > > > exists, but doesn't  
> > > > > at
> > > > > 
> > > > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> > > > > 99)  
> > > > > at org.elasticsearch.index.gateway.IndexShardGatewayService  
> > > > > $1.run(IndexShardGatewayService.java:179)  
> > > > > at
> > > > > 
> > > > > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
> > > > > 1110)  
> > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > $Worker.run(ThreadPoolExecutor.java:603)  
> > > > > at java.lang.Thread.run(Thread.java:679)
> > > > > 
> > > > > Config Files:
> > > > > 
> > > > > Node01 (Master)
> > > > > 
> > > > > cluster:  
> > > > > name: gnulinux
> > > > > 
> > > > > node.name: "linux Ubu2"  
> > > > > node.master: true  
> > > > > node.data: true  
> > > > > node.rack: rack01
> > > > > 
> > > > > network:  
> > > > > bindHost: 192.168.2.10  
> > > > > publishHost: 192.168.2.10
> > > > > 
> > > > > index.engine.robin.refreshInterval: -1  
> > > > > index.gateway.snapshot\_interval: -1  
> > > > > index.gateway.type: local  
> > > > > index.number\_of\_shards: 5  
> > > > > index.number\_of\_replicas: 1
> > > > > 
> > > > > gateway.recover\_after\_nodes: 2  
> > > > > gateway.recover\_after\_time: 5m  
> > > > > gateway.expected\_nodes: 2  
> > > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > > indices.recovery.concurrent\_streams: 5
> > > > > 
> > > > > index:  
> > > > > store:  
> > > > > fs:  
> > > > > memory:  
> > > > > enabled: true  
> > > > > discovery:  
> > > > > jgroups:  
> > > > > config: tcp  
> > > > > bind\_port: 9700  
> > > > > bind\_address: 192.168.2.10  
> > > > > tcpping:  
> > > > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]
> > > > > 
> > > > > Node02:
> > > > > 
> > > > > cluster:  
> > > > > name: gnulinux
> > > > > 
> > > > > node.name: "linux Ubu1"  
> > > > > node.master: false  
> > > > > node.data: true  
> > > > > node.rack: rack01
> > > > > 
> > > > > network:  
> > > > > bindHost: 192.168.2.11  
> > > > > publishHost: 192.168.2.11
> > > > > 
> > > > > index.engine.robin.refreshInterval: -1  
> > > > > index.gateway.snapshot\_interval: -1  
> > > > > index.gateway.type: local  
> > > > > index.number\_of\_shards: 5  
> > > > > index.number\_of\_replicas: 1
> > > > > 
> > > > > gateway.recover\_after\_nodes: 2  
> > > > > gateway.recover\_after\_time: 5m  
> > > > > gateway.expected\_nodes: 2  
> > > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > > indices.recovery.concurrent\_streams: 5
> > > > > 
> > > > > index:  
> > > > > store:  
> > > > > fs:  
> > > > > memory:  
> > > > > enabled: true  
> > > > > discovery:  
> > > > > jgroups:  
> > > > > config: tcp  
> > > > > bind\_port: 9700  
> > > > > bind\_address: 192.168.2.11  
> > > > > tcpping:  
> > > > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

---

<div class="post-metadata">

**Author:** ![Viji\_Nair](https://avatars.discourse-cdn.com/v4/letter/v/7ab992/32.png) [@Viji\_Nair](https://discuss.elastic.co/u/Viji_Nair)\
**Post date:** [October 27, 2011, 11:09am UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/7 "2011-10-27T11:09:15Z")

</div>

I am not sure what exactly went wrong. I upgraded to the latest version of  
ES today (0.18.1) and everything started working fine, even after multiple  
stop/start of the instances my cluster seems stable in all aspects.

1. Downloaded the latest ES and Service Wrapper binaries.
2. Copied the config file form old setup
3. Started the cluster and added some data
4. Restarted the instances
5. Cluster is stable and "green"

Cheers,  
Viji

On Sun, Oct 23, 2011 at 5:45 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> The error comes when a shard is allocated on a node where it expects the  
> index to exists, but its not there. Maybe you can somehow try and recreate  
> it locally (you can easily start 2 nodes locally on your machine), see if it  
> happens then. If so, gist the steps you use and I can check it.
> 
> On Sat, Oct 22, 2011 at 5:26 AM, Viji Nair [viji@linux.com](mailto:viji@linux.com) wrote:
> 
> > Yes, I am sure. Deleted the old index, reconfigured freshly as explained,  
> > added data, tested , and restarted the nodes. No deletion in-between.
> > 
> > On Sat, Oct 22, 2011 at 5:23 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > > Are you sure that you don't delete the index content between restarts?
> > > 
> > > On Fri, Oct 21, 2011 at 10:09 PM, Viji Nair [viji@linux.com](mailto:viji@linux.com) wrote:
> > > 
> > > > Hi,
> > > > 
> > > > Yes, I missed it. But I don't know how, the cluster API was reporting  
> > > > "green" and showing the total number of nodes as "2"
> > > > 
> > > > Now, I changed to zen discovery and deleted all the existing indexes.  
> > > > The node discovery happens properly, but after adding index (this time  
> > > > followed the twitter example) and putting some data, a subsequent reboot is  
> > > > giving the same issue. Please find the steps I have followed.
> > > > 
> > > > 1. Installed Java
> > > > 
> > > > # java -version
> > > > 
> > > > java version "1.6.0\_29"  
> > > > Java(TM) SE Runtime Environment (build 1.6.0\_29-b11)  
> > > > Java HotSpot(TM) 64-Bit Server VM (build 20.4-b02, mixed mode)
> > > > 
> > > > 1. This is an ubuntu 64 bit machine (11.04)
> > > > 
> > > > # uname -a
> > > > 
> > > > Linux ubu-ser 2.6.38-8-generic #42-Ubuntu SMP Mon Apr 11 03:31:24 UTC  
> > > > 2011 x86\_64 x86\_64 x86\_64 GNU/Linux
> > > > 
> > > > 1. Installed Elastic Search 0.17.9 and Service Wrapper.
> > > > 
> > > > 2. Configured a two node cluster
> > > > 
> > > > - 
> > > > 
> > > > Node01 Configuration\*
> > > > 
> > > > # cat /root/elasticsearch-0.17.8/config/elasticsearch.yml
> > > > 
> > > > cluster:  
> > > > name: gnulinux
> > > > 
> > > > node.name: "Ubu2"
> > > > 
> > > > node.master: true  
> > > > node.data: true  
> > > > node.rack: rack01
> > > > 
> > > > network:  
> > > > bindHost: 192.168.2.10  
> > > > publishHost: 192.168.2.10
> > > > 
> > > > index.engine.robin.refreshInterval: -1  
> > > > index.gateway.snapshot\_interval: -1  
> > > > index.gateway.type: local  
> > > > index.number\_of\_shards: 5  
> > > > index.number\_of\_replicas: 1
> > > > 
> > > > gateway.recover\_after\_nodes: 2  
> > > > gateway.recover\_after\_time: 5m  
> > > > gateway.expected\_nodes: 2  
> > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > indices.recovery.concurrent\_streams: 5
> > > > 
> > > > index:  
> > > > store:  
> > > > fs:  
> > > > memory:  
> > > > enabled: true
> > > > 
> > > > discovery:  
> > > > zen:  
> > > > ping\_timeout: 30s  
> > > > ping:  
> > > > multicast:  
> > > > enabled: false  
> > > > unicast:  
> > > > enabled: true  
> > > > hosts: 192.168.2.10, 192.168.2.11  
> > > > fd:  
> > > > ping\_retries: 10  
> > > > ping\_interval: 5s  
> > > > ping\_timeout: 30s
> > > > 
> > > > _Node02 Configuration_
> > > > 
> > > > #cat /root/elasticsearch-0.17.8/config/elasticsearch.yml  
> > > > cluster:  
> > > > name: gnulinux
> > > > 
> > > > node.name: "Ubu1"
> > > > 
> > > > node.master: false  
> > > > node.data: true  
> > > > node.rack: rack01
> > > > 
> > > > network:  
> > > > bindHost: 192.168.2.11  
> > > > publishHost: 192.168.2.11
> > > > 
> > > > index.engine.robin.refreshInterval: -1  
> > > > index.gateway.snapshot\_interval: -1  
> > > > index.gateway.type: local  
> > > > index.number\_of\_shards: 5  
> > > > index.number\_of\_replicas: 1
> > > > 
> > > > gateway.recover\_after\_nodes: 2  
> > > > gateway.recover\_after\_time: 5m  
> > > > gateway.expected\_nodes: 2  
> > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > indices.recovery.concurrent\_streams: 5
> > > > 
> > > > index:  
> > > > store:  
> > > > fs:  
> > > > memory:  
> > > > enabled: true
> > > > 
> > > > discovery:  
> > > > zen:  
> > > > ping\_timeout: 30s  
> > > > ping:  
> > > > multicast:  
> > > > enabled: false  
> > > > unicast:  
> > > > enabled: true  
> > > > hosts: 192.168.2.10, 192.168.2.11  
> > > > fd:  
> > > > ping\_retries: 10  
> > > > ping\_interval: 5s  
> > > > ping\_timeout: 30s
> > > > 
> > > > 1. Started both the nodes and checked the status, both were up. Verified  
> > > > the log file as well. Everything was fine till this step.
> > > > 
> > > > # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> > > > 
> > > > {  
> > > > "cluster\_name" : "gnulinux",  
> > > > "status" : "green",  
> > > > "timed\_out" : false,  
> > > > "number\_of\_nodes" : 2,  
> > > > "number\_of\_data\_nodes" : 2,  
> > > > "active\_primary\_shards" : 0,  
> > > > "active\_shards" : 0,  
> > > > "relocating\_shards" : 0,  
> > > > "initializing\_shards" : 0,  
> > > > "unassigned\_shards" : 0  
> > > > }
> > > > 
> > > > # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> > > > 
> > > > {  
> > > > "cluster\_name" : "gnulinux",  
> > > > "status" : "green",  
> > > > "timed\_out" : false,  
> > > > "number\_of\_nodes" : 2,  
> > > > "number\_of\_data\_nodes" : 2,  
> > > > "active\_primary\_shards" : 0,  
> > > > "active\_shards" : 0,  
> > > > "relocating\_shards" : 0,  
> > > > "initializing\_shards" : 0,  
> > > > "unassigned\_shards" : 0  
> > > > }
> > > > 
> > > > 1. Added some data and restarted the nodes, the cluster status is red  
> > > > and log file is giving the same error.
> > > > 
> > > > # curl -XGET '[http://192.168.2.10:9200/\_cluster/health?pretty=true](http://192.168.2.10:9200/_cluster/health?pretty=true)'
> > > > 
> > > > {  
> > > > "cluster\_name" : "gnulinux",  
> > > > "status" : "red",  
> > > > "timed\_out" : false,  
> > > > "number\_of\_nodes" : 2,  
> > > > "number\_of\_data\_nodes" : 2,  
> > > > "active\_primary\_shards" : 0,  
> > > > "active\_shards" : 0,  
> > > > "relocating\_shards" : 0,  
> > > > "initializing\_shards" : 5,  
> > > > "unassigned\_shards" : 5  
> > > > }
> > > > 
> > > > # curl -XGET '[http://192.168.2.11:9200/\_cluster/health?pretty=true](http://192.168.2.11:9200/_cluster/health?pretty=true)'
> > > > 
> > > > {  
> > > > "cluster\_name" : "gnulinux",  
> > > > "status" : "red",  
> > > > "timed\_out" : false,  
> > > > "number\_of\_nodes" : 2,  
> > > > "number\_of\_data\_nodes" : 2,  
> > > > "active\_primary\_shards" : 0,  
> > > > "active\_shards" : 0,  
> > > > "relocating\_shards" : 0,  
> > > > "initializing\_shards" : 5,  
> > > > "unassigned\_shards" : 5
> > > > 
> > > > [2011-10-22 01:18:46,703][WARN][cluster.action.shard] [Ubu1]  
> > > > sending failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > > s[INITIALIZING], reason [Failed to start shard, message  
> > > > [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> > > > recovery (post api), should exists, but doesn't]]]  
> > > > [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1]  
> > > > sending failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > > s[INITIALIZING], reason [Failed to start shard, message  
> > > > [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> > > > recovery (post api), should exists, but doesn't]]]  
> > > > [2011-10-22 01:18:46,704][WARN][cluster.action.shard] [Ubu1]  
> > > > sending failed shard for [twitter][0], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > > s[INITIALIZING], reason [Failed to start shard, message  
> > > > [IndexShardGatewayRecoveryException[[twitter][0] shard allocated for local  
> > > > recovery (post api), should exists, but doesn't]]]  
> > > > [2011-10-22 01:18:46,725][WARN][indices.cluster] [Ubu1]  
> > > > [twitter][1] failed to start shard  
> > > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > > [twitter][1] shard allocated for local recovery (post api), should exists,  
> > > > but doesn't
> > > > 
> > > > ```
> > > > at
> > > > 
> > > > ```
> > > > 
> > > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> > > > at  
> > > > org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> > > > at  
> > > > java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > at  
> > > > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> > > > at java.lang.Thread.run(Thread.java:662)  
> > > > [2011-10-22 01:18:46,736][WARN][indices.cluster] [Ubu1]  
> > > > [twitter][4] failed to start shard  
> > > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > > [twitter][4] shard allocated for local recovery (post api), should exists,  
> > > > but doesn't
> > > > 
> > > > ```
> > > > at
> > > > 
> > > > ```
> > > > 
> > > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:99)  
> > > > at  
> > > > org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:179)  
> > > > at  
> > > > java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > at  
> > > > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)  
> > > > at java.lang.Thread.run(Thread.java:662)  
> > > > [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1]  
> > > > sending failed shard for [twitter][4], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > > s[INITIALIZING], reason [Failed to start shard, message  
> > > > [IndexShardGatewayRecoveryException[[twitter][4] shard allocated for local  
> > > > recovery (post api), should exists, but doesn't]]]  
> > > > [2011-10-22 01:18:46,737][WARN][cluster.action.shard] [Ubu1]  
> > > > sending failed shard for [twitter][1], node[fgELpN11R6m2XbIKdLHYgg], [P],  
> > > > s[INITIALIZING], reason [Failed to start shard, message  
> > > > [IndexShardGatewayRecoveryException[[twitter][1] shard allocated for local  
> > > > recovery (post api), should exists, but doesn't]]]
> > > > 
> > > > Thanks  
> > > > Viji
> > > > 
> > > > On Fri, Oct 21, 2011 at 11:57 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > > 
> > > > > Are you sure the two nodes find each other? The configuration you have  
> > > > > configure jgroups for discovery, which was removed in version 0.6 ....
> > > > > 
> > > > > On Fri, Oct 21, 2011 at 11:53 AM, gnulinux [vijivijayakumar@gmail.com](mailto:vijivijayakumar@gmail.com)wrote:
> > > > > 
> > > > > > Hi
> > > > > > 
> > > > > > I am evaluating Elasticsearch (0.17.8) for a spatial search platform.  
> > > > > > I was able to setup a two-node cluster and everything was working  
> > > > > > fine. But after rebooting both the nodes, I am getting the following  
> > > > > > error on both.
> > > > > > 
> > > > > > [2011-10-19 06:02:45,243][WARN][indices.cluster] [linux  
> > > > > > Ubu2] [books][1] failed to start shard  
> > > > > > org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException:  
> > > > > > [books][1] shard allocated for local recovery (post api), should  
> > > > > > exists, but doesn't  
> > > > > > at
> > > > > > 
> > > > > > org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:  
> > > > > > 99)  
> > > > > > at org.elasticsearch.index.gateway.IndexShardGatewayService  
> > > > > > $1.run(IndexShardGatewayService.java:179)  
> > > > > > at
> > > > > > 
> > > > > > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:  
> > > > > > 1110)  
> > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > $Worker.run(ThreadPoolExecutor.java:603)  
> > > > > > at java.lang.Thread.run(Thread.java:679)
> > > > > > 
> > > > > > Config Files:
> > > > > > 
> > > > > > Node01 (Master)
> > > > > > 
> > > > > > cluster:  
> > > > > > name: gnulinux
> > > > > > 
> > > > > > node.name: "linux Ubu2"  
> > > > > > node.master: true  
> > > > > > node.data: true  
> > > > > > node.rack: rack01
> > > > > > 
> > > > > > network:  
> > > > > > bindHost: 192.168.2.10  
> > > > > > publishHost: 192.168.2.10
> > > > > > 
> > > > > > index.engine.robin.refreshInterval: -1  
> > > > > > index.gateway.snapshot\_interval: -1  
> > > > > > index.gateway.type: local  
> > > > > > index.number\_of\_shards: 5  
> > > > > > index.number\_of\_replicas: 1
> > > > > > 
> > > > > > gateway.recover\_after\_nodes: 2  
> > > > > > gateway.recover\_after\_time: 5m  
> > > > > > gateway.expected\_nodes: 2  
> > > > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > > > indices.recovery.concurrent\_streams: 5
> > > > > > 
> > > > > > index:  
> > > > > > store:  
> > > > > > fs:  
> > > > > > memory:  
> > > > > > enabled: true  
> > > > > > discovery:  
> > > > > > jgroups:  
> > > > > > config: tcp  
> > > > > > bind\_port: 9700  
> > > > > > bind\_address: 192.168.2.10  
> > > > > > tcpping:  
> > > > > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]
> > > > > > 
> > > > > > Node02:
> > > > > > 
> > > > > > cluster:  
> > > > > > name: gnulinux
> > > > > > 
> > > > > > node.name: "linux Ubu1"  
> > > > > > node.master: false  
> > > > > > node.data: true  
> > > > > > node.rack: rack01
> > > > > > 
> > > > > > network:  
> > > > > > bindHost: 192.168.2.11  
> > > > > > publishHost: 192.168.2.11
> > > > > > 
> > > > > > index.engine.robin.refreshInterval: -1  
> > > > > > index.gateway.snapshot\_interval: -1  
> > > > > > index.gateway.type: local  
> > > > > > index.number\_of\_shards: 5  
> > > > > > index.number\_of\_replicas: 1
> > > > > > 
> > > > > > gateway.recover\_after\_nodes: 2  
> > > > > > gateway.recover\_after\_time: 5m  
> > > > > > gateway.expected\_nodes: 2  
> > > > > > cluster.routing.allocation.node\_initial\_primaries\_recoveries: 4  
> > > > > > cluster.routing.allocation.node\_concurrent\_recoveries: 2  
> > > > > > indices.recovery.concurrent\_streams: 5
> > > > > > 
> > > > > > index:  
> > > > > > store:  
> > > > > > fs:  
> > > > > > memory:  
> > > > > > enabled: true  
> > > > > > discovery:  
> > > > > > jgroups:  
> > > > > > config: tcp  
> > > > > > bind\_port: 9700  
> > > > > > bind\_address: 192.168.2.11  
> > > > > > tcpping:  
> > > > > > initial\_hosts: 192.168.2.10[9700], 192.168.2.11[9700]

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:50am UTC](https://discuss.elastic.co/t/failed-to-start-shard/5656/8 "2017-07-06T03:50:40Z")

</div>


