# Data lost after full cluster restart

**URL:** <https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906>\
**Category:** Elasticsearch\
**Created:** [July 20, 2011, 11:53am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906 "2011-07-20T11:53:02Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [July 20, 2011, 11:53am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/1 "2011-07-20T11:53:02Z")

</div>

10 nodes, replication factor 3, local storage, 10 data nodes, 10 super  
clients.  
Please let me know if you need more info.

Health status:  
{  
"cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
"status" : "red",  
"timed\_out" : false,  
"number\_of\_nodes" : 20,  
"number\_of\_data\_nodes" : 10,  
"active\_primary\_shards" : 9,  
"active\_shards" : 36,  
"relocating\_shards" : 0,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 4  
}  
State status:  
"7" : [ {  
"state" : "UNASSIGNED",  
"primary" : false,  
"node" : null,  
"relocating\_node" : null,  
"shard" : 7,  
"index" : "fs"  
}, {  
"state" : "UNASSIGNED",  
"primary" : false,  
"node" : null,  
"relocating\_node" : null,  
"shard" : 7,  
"index" : "fs"  
}, {  
"state" : "UNASSIGNED",  
"primary" : false,  
"node" : null,  
"relocating\_node" : null,  
"shard" : 7,  
"index" : "fs"  
}, {  
"state" : "UNASSIGNED",  
"primary" : true,  
"node" : null,  
"relocating\_node" : null,  
"shard" : 7,  
"index" : "fs"  
} ],

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 20, 2011, 4:47pm UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/2 "2011-07-20T16:47:49Z")

</div>

Can you gist your config? Also, if you set: gateway.local to TRACE, it will  
print all the allocation information (on the elected master) and it can  
possibly give us info as to why those shards are not allocated.

On Wed, Jul 20, 2011 at 2:53 PM, vadim [vpunski@gmail.com](mailto:vpunski@gmail.com) wrote:

> 10 nodes, replication factor 3, local storage, 10 data nodes, 10 super  
> clients.  
> Please let me know if you need more info.
> 
> Health status:  
> {  
> "cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 20,  
> "number\_of\_data\_nodes" : 10,  
> "active\_primary\_shards" : 9,  
> "active\_shards" : 36,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 0,  
> "unassigned\_shards" : 4  
> }  
> State status:  
> "7" : [ {  
> "state" : "UNASSIGNED",  
> "primary" : false,  
> "node" : null,  
> "relocating\_node" : null,  
> "shard" : 7,  
> "index" : "fs"  
> }, {  
> "state" : "UNASSIGNED",  
> "primary" : false,  
> "node" : null,  
> "relocating\_node" : null,  
> "shard" : 7,  
> "index" : "fs"  
> }, {  
> "state" : "UNASSIGNED",  
> "primary" : false,  
> "node" : null,  
> "relocating\_node" : null,  
> "shard" : 7,  
> "index" : "fs"  
> }, {  
> "state" : "UNASSIGNED",  
> "primary" : true,  
> "node" : null,  
> "relocating\_node" : null,  
> "shard" : 7,  
> "index" : "fs"  
> } ],

---

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [July 21, 2011, 8:11am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/3 "2011-07-21T08:11:49Z")

</div>

From tracing gateway.local, the only message related bad 7-th shard is  
repeated many times:  
.  
.  
[2011-07-21 10:53:53,745][INFO][cluster.service] [Delphi]  
new\_master [Delphi][apBvicCZT4mzWZrm6wZ16Q][inet[/10.11.40.238:9300]],  
reason: zen-disco-join (elected\_as\_master)  
.  
.  
[2011-07-21 10:59:57,551][DEBUG][gateway.local] [Delphi]  
[fs][7]: not allocating, number\_of\_allocated\_shards\_found [2],  
required\_number [3]

My config is:

* * *

cluster.name : MY\_CLUSTER

gateway:  
recover\_after\_nodes: 8  
recover\_after\_time: 5m  
expected\_nodes: 10

index.compound\_format : false  
index.refresh\_interval : 10s  
index.term\_index\_interval: 30

discovery.zen.ping.unicast:  
hosts: node01:9300,node02:9300,node03:9300

* * *

On Jul 20, 7:47 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Can you gist your config? Also, if you set: gateway.local to TRACE, it will  
> print all the allocation information (on the elected master) and it can  
> possibly give us info as to why those shards are not allocated.
> 
> On Wed, Jul 20, 2011 at 2:53 PM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > 10 nodes, replication factor 3, local storage, 10 data nodes, 10 super  
> > clients.  
> > Please let me know if you need more info.
> 
> > Health status:  
> > {  
> > "cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
> > "status" : "red",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 20,  
> > "number\_of\_data\_nodes" : 10,  
> > "active\_primary\_shards" : 9,  
> > "active\_shards" : 36,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 0,  
> > "unassigned\_shards" : 4  
> > }  
> > State status:  
> > "7" : [ {  
> > "state" : "UNASSIGNED",  
> > "primary" : false,  
> > "node" : null,  
> > "relocating\_node" : null,  
> > "shard" : 7,  
> > "index" : "fs"  
> > }, {  
> > "state" : "UNASSIGNED",  
> > "primary" : false,  
> > "node" : null,  
> > "relocating\_node" : null,  
> > "shard" : 7,  
> > "index" : "fs"  
> > }, {  
> > "state" : "UNASSIGNED",  
> > "primary" : false,  
> > "node" : null,  
> > "relocating\_node" : null,  
> > "shard" : 7,  
> > "index" : "fs"  
> > }, {  
> > "state" : "UNASSIGNED",  
> > "primary" : true,  
> > "node" : null,  
> > "relocating\_node" : null,  
> > "shard" : 7,  
> > "index" : "fs"  
> > } ],

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 21, 2011, 5:56pm UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/4 "2011-07-21T17:56:31Z")

</div>

It seems like the gateway only finds 2 shards (out of 4) for that shard  
group. By default, it wants to find a quorum, you can change that by  
settings gateway.local.initial\_shards to a different value than `quorum`,  
for example: 2, and then it will recover.

Another question is why there are only 2. When you created the index, was it  
create and all shards were allocated before the cluster was restarted?

On Thu, Jul 21, 2011 at 11:11 AM, vadim [vpunski@gmail.com](mailto:vpunski@gmail.com) wrote:

> From tracing gateway.local, the only message related bad 7-th shard is  
> repeated many times:  
> .  
> .  
> [2011-07-21 10:53:53,745][INFO][cluster.service] [Delphi]  
> new\_master [Delphi][apBvicCZT4mzWZrm6wZ16Q][inet[/10.11.40.238:9300]],  
> reason: zen-disco-join (elected\_as\_master)  
> .  
> .  
> [2011-07-21 10:59:57,551][DEBUG][gateway.local] [Delphi]  
> [fs][7]: not allocating, number\_of\_allocated\_shards\_found [2],  
> required\_number [3]
> 
> My config is:
> 
> * * *
> 
> cluster.name : MY\_CLUSTER
> 
> gateway:  
> recover\_after\_nodes: 8  
> recover\_after\_time: 5m  
> expected\_nodes: 10
> 
> index.compound\_format : false  
> index.refresh\_interval : 10s  
> index.term\_index\_interval: 30
> 
> discovery.zen.ping.unicast:  
> hosts: node01:9300,node02:9300,node03:9300
> 
> * * *
> 
> On Jul 20, 7:47 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Can you gist your config? Also, if you set: gateway.local to TRACE, it  
> > will  
> > print all the allocation information (on the elected master) and it can  
> > possibly give us info as to why those shards are not allocated.
> > 
> > On Wed, Jul 20, 2011 at 2:53 PM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > 
> > > 10 nodes, replication factor 3, local storage, 10 data nodes, 10 super  
> > > clients.  
> > > Please let me know if you need more info.
> > 
> > > Health status:  
> > > {  
> > > "cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
> > > "status" : "red",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 20,  
> > > "number\_of\_data\_nodes" : 10,  
> > > "active\_primary\_shards" : 9,  
> > > "active\_shards" : 36,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 0,  
> > > "unassigned\_shards" : 4  
> > > }  
> > > State status:  
> > > "7" : [ {  
> > > "state" : "UNASSIGNED",  
> > > "primary" : false,  
> > > "node" : null,  
> > > "relocating\_node" : null,  
> > > "shard" : 7,  
> > > "index" : "fs"  
> > > }, {  
> > > "state" : "UNASSIGNED",  
> > > "primary" : false,  
> > > "node" : null,  
> > > "relocating\_node" : null,  
> > > "shard" : 7,  
> > > "index" : "fs"  
> > > }, {  
> > > "state" : "UNASSIGNED",  
> > > "primary" : false,  
> > > "node" : null,  
> > > "relocating\_node" : null,  
> > > "shard" : 7,  
> > > "index" : "fs"  
> > > }, {  
> > > "state" : "UNASSIGNED",  
> > > "primary" : true,  
> > > "node" : null,  
> > > "relocating\_node" : null,  
> > > "shard" : 7,  
> > > "index" : "fs"  
> > > } ],

---

<div class="post-metadata">

**Author:** ![Michel\_Conrad](https://avatars.discourse-cdn.com/v4/letter/m/5e9695/32.png) [@Michel\_Conrad](https://discuss.elastic.co/u/Michel_Conrad)\
**Post date:** [July 22, 2011, 8:21am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/5 "2011-07-22T08:21:27Z")

</div>

Could it be that you managed to start two instances of elasticsearch  
on one of your servers?  
In that case the 2 shards missing could have been allocated to the  
second running instance,  
which would explain elasticsearch couldn't find the shards after  
restarting the cluster (starting only  
one instance of es on every server).

By looking at your data directory in the folder nodes there should  
only be a directory called 0. If there  
are multiple directories 0,1,2... you have been starting multiple  
nodes on a server and the missing shards  
may have been allocated to another node on the same server.

On Thu, Jul 21, 2011 at 7:56 PM, Shay Banon  
[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com) wrote:

> It seems like the gateway only finds 2 shards (out of 4) for that shard  
> group. By default, it wants to find a quorum, you can change that by  
> settings gateway.local.initial\_shards to a different value than `quorum`,  
> for example: 2, and then it will recover.  
> Another question is why there are only 2. When you created the index, was it  
> create and all shards were allocated before the cluster was restarted?
> 
> On Thu, Jul 21, 2011 at 11:11 AM, vadim [vpunski@gmail.com](mailto:vpunski@gmail.com) wrote:
> 
> > From tracing gateway.local, the only message related bad 7-th shard is  
> > repeated many times:  
> > .  
> > .  
> > [2011-07-21 10:53:53,745][INFO][cluster.service] [Delphi]  
> > new\_master [Delphi][apBvicCZT4mzWZrm6wZ16Q][inet[/10.11.40.238:9300]],  
> > reason: zen-disco-join (elected\_as\_master)  
> > .  
> > .  
> > [2011-07-21 10:59:57,551][DEBUG][gateway.local] [Delphi]  
> > [fs][7]: not allocating, number\_of\_allocated\_shards\_found [2],  
> > required\_number [3]
> > 
> > My config is:
> > 
> > * * *
> > 
> > cluster.name : MY\_CLUSTER
> > 
> > gateway:  
> > recover\_after\_nodes: 8  
> > recover\_after\_time: 5m  
> > expected\_nodes: 10
> > 
> > index.compound\_format : false  
> > index.refresh\_interval : 10s  
> > index.term\_index\_interval: 30
> > 
> > discovery.zen.ping.unicast:  
> > hosts: node01:9300,node02:9300,node03:9300
> > 
> > * * *
> > 
> > On Jul 20, 7:47 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Can you gist your config? Also, if you set: gateway.local to TRACE, it  
> > > will  
> > > print all the allocation information (on the elected master) and it can  
> > > possibly give us info as to why those shards are not allocated.
> > > 
> > > On Wed, Jul 20, 2011 at 2:53 PM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > > 
> > > > 10 nodes, replication factor 3, local storage, 10 data nodes, 10 super  
> > > > clients.  
> > > > Please let me know if you need more info.
> > > 
> > > > Health status:  
> > > > {  
> > > > "cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
> > > > "status" : "red",  
> > > > "timed\_out" : false,  
> > > > "number\_of\_nodes" : 20,  
> > > > "number\_of\_data\_nodes" : 10,  
> > > > "active\_primary\_shards" : 9,  
> > > > "active\_shards" : 36,  
> > > > "relocating\_shards" : 0,  
> > > > "initializing\_shards" : 0,  
> > > > "unassigned\_shards" : 4  
> > > > }  
> > > > State status:  
> > > > "7" : [ {  
> > > > "state" : "UNASSIGNED",  
> > > > "primary" : false,  
> > > > "node" : null,  
> > > > "relocating\_node" : null,  
> > > > "shard" : 7,  
> > > > "index" : "fs"  
> > > > }, {  
> > > > "state" : "UNASSIGNED",  
> > > > "primary" : false,  
> > > > "node" : null,  
> > > > "relocating\_node" : null,  
> > > > "shard" : 7,  
> > > > "index" : "fs"  
> > > > }, {  
> > > > "state" : "UNASSIGNED",  
> > > > "primary" : false,  
> > > > "node" : null,  
> > > > "relocating\_node" : null,  
> > > > "shard" : 7,  
> > > > "index" : "fs"  
> > > > }, {  
> > > > "state" : "UNASSIGNED",  
> > > > "primary" : true,  
> > > > "node" : null,  
> > > > "relocating\_node" : null,  
> > > > "shard" : 7,  
> > > > "index" : "fs"  
> > > > } ],

---

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [July 24, 2011, 7:33am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/6 "2011-07-24T07:33:36Z")

</div>

No, there is no multiple instances running on the same node...  
Setting gateway.local.initial\_shards=2 solved the problem.

In order to summarise, several questions remain:

1. Is it possible to see the real reason why 2 of 4 shards wasn't  
detected during system start? (checksum, bad file length, io error,  
etc...?)
2. Does it mean that only 1 node may have "bad" shard during system  
start (using default configuration and replication factor 3 for [fs]  
index). In case there are more than one, the cluster will never be  
"green"?
3. Should this parameter be set together with replication\_factor of  
the index, order to configure not only "runtime recovery", but also  
"system start up recovery" ?
4. In case several indexes exist in the system with different  
replication factors, do we need initial\_shards parameter configured  
for each one separately, and current system wide parameter may be  
problematic in case replication\_factor + 1 \< initial\_shards ?

On Jul 22, 11:21 am, Michel Conrad [michel.con...@trendiction.com](mailto:michel.con...@trendiction.com)  
wrote:

> Could it be that you managed to start two instances of elasticsearch  
> on one of your servers?  
> In that case the 2 shards missing could have been allocated to the  
> second running instance,  
> which would explain elasticsearch couldn't find the shards after  
> restarting the cluster (starting only  
> one instance of es on every server).
> 
> By looking at your data directory in the folder nodes there should  
> only be a directory called 0. If there  
> are multiple directories 0,1,2... you have been starting multiple  
> nodes on a server and the missing shards  
> may have been allocated to another node on the same server.
> 
> On Thu, Jul 21, 2011 at 7:56 PM, Shay Banon
> 
> [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > It seems like the gateway only finds 2 shards (out of 4) for that shard  
> > group. By default, it wants to find a quorum, you can change that by  
> > settings gateway.local.initial\_shards to a different value than `quorum`,  
> > for example: 2, and then it will recover.  
> > Another question is why there are only 2. When you created the index, was it  
> > create and all shards were allocated before the cluster was restarted?
> 
> > On Thu, Jul 21, 2011 at 11:11 AM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > > From tracing gateway.local, the only message related bad 7-th shard is  
> > > repeated many times:  
> > > .  
> > > .  
> > > [2011-07-21 10:53:53,745][INFO][cluster.service] [Delphi]  
> > > new\_master [Delphi][apBvicCZT4mzWZrm6wZ16Q][inet[/10.11.40.238:9300]],  
> > > reason: zen-disco-join (elected\_as\_master)  
> > > .  
> > > .  
> > > [2011-07-21 10:59:57,551][DEBUG][gateway.local] [Delphi]  
> > > [fs][7]: not allocating, number\_of\_allocated\_shards\_found [2],  
> > > required\_number [3]
> 
> > > My config is:
> > > 
> > > * * *
> > > 
> > > cluster.name : MY\_CLUSTER
> 
> > > gateway:  
> > > recover\_after\_nodes: 8  
> > > recover\_after\_time: 5m  
> > > expected\_nodes: 10
> 
> > > index.compound\_format : false  
> > > index.refresh\_interval : 10s  
> > > index.term\_index\_interval: 30
> 
> > > discovery.zen.ping.unicast:  
> > > hosts: node01:9300,node02:9300,node03:9300
> > > 
> > > * * *
> 
> > > On Jul 20, 7:47 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Can you gist your config? Also, if you set: gateway.local to TRACE, it  
> > > > will  
> > > > print all the allocation information (on the elected master) and it can  
> > > > possibly give us info as to why those shards are not allocated.
> 
> > > > On Wed, Jul 20, 2011 at 2:53 PM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > > > 
> > > > > 10 nodes, replication factor 3, local storage, 10 data nodes, 10 super  
> > > > > clients.  
> > > > > Please let me know if you need more info.
> 
> > > > > Health status:  
> > > > > {  
> > > > > "cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
> > > > > "status" : "red",  
> > > > > "timed\_out" : false,  
> > > > > "number\_of\_nodes" : 20,  
> > > > > "number\_of\_data\_nodes" : 10,  
> > > > > "active\_primary\_shards" : 9,  
> > > > > "active\_shards" : 36,  
> > > > > "relocating\_shards" : 0,  
> > > > > "initializing\_shards" : 0,  
> > > > > "unassigned\_shards" : 4  
> > > > > }  
> > > > > State status:  
> > > > > "7" : [ {  
> > > > > "state" : "UNASSIGNED",  
> > > > > "primary" : false,  
> > > > > "node" : null,  
> > > > > "relocating\_node" : null,  
> > > > > "shard" : 7,  
> > > > > "index" : "fs"  
> > > > > }, {  
> > > > > "state" : "UNASSIGNED",  
> > > > > "primary" : false,  
> > > > > "node" : null,  
> > > > > "relocating\_node" : null,  
> > > > > "shard" : 7,  
> > > > > "index" : "fs"  
> > > > > }, {  
> > > > > "state" : "UNASSIGNED",  
> > > > > "primary" : false,  
> > > > > "node" : null,  
> > > > > "relocating\_node" : null,  
> > > > > "shard" : 7,  
> > > > > "index" : "fs"  
> > > > > }, {  
> > > > > "state" : "UNASSIGNED",  
> > > > > "primary" : true,  
> > > > > "node" : null,  
> > > > > "relocating\_node" : null,  
> > > > > "shard" : 7,  
> > > > > "index" : "fs"  
> > > > > } ],

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 26, 2011, 6:16am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/7 "2011-07-26T06:16:22Z")

</div>

On Sun, Jul 24, 2011 at 10:33 AM, vadim [vpunski@gmail.com](mailto:vpunski@gmail.com) wrote:

> No, there is no multiple instances running on the same node...  
> Setting gateway.local.initial\_shards=2 solved the problem.
> 
> In order to summarise, several questions remain:
> 
> 1. Is it possible to see the real reason why 2 of 4 shards wasn't  
> detected during system start? (checksum, bad file length, io error,  
> etc...?)

It should be in the trace logging of the master node when it does full  
recovery (at least where it found which shards).

> 1. Does it mean that only 1 node may have "bad" shard during system  
> start (using default configuration and replication factor 3 for [fs]  
> index). In case there are more than one, the cluster will never be  
> "green"?

I don't know what you mean by replication factor, and I don't want to  
confuse people. For an index with "number\_of\_replicas" set to 2 (meaning an  
"additional" 2 replicas per shard), and the default quorum size shards to  
exists in order to recover, then yes, a quorum of 3 (a shard and 2 replicas)  
is 2.

> 1. Should this parameter be set together with replication\_factor of  
> the index, order to configure not only "runtime recovery", but also  
> "system start up recovery" ?

Quorum should be good enough, unless you are after something different? I  
can add a naming convention for "quorum-1", which can simplify things.

> 1. In case several indexes exist in the system with different  
> replication factors, do we need initial\_shards parameter configured  
> for each one separately, and current system wide parameter may be  
> problematic in case replication\_factor + 1 \< initial\_shards ?

Explicit value for initial\_shards can be problematic, yes, for cases where  
indices have different number of replicas. We can make it an index level  
settings as well, though I think "quorum-1" is good enough.

> On Jul 22, 11:21 am, Michel Conrad [michel.con...@trendiction.com](mailto:michel.con...@trendiction.com)  
> wrote:
> 
> > Could it be that you managed to start two instances of elasticsearch  
> > on one of your servers?  
> > In that case the 2 shards missing could have been allocated to the  
> > second running instance,  
> > which would explain elasticsearch couldn't find the shards after  
> > restarting the cluster (starting only  
> > one instance of es on every server).
> > 
> > By looking at your data directory in the folder nodes there should  
> > only be a directory called 0. If there  
> > are multiple directories 0,1,2... you have been starting multiple  
> > nodes on a server and the missing shards  
> > may have been allocated to another node on the same server.
> > 
> > On Thu, Jul 21, 2011 at 7:56 PM, Shay Banon
> > 
> > [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > It seems like the gateway only finds 2 shards (out of 4) for that shard  
> > > group. By default, it wants to find a quorum, you can change that by  
> > > settings gateway.local.initial\_shards to a different value than  
> > > `quorum`,  
> > > for example: 2, and then it will recover.  
> > > Another question is why there are only 2. When you created the index,  
> > > was it  
> > > create and all shards were allocated before the cluster was restarted?
> > 
> > > On Thu, Jul 21, 2011 at 11:11 AM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > 
> > > > From tracing gateway.local, the only message related bad 7-th shard is  
> > > > repeated many times:  
> > > > .  
> > > > .  
> > > > [2011-07-21 10:53:53,745][INFO][cluster.service] [Delphi]  
> > > > new\_master [Delphi][apBvicCZT4mzWZrm6wZ16Q][inet[/10.11.40.238:9300  
> > > > ]],  
> > > > reason: zen-disco-join (elected\_as\_master)  
> > > > .  
> > > > .  
> > > > [2011-07-21 10:59:57,551][DEBUG][gateway.local] [Delphi]  
> > > > [fs][7]: not allocating, number\_of\_allocated\_shards\_found [2],  
> > > > required\_number [3]
> > 
> > > > My config is:
> > > > 
> > > > * * *
> > > > 
> > > > cluster.name : MY\_CLUSTER
> > 
> > > > gateway:  
> > > > recover\_after\_nodes: 8  
> > > > recover\_after\_time: 5m  
> > > > expected\_nodes: 10
> > 
> > > > index.compound\_format : false  
> > > > index.refresh\_interval : 10s  
> > > > index.term\_index\_interval: 30
> > 
> > > > discovery.zen.ping.unicast:  
> > > > hosts: node01:9300,node02:9300,node03:9300
> > > > 
> > > > * * *
> > 
> > > > On Jul 20, 7:47 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > 
> > > > > Can you gist your config? Also, if you set: gateway.local to TRACE,  
> > > > > it  
> > > > > will  
> > > > > print all the allocation information (on the elected master) and it  
> > > > > can  
> > > > > possibly give us info as to why those shards are not allocated.
> > 
> > > > > On Wed, Jul 20, 2011 at 2:53 PM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > > > > 
> > > > > > 10 nodes, replication factor 3, local storage, 10 data nodes, 10  
> > > > > > super  
> > > > > > clients.  
> > > > > > Please let me know if you need more info.
> > 
> > > > > > Health status:  
> > > > > > {  
> > > > > > "cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
> > > > > > "status" : "red",  
> > > > > > "timed\_out" : false,  
> > > > > > "number\_of\_nodes" : 20,  
> > > > > > "number\_of\_data\_nodes" : 10,  
> > > > > > "active\_primary\_shards" : 9,  
> > > > > > "active\_shards" : 36,  
> > > > > > "relocating\_shards" : 0,  
> > > > > > "initializing\_shards" : 0,  
> > > > > > "unassigned\_shards" : 4  
> > > > > > }  
> > > > > > State status:  
> > > > > > "7" : [ {  
> > > > > > "state" : "UNASSIGNED",  
> > > > > > "primary" : false,  
> > > > > > "node" : null,  
> > > > > > "relocating\_node" : null,  
> > > > > > "shard" : 7,  
> > > > > > "index" : "fs"  
> > > > > > }, {  
> > > > > > "state" : "UNASSIGNED",  
> > > > > > "primary" : false,  
> > > > > > "node" : null,  
> > > > > > "relocating\_node" : null,  
> > > > > > "shard" : 7,  
> > > > > > "index" : "fs"  
> > > > > > }, {  
> > > > > > "state" : "UNASSIGNED",  
> > > > > > "primary" : false,  
> > > > > > "node" : null,  
> > > > > > "relocating\_node" : null,  
> > > > > > "shard" : 7,  
> > > > > > "index" : "fs"  
> > > > > > }, {  
> > > > > > "state" : "UNASSIGNED",  
> > > > > > "primary" : true,  
> > > > > > "node" : null,  
> > > > > > "relocating\_node" : null,  
> > > > > > "shard" : 7,  
> > > > > > "index" : "fs"  
> > > > > > } ],

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 26, 2011, 8:43am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/8 "2011-07-26T08:43:33Z")

</div>

Here are the issues that spawned out of this discussion:

- [Local Gateway: Allow to set gateway.local.initial\_shards to `quorum-1` · Issue #1160 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/1160)
- [Index Settings: Add `index.recovery.initial_shards` controlling the number of shards to exists when using local gateway · Issue #1163 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/1163)

Simple to implement, will be in 0.17.2.

On Tue, Jul 26, 2011 at 9:16 AM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> On Sun, Jul 24, 2011 at 10:33 AM, vadim [vpunski@gmail.com](mailto:vpunski@gmail.com) wrote:
> 
> > No, there is no multiple instances running on the same node...  
> > Setting gateway.local.initial\_shards=2 solved the problem.
> > 
> > In order to summarise, several questions remain:
> > 
> > 1. Is it possible to see the real reason why 2 of 4 shards wasn't  
> > detected during system start? (checksum, bad file length, io error,  
> > etc...?)
> 
> It should be in the trace logging of the master node when it does full  
> recovery (at least where it found which shards).
> 
> > 1. Does it mean that only 1 node may have "bad" shard during system  
> > start (using default configuration and replication factor 3 for [fs]  
> > index). In case there are more than one, the cluster will never be  
> > "green"?
> 
> I don't know what you mean by replication factor, and I don't want to  
> confuse people. For an index with "number\_of\_replicas" set to 2 (meaning an  
> "additional" 2 replicas per shard), and the default quorum size shards to  
> exists in order to recover, then yes, a quorum of 3 (a shard and 2 replicas)  
> is 2.
> 
> > 1. Should this parameter be set together with replication\_factor of  
> > the index, order to configure not only "runtime recovery", but also  
> > "system start up recovery" ?
> 
> Quorum should be good enough, unless you are after something different? I  
> can add a naming convention for "quorum-1", which can simplify things.
> 
> > 1. In case several indexes exist in the system with different  
> > replication factors, do we need initial\_shards parameter configured  
> > for each one separately, and current system wide parameter may be  
> > problematic in case replication\_factor + 1 \< initial\_shards ?
> 
> Explicit value for initial\_shards can be problematic, yes, for cases where  
> indices have different number of replicas. We can make it an index level  
> settings as well, though I think "quorum-1" is good enough.
> 
> > On Jul 22, 11:21 am, Michel Conrad [michel.con...@trendiction.com](mailto:michel.con...@trendiction.com)  
> > wrote:
> > 
> > > Could it be that you managed to start two instances of elasticsearch  
> > > on one of your servers?  
> > > In that case the 2 shards missing could have been allocated to the  
> > > second running instance,  
> > > which would explain elasticsearch couldn't find the shards after  
> > > restarting the cluster (starting only  
> > > one instance of es on every server).
> > > 
> > > By looking at your data directory in the folder nodes there should  
> > > only be a directory called 0. If there  
> > > are multiple directories 0,1,2... you have been starting multiple  
> > > nodes on a server and the missing shards  
> > > may have been allocated to another node on the same server.
> > > 
> > > On Thu, Jul 21, 2011 at 7:56 PM, Shay Banon
> > > 
> > > [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > It seems like the gateway only finds 2 shards (out of 4) for that  
> > > > shard  
> > > > group. By default, it wants to find a quorum, you can change that by  
> > > > settings gateway.local.initial\_shards to a different value than  
> > > > `quorum`,  
> > > > for example: 2, and then it will recover.  
> > > > Another question is why there are only 2. When you created the index,  
> > > > was it  
> > > > create and all shards were allocated before the cluster was restarted?
> > > 
> > > > On Thu, Jul 21, 2011 at 11:11 AM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > > 
> > > > > From tracing gateway.local, the only message related bad 7-th shard  
> > > > > is  
> > > > > repeated many times:  
> > > > > .  
> > > > > .  
> > > > > [2011-07-21 10:53:53,745][INFO][cluster.service] [Delphi]  
> > > > > new\_master [Delphi][apBvicCZT4mzWZrm6wZ16Q][inet[/10.11.40.238:9300  
> > > > > ]],  
> > > > > reason: zen-disco-join (elected\_as\_master)  
> > > > > .  
> > > > > .  
> > > > > [2011-07-21 10:59:57,551][DEBUG][gateway.local] [Delphi]  
> > > > > [fs][7]: not allocating, number\_of\_allocated\_shards\_found [2],  
> > > > > required\_number [3]
> > > 
> > > > > My config is:
> > > > > 
> > > > > * * *
> > > > > 
> > > > > cluster.name : MY\_CLUSTER
> > > 
> > > > > gateway:  
> > > > > recover\_after\_nodes: 8  
> > > > > recover\_after\_time: 5m  
> > > > > expected\_nodes: 10
> > > 
> > > > > index.compound\_format : false  
> > > > > index.refresh\_interval : 10s  
> > > > > index.term\_index\_interval: 30
> > > 
> > > > > discovery.zen.ping.unicast:  
> > > > > hosts: node01:9300,node02:9300,node03:9300
> > > > > 
> > > > > * * *
> > > 
> > > > > On Jul 20, 7:47 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > 
> > > > > > Can you gist your config? Also, if you set: gateway.local to TRACE,  
> > > > > > it  
> > > > > > will  
> > > > > > print all the allocation information (on the elected master) and it  
> > > > > > can  
> > > > > > possibly give us info as to why those shards are not allocated.
> > > 
> > > > > > On Wed, Jul 20, 2011 at 2:53 PM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > > > > > 
> > > > > > > 10 nodes, replication factor 3, local storage, 10 data nodes, 10  
> > > > > > > super  
> > > > > > > clients.  
> > > > > > > Please let me know if you need more info.
> > > 
> > > > > > > Health status:  
> > > > > > > {  
> > > > > > > "cluster\_name" : "CMWELL\_INDEX\_PRODUCTION\_CLUSTER",  
> > > > > > > "status" : "red",  
> > > > > > > "timed\_out" : false,  
> > > > > > > "number\_of\_nodes" : 20,  
> > > > > > > "number\_of\_data\_nodes" : 10,  
> > > > > > > "active\_primary\_shards" : 9,  
> > > > > > > "active\_shards" : 36,  
> > > > > > > "relocating\_shards" : 0,  
> > > > > > > "initializing\_shards" : 0,  
> > > > > > > "unassigned\_shards" : 4  
> > > > > > > }  
> > > > > > > State status:  
> > > > > > > "7" : [ {  
> > > > > > > "state" : "UNASSIGNED",  
> > > > > > > "primary" : false,  
> > > > > > > "node" : null,  
> > > > > > > "relocating\_node" : null,  
> > > > > > > "shard" : 7,  
> > > > > > > "index" : "fs"  
> > > > > > > }, {  
> > > > > > > "state" : "UNASSIGNED",  
> > > > > > > "primary" : false,  
> > > > > > > "node" : null,  
> > > > > > > "relocating\_node" : null,  
> > > > > > > "shard" : 7,  
> > > > > > > "index" : "fs"  
> > > > > > > }, {  
> > > > > > > "state" : "UNASSIGNED",  
> > > > > > > "primary" : false,  
> > > > > > > "node" : null,  
> > > > > > > "relocating\_node" : null,  
> > > > > > > "shard" : 7,  
> > > > > > > "index" : "fs"  
> > > > > > > }, {  
> > > > > > > "state" : "UNASSIGNED",  
> > > > > > > "primary" : true,  
> > > > > > > "node" : null,  
> > > > > > > "relocating\_node" : null,  
> > > > > > > "shard" : 7,  
> > > > > > > "index" : "fs"  
> > > > > > > } ],

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:59am UTC](https://discuss.elastic.co/t/data-lost-after-full-cluster-restart/4906/9 "2017-07-06T03:59:24Z")

</div>


