# ES Ate My Shards/Indexes

**URL:** <https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738>\
**Category:** Elasticsearch\
**Created:** [February 17, 2012, 6:34pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738 "2012-02-17T18:34:23Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman/32/2224_2.png) [@Kenneth\_Loafman](https://discuss.elastic.co/u/Kenneth_Loafman)\
**Post date:** [February 17, 2012, 6:34pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/1 "2012-02-17T18:34:23Z")

</div>

Hi,

We're on ES 18.6 driving a 4-node cluster on RackSpace.

Last night we had a nework outage on two of the nodes and our 4-node  
cluster morphed into 2 2-node clusters. I think that's what happened  
anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
and the cluster reformed into one, however, it's sticking on getting out of  
red.

{  
"cluster\_name" : "elasticsearch",  
"status" : "red",  
"timed\_out" : false,  
"number\_of\_nodes" : 4,  
"number\_of\_data\_nodes" : 4,  
"active\_primary\_shards" : 234,  
"active\_shards" : 426,  
"relocating\_shards" : 0,  
"initializing\_shards" : 8,  
"unassigned\_shards" : 126  
}

It will stay in this mode for a long time and the '\_status' will show a  
bunch of

```
"failures" : [ {
  "index" : "co0181ca0711",
  "shard" : 1,
  "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1]

```

]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
nested: IndexMissingException[[co0181ca0711] missing]; "  
}, {

which seem to come and go, but not get initialized. With 2 shards and 1  
replica, it seems that it should be able to recover the missing index from  
the other shard or the replica, but it sticks at this point until I  
manually delete what's left of the index. Was this due to the split-brain  
issue or is this just a limitation of ES? Is there a way to recover the  
missing index from the replica? How do I find the replicas?

...Thanks,  
...Ken

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 17, 2012, 6:45pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/2 "2012-02-17T18:45:11Z")

</div>

Do you see anything in the logs? It seems like there are 8 initializing shards.

On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:

> Hi,
> 
> We're on ES 18.6 driving a 4-node cluster on RackSpace.
> 
> Last night we had a nework outage on two of the nodes and our 4-node cluster morphed into 2 2-node clusters. I think that's what happened anyway. We shut all 4 nodes down cleanly, brought them up one at a time and the cluster reformed into one, however, it's sticking on getting out of red.
> 
> {  
> "cluster\_name" : "elasticsearch",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 4,  
> "number\_of\_data\_nodes" : 4,  
> "active\_primary\_shards" : 234,  
> "active\_shards" : 426,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 8,  
> "unassigned\_shards" : 126  
> }
> 
> It will stay in this mode for a long time and the '\_status' will show a bunch of
> 
> ```
> "failures" : [ { 
> "index" : "co0181ca0711",
> "shard" : 1,
> "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]]; nested: IndexMissingException[[co0181ca0711] missing]; "
> }, {
> 
> ```
> 
> which seem to come and go, but not get initialized. With 2 shards and 1 replica, it seems that it should be able to recover the missing index from the other shard or the replica, but it sticks at this point until I manually delete what's left of the index. Was this due to the split-brain issue or is this just a limitation of ES? Is there a way to recover the missing index from the replica? How do I find the replicas?
> 
> ...Thanks,  
> ...Ken

---

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman_2/32/2224_2.png) [@Kenneth\_Loafman\_2](https://discuss.elastic.co/u/Kenneth_Loafman_2)\
**Post date:** [February 17, 2012, 6:49pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/3 "2012-02-17T18:49:09Z")

</div>

Yes, a whole bunch of messages repeating like this:

[2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot]  
received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ],  
[P], s[INITIALIZING], reason [Failed to start shard, message  
[IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for  
local recovery (post api), should exists, but doesn't]]]

On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Do you see anything in the logs? It seems like there are 8 initializing  
> shards.
> 
> On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> 
> Hi,
> 
> We're on ES 18.6 driving a 4-node cluster on RackSpace.
> 
> Last night we had a nework outage on two of the nodes and our 4-node  
> cluster morphed into 2 2-node clusters. I think that's what happened  
> anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
> and the cluster reformed into one, however, it's sticking on getting out of  
> red.
> 
> {  
> "cluster\_name" : "elasticsearch",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 4,  
> "number\_of\_data\_nodes" : 4,  
> "active\_primary\_shards" : 234,  
> "active\_shards" : 426,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 8,  
> "unassigned\_shards" : 126  
> }
> 
> It will stay in this mode for a long time and the '\_status' will show a  
> bunch of
> 
> ```
> "failures" : [ {
> "index" : "co0181ca0711",
> "shard" : 1,
> "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1]
> 
> ```
> 
> ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
> nested: IndexMissingException[[co0181ca0711] missing]; "  
> }, {
> 
> which seem to come and go, but not get initialized. With 2 shards and 1  
> replica, it seems that it should be able to recover the missing index from  
> the other shard or the replica, but it sticks at this point until I  
> manually delete what's left of the index. Was this due to the split-brain  
> issue or is this just a limitation of ES? Is there a way to recover the  
> missing index from the replica? How do I find the replicas?
> 
> ...Thanks,  
> ...Ken

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 17, 2012, 6:54pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/4 "2012-02-17T18:54:17Z")

</div>

This means that the shard was supposed to exist on that node, but it can't be found, are you sure nothing was deleted?

On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:

> Yes, a whole bunch of messages repeating like this:
> 
> [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot] received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ], [P], s[INITIALIZING], reason [Failed to start shard, message [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for local recovery (post api), should exists, but doesn't]]]
> 
> On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> 
> > Do you see anything in the logs? It seems like there are 8 initializing shards.
> > 
> > On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> > 
> > > Hi,
> > > 
> > > We're on ES 18.6 driving a 4-node cluster on RackSpace.
> > > 
> > > Last night we had a nework outage on two of the nodes and our 4-node cluster morphed into 2 2-node clusters. I think that's what happened anyway. We shut all 4 nodes down cleanly, brought them up one at a time and the cluster reformed into one, however, it's sticking on getting out of red.
> > > 
> > > {  
> > > "cluster\_name" : "elasticsearch",  
> > > "status" : "red",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 4,  
> > > "number\_of\_data\_nodes" : 4,  
> > > "active\_primary\_shards" : 234,  
> > > "active\_shards" : 426,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 8,  
> > > "unassigned\_shards" : 126  
> > > }
> > > 
> > > It will stay in this mode for a long time and the '\_status' will show a bunch of
> > > 
> > > ```
> > > "failures" : [ { 
> > > "index" : "co0181ca0711",
> > > "shard" : 1,
> > > "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]]; nested: IndexMissingException[[co0181ca0711] missing]; "
> > > }, {
> > > 
> > > ```
> > > 
> > > which seem to come and go, but not get initialized. With 2 shards and 1 replica, it seems that it should be able to recover the missing index from the other shard or the replica, but it sticks at this point until I manually delete what's left of the index. Was this due to the split-brain issue or is this just a limitation of ES? Is there a way to recover the missing index from the replica? How do I find the replicas?
> > > 
> > > ...Thanks,  
> > > ...Ken

---

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman_2/32/2224_2.png) [@Kenneth\_Loafman\_2](https://discuss.elastic.co/u/Kenneth_Loafman_2)\
**Post date:** [February 17, 2012, 6:58pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/5 "2012-02-17T18:58:21Z")

</div>

Nothing was deleted manually or through curl. Is this recoverable at all?  
What happened? Was this because of the split-cluster condition?

On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> This means that the shard was supposed to exist on that node, but it  
> can't be found, are you sure nothing was deleted?
> 
> On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> 
> Yes, a whole bunch of messages repeating like this:
> 
> [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot]  
> received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ],  
> [P], s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for  
> local recovery (post api), should exists, but doesn't]]]
> 
> On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> Do you see anything in the logs? It seems like there are 8 initializing  
> shards.
> 
> On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> 
> Hi,
> 
> We're on ES 18.6 driving a 4-node cluster on RackSpace.
> 
> Last night we had a nework outage on two of the nodes and our 4-node  
> cluster morphed into 2 2-node clusters. I think that's what happened  
> anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
> and the cluster reformed into one, however, it's sticking on getting out of  
> red.
> 
> {  
> "cluster\_name" : "elasticsearch",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 4,  
> "number\_of\_data\_nodes" : 4,  
> "active\_primary\_shards" : 234,  
> "active\_shards" : 426,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 8,  
> "unassigned\_shards" : 126  
> }
> 
> It will stay in this mode for a long time and the '\_status' will show a  
> bunch of
> 
> ```
> "failures" : [ {
> "index" : "co0181ca0711",
> "shard" : 1,
> "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1]
> 
> ```
> 
> ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
> nested: IndexMissingException[[co0181ca0711] missing]; "  
> }, {
> 
> which seem to come and go, but not get initialized. With 2 shards and 1  
> replica, it seems that it should be able to recover the missing index from  
> the other shard or the replica, but it sticks at this point until I  
> manually delete what's left of the index. Was this due to the split-brain  
> issue or is this just a limitation of ES? Is there a way to recover the  
> missing index from the replica? How do I find the replicas?
> 
> ...Thanks,  
> ...Ken

---

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman_2/32/2224_2.png) [@Kenneth\_Loafman\_2](https://discuss.elastic.co/u/Kenneth_Loafman_2)\
**Post date:** [February 17, 2012, 7:55pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/6 "2012-02-17T19:55:25Z")

</div>

Hmm, something else is going on. Two of the indexes that were OK  
originally are now showing IndexShardMissingException.

ES is indeed hungry!

On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:

> Nothing was deleted manually or through curl. Is this recoverable at all?  
> What happened? Was this because of the split-cluster condition?
> 
> On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > This means that the shard was supposed to exist on that node, but it  
> > can't be found, are you sure nothing was deleted?
> > 
> > On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> > 
> > Yes, a whole bunch of messages repeating like this:
> > 
> > [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot]  
> > received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ],  
> > [P], s[INITIALIZING], reason [Failed to start shard, message  
> > [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for  
> > local recovery (post api), should exists, but doesn't]]]
> > 
> > On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > Do you see anything in the logs? It seems like there are 8 initializing  
> > shards.
> > 
> > On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> > 
> > Hi,
> > 
> > We're on ES 18.6 driving a 4-node cluster on RackSpace.
> > 
> > Last night we had a nework outage on two of the nodes and our 4-node  
> > cluster morphed into 2 2-node clusters. I think that's what happened  
> > anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
> > and the cluster reformed into one, however, it's sticking on getting out of  
> > red.
> > 
> > {  
> > "cluster\_name" : "elasticsearch",  
> > "status" : "red",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 4,  
> > "number\_of\_data\_nodes" : 4,  
> > "active\_primary\_shards" : 234,  
> > "active\_shards" : 426,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 8,  
> > "unassigned\_shards" : 126  
> > }
> > 
> > It will stay in this mode for a long time and the '\_status' will show a  
> > bunch of
> > 
> > ```
> > "failures" : [ {
> > "index" : "co0181ca0711",
> > "shard" : 1,
> > "reason" :
> > 
> > ```
> > 
> > "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested:  
> > RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
> > nested: IndexMissingException[[co0181ca0711] missing]; "  
> > }, {
> > 
> > which seem to come and go, but not get initialized. With 2 shards and 1  
> > replica, it seems that it should be able to recover the missing index from  
> > the other shard or the replica, but it sticks at this point until I  
> > manually delete what's left of the index. Was this due to the split-brain  
> > issue or is this just a limitation of ES? Is there a way to recover the  
> > missing index from the replica? How do I find the replicas?
> > 
> > ...Thanks,  
> > ...Ken

---

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman_2/32/2224_2.png) [@Kenneth\_Loafman\_2](https://discuss.elastic.co/u/Kenneth_Loafman_2)\
**Post date:** [February 17, 2012, 9:32pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/7 "2012-02-17T21:32:13Z")

</div>

Any ideas? I've shut down again, forced fsck on next boot, rebooted and  
restarted. No real problems found, so we can rule that out. The logs are  
too big to gist. What would you need from them if I could find it?

On Fri, Feb 17, 2012 at 1:55 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:

> Hmm, something else is going on. Two of the indexes that were OK  
> originally are now showing IndexShardMissingException.
> 
> ES is indeed hungry!
> 
> On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:
> 
> > Nothing was deleted manually or through curl. Is this recoverable at  
> > all? What happened? Was this because of the split-cluster condition?
> > 
> > On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > > This means that the shard was supposed to exist on that node, but it  
> > > can't be found, are you sure nothing was deleted?
> > > 
> > > On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> > > 
> > > Yes, a whole bunch of messages repeating like this:
> > > 
> > > [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot]  
> > > received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ],  
> > > [P], s[INITIALIZING], reason [Failed to start shard, message  
> > > [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for  
> > > local recovery (post api), should exists, but doesn't]]]
> > > 
> > > On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > 
> > > Do you see anything in the logs? It seems like there are 8 initializing  
> > > shards.
> > > 
> > > On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> > > 
> > > Hi,
> > > 
> > > We're on ES 18.6 driving a 4-node cluster on RackSpace.
> > > 
> > > Last night we had a nework outage on two of the nodes and our 4-node  
> > > cluster morphed into 2 2-node clusters. I think that's what happened  
> > > anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
> > > and the cluster reformed into one, however, it's sticking on getting out of  
> > > red.
> > > 
> > > {  
> > > "cluster\_name" : "elasticsearch",  
> > > "status" : "red",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 4,  
> > > "number\_of\_data\_nodes" : 4,  
> > > "active\_primary\_shards" : 234,  
> > > "active\_shards" : 426,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 8,  
> > > "unassigned\_shards" : 126  
> > > }
> > > 
> > > It will stay in this mode for a long time and the '\_status' will show a  
> > > bunch of
> > > 
> > > ```
> > > "failures" : [ {
> > > "index" : "co0181ca0711",
> > > "shard" : 1,
> > > "reason" :
> > > 
> > > ```
> > > 
> > > "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested:  
> > > RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
> > > nested: IndexMissingException[[co0181ca0711] missing]; "  
> > > }, {
> > > 
> > > which seem to come and go, but not get initialized. With 2 shards and 1  
> > > replica, it seems that it should be able to recover the missing index from  
> > > the other shard or the replica, but it sticks at this point until I  
> > > manually delete what's left of the index. Was this due to the split-brain  
> > > issue or is this just a limitation of ES? Is there a way to recover the  
> > > missing index from the replica? How do I find the replicas?
> > > 
> > > ...Thanks,  
> > > ...Ken

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 17, 2012, 10:35pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/8 "2012-02-17T22:35:03Z")

</div>

What I meant by data deleted is that some data was deleted from the file system by any chance? I suggest you start to delete the problematic indexes that hold the problematic shards when the cluster is up.

On Friday, February 17, 2012 at 11:32 PM, Kenneth Loafman wrote:

> Any ideas? I've shut down again, forced fsck on next boot, rebooted and restarted. No real problems found, so we can rule that out. The logs are too big to gist. What would you need from them if I could find it?
> 
> On Fri, Feb 17, 2012 at 1:55 PM, Kenneth Loafman \<[kenneth@loafman.com](mailto:kenneth@loafman.com) ([mailto:kenneth@loafman.com](mailto:kenneth@loafman.com))\> wrote:
> 
> > Hmm, something else is going on. Two of the indexes that were OK originally are now showing IndexShardMissingException.
> > 
> > ES is indeed hungry!
> > 
> > On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman \<[kenneth@loafman.com](mailto:kenneth@loafman.com) ([mailto:kenneth@loafman.com](mailto:kenneth@loafman.com))\> wrote:
> > 
> > > Nothing was deleted manually or through curl. Is this recoverable at all? What happened? Was this because of the split-cluster condition?
> > > 
> > > On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > > 
> > > > This means that the shard was supposed to exist on that node, but it can't be found, are you sure nothing was deleted?
> > > > 
> > > > On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> > > > 
> > > > > Yes, a whole bunch of messages repeating like this:
> > > > > 
> > > > > [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot] received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ], [P], s[INITIALIZING], reason [Failed to start shard, message [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for local recovery (post api), should exists, but doesn't]]]
> > > > > 
> > > > > On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > > > > 
> > > > > > Do you see anything in the logs? It seems like there are 8 initializing shards.
> > > > > > 
> > > > > > On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> > > > > > 
> > > > > > > Hi,
> > > > > > > 
> > > > > > > We're on ES 18.6 driving a 4-node cluster on RackSpace.
> > > > > > > 
> > > > > > > Last night we had a nework outage on two of the nodes and our 4-node cluster morphed into 2 2-node clusters. I think that's what happened anyway. We shut all 4 nodes down cleanly, brought them up one at a time and the cluster reformed into one, however, it's sticking on getting out of red.
> > > > > > > 
> > > > > > > {  
> > > > > > > "cluster\_name" : "elasticsearch",  
> > > > > > > "status" : "red",  
> > > > > > > "timed\_out" : false,  
> > > > > > > "number\_of\_nodes" : 4,  
> > > > > > > "number\_of\_data\_nodes" : 4,  
> > > > > > > "active\_primary\_shards" : 234,  
> > > > > > > "active\_shards" : 426,  
> > > > > > > "relocating\_shards" : 0,  
> > > > > > > "initializing\_shards" : 8,  
> > > > > > > "unassigned\_shards" : 126  
> > > > > > > }
> > > > > > > 
> > > > > > > It will stay in this mode for a long time and the '\_status' will show a bunch of
> > > > > > > 
> > > > > > > ```
> > > > > > > "failures" : [ { 
> > > > > > > "index" : "co0181ca0711",
> > > > > > > "shard" : 1,
> > > > > > > "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]]; nested: IndexMissingException[[co0181ca0711] missing]; "
> > > > > > > }, {
> > > > > > > 
> > > > > > > ```
> > > > > > > 
> > > > > > > which seem to come and go, but not get initialized. With 2 shards and 1 replica, it seems that it should be able to recover the missing index from the other shard or the replica, but it sticks at this point until I manually delete what's left of the index. Was this due to the split-brain issue or is this just a limitation of ES? Is there a way to recover the missing index from the replica? How do I find the replicas?
> > > > > > > 
> > > > > > > ...Thanks,  
> > > > > > > ...Ken

---

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman_2/32/2224_2.png) [@Kenneth\_Loafman\_2](https://discuss.elastic.co/u/Kenneth_Loafman_2)\
**Post date:** [February 17, 2012, 10:58pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/9 "2012-02-17T22:58:13Z")

</div>

No data was deleted from the filesystem. What I found when I looked was an  
empty directory where the shard should have been.

On Fri, Feb 17, 2012 at 4:35 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> What I meant by data deleted is that some data was deleted from the file  
> system by any chance? I suggest you start to delete the problematic indexes  
> that hold the problematic shards when the cluster is up.
> 
> On Friday, February 17, 2012 at 11:32 PM, Kenneth Loafman wrote:
> 
> Any ideas? I've shut down again, forced fsck on next boot, rebooted and  
> restarted. No real problems found, so we can rule that out. The logs are  
> too big to gist. What would you need from them if I could find it?
> 
> On Fri, Feb 17, 2012 at 1:55 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:
> 
> Hmm, something else is going on. Two of the indexes that were OK  
> originally are now showing IndexShardMissingException.
> 
> ES is indeed hungry!
> 
> On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:
> 
> Nothing was deleted manually or through curl. Is this recoverable at all?  
> What happened? Was this because of the split-cluster condition?
> 
> On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> This means that the shard was supposed to exist on that node, but it  
> can't be found, are you sure nothing was deleted?
> 
> On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> 
> Yes, a whole bunch of messages repeating like this:
> 
> [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot]  
> received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ],  
> [P], s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for  
> local recovery (post api), should exists, but doesn't]]]
> 
> On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> Do you see anything in the logs? It seems like there are 8 initializing  
> shards.
> 
> On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> 
> Hi,
> 
> We're on ES 18.6 driving a 4-node cluster on RackSpace.
> 
> Last night we had a nework outage on two of the nodes and our 4-node  
> cluster morphed into 2 2-node clusters. I think that's what happened  
> anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
> and the cluster reformed into one, however, it's sticking on getting out of  
> red.
> 
> {  
> "cluster\_name" : "elasticsearch",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 4,  
> "number\_of\_data\_nodes" : 4,  
> "active\_primary\_shards" : 234,  
> "active\_shards" : 426,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 8,  
> "unassigned\_shards" : 126  
> }
> 
> It will stay in this mode for a long time and the '\_status' will show a  
> bunch of
> 
> ```
> "failures" : [ {
> "index" : "co0181ca0711",
> "shard" : 1,
> "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1]
> 
> ```
> 
> ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
> nested: IndexMissingException[[co0181ca0711] missing]; "  
> }, {
> 
> which seem to come and go, but not get initialized. With 2 shards and 1  
> replica, it seems that it should be able to recover the missing index from  
> the other shard or the replica, but it sticks at this point until I  
> manually delete what's left of the index. Was this due to the split-brain  
> issue or is this just a limitation of ES? Is there a way to recover the  
> missing index from the replica? How do I find the replicas?
> 
> ...Thanks,  
> ...Ken

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 20, 2012, 12:56pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/10 "2012-02-20T12:56:49Z")

</div>

Thats strange… . In 0.19, we have a better storage system for local gateway, where state is stored within each index/shard, instead of globally on the node level. I am still not sure what caused the data to be removed though…, elasticsearch does not remove data on its own unless instructed to.

On Saturday, February 18, 2012 at 12:58 AM, Kenneth Loafman wrote:

> No data was deleted from the filesystem. What I found when I looked was an empty directory where the shard should have been.
> 
> On Fri, Feb 17, 2012 at 4:35 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> 
> > What I meant by data deleted is that some data was deleted from the file system by any chance? I suggest you start to delete the problematic indexes that hold the problematic shards when the cluster is up.
> > 
> > On Friday, February 17, 2012 at 11:32 PM, Kenneth Loafman wrote:
> > 
> > > Any ideas? I've shut down again, forced fsck on next boot, rebooted and restarted. No real problems found, so we can rule that out. The logs are too big to gist. What would you need from them if I could find it?
> > > 
> > > On Fri, Feb 17, 2012 at 1:55 PM, Kenneth Loafman \<[kenneth@loafman.com](mailto:kenneth@loafman.com) ([mailto:kenneth@loafman.com](mailto:kenneth@loafman.com))\> wrote:
> > > 
> > > > Hmm, something else is going on. Two of the indexes that were OK originally are now showing IndexShardMissingException.
> > > > 
> > > > ES is indeed hungry!
> > > > 
> > > > On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman \<[kenneth@loafman.com](mailto:kenneth@loafman.com) ([mailto:kenneth@loafman.com](mailto:kenneth@loafman.com))\> wrote:
> > > > 
> > > > > Nothing was deleted manually or through curl. Is this recoverable at all? What happened? Was this because of the split-cluster condition?
> > > > > 
> > > > > On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > > > > 
> > > > > > This means that the shard was supposed to exist on that node, but it can't be found, are you sure nothing was deleted?
> > > > > > 
> > > > > > On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> > > > > > 
> > > > > > > Yes, a whole bunch of messages repeating like this:
> > > > > > > 
> > > > > > > [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot] received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ], [P], s[INITIALIZING], reason [Failed to start shard, message [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for local recovery (post api), should exists, but doesn't]]]
> > > > > > > 
> > > > > > > On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > > > > > > 
> > > > > > > > Do you see anything in the logs? It seems like there are 8 initializing shards.
> > > > > > > > 
> > > > > > > > On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> > > > > > > > 
> > > > > > > > > Hi,
> > > > > > > > > 
> > > > > > > > > We're on ES 18.6 driving a 4-node cluster on RackSpace.
> > > > > > > > > 
> > > > > > > > > Last night we had a nework outage on two of the nodes and our 4-node cluster morphed into 2 2-node clusters. I think that's what happened anyway. We shut all 4 nodes down cleanly, brought them up one at a time and the cluster reformed into one, however, it's sticking on getting out of red.
> > > > > > > > > 
> > > > > > > > > {  
> > > > > > > > > "cluster\_name" : "elasticsearch",  
> > > > > > > > > "status" : "red",  
> > > > > > > > > "timed\_out" : false,  
> > > > > > > > > "number\_of\_nodes" : 4,  
> > > > > > > > > "number\_of\_data\_nodes" : 4,  
> > > > > > > > > "active\_primary\_shards" : 234,  
> > > > > > > > > "active\_shards" : 426,  
> > > > > > > > > "relocating\_shards" : 0,  
> > > > > > > > > "initializing\_shards" : 8,  
> > > > > > > > > "unassigned\_shards" : 126  
> > > > > > > > > }
> > > > > > > > > 
> > > > > > > > > It will stay in this mode for a long time and the '\_status' will show a bunch of
> > > > > > > > > 
> > > > > > > > > ```
> > > > > > > > > "failures" : [ {  
> > > > > > > > > "index" : "co0181ca0711",
> > > > > > > > > "shard" : 1,
> > > > > > > > > "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]]; nested: IndexMissingException[[co0181ca0711] missing]; "
> > > > > > > > > }, {
> > > > > > > > > 
> > > > > > > > > ```
> > > > > > > > > 
> > > > > > > > > which seem to come and go, but not get initialized. With 2 shards and 1 replica, it seems that it should be able to recover the missing index from the other shard or the replica, but it sticks at this point until I manually delete what's left of the index. Was this due to the split-brain issue or is this just a limitation of ES? Is there a way to recover the missing index from the replica? How do I find the replicas?
> > > > > > > > > 
> > > > > > > > > ...Thanks,  
> > > > > > > > > ...Ken

---

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman_2/32/2224_2.png) [@Kenneth\_Loafman\_2](https://discuss.elastic.co/u/Kenneth_Loafman_2)\
**Post date:** [February 20, 2012, 3:34pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/11 "2012-02-20T15:34:13Z")

</div>

But data will be removed if a shard relocates, right? So what happens in a  
split-brain situation when the brain is put back together?

On Mon, Feb 20, 2012 at 6:56 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Thats strange… . In 0.19, we have a better storage system for local  
> gateway, where state is stored within each index/shard, instead of globally  
> on the node level. I am still not sure what caused the data to be removed  
> though…, elasticsearch does not remove data on its own unless instructed  
> to.
> 
> On Saturday, February 18, 2012 at 12:58 AM, Kenneth Loafman wrote:
> 
> No data was deleted from the filesystem. What I found when I looked was  
> an empty directory where the shard should have been.
> 
> On Fri, Feb 17, 2012 at 4:35 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> What I meant by data deleted is that some data was deleted from the file  
> system by any chance? I suggest you start to delete the problematic indexes  
> that hold the problematic shards when the cluster is up.
> 
> On Friday, February 17, 2012 at 11:32 PM, Kenneth Loafman wrote:
> 
> Any ideas? I've shut down again, forced fsck on next boot, rebooted and  
> restarted. No real problems found, so we can rule that out. The logs are  
> too big to gist. What would you need from them if I could find it?
> 
> On Fri, Feb 17, 2012 at 1:55 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:
> 
> Hmm, something else is going on. Two of the indexes that were OK  
> originally are now showing IndexShardMissingException.
> 
> ES is indeed hungry!
> 
> On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:
> 
> Nothing was deleted manually or through curl. Is this recoverable at all?  
> What happened? Was this because of the split-cluster condition?
> 
> On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> This means that the shard was supposed to exist on that node, but it  
> can't be found, are you sure nothing was deleted?
> 
> On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> 
> Yes, a whole bunch of messages repeating like this:
> 
> [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot]  
> received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ],  
> [P], s[INITIALIZING], reason [Failed to start shard, message  
> [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for  
> local recovery (post api), should exists, but doesn't]]]
> 
> On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> Do you see anything in the logs? It seems like there are 8 initializing  
> shards.
> 
> On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> 
> Hi,
> 
> We're on ES 18.6 driving a 4-node cluster on RackSpace.
> 
> Last night we had a nework outage on two of the nodes and our 4-node  
> cluster morphed into 2 2-node clusters. I think that's what happened  
> anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
> and the cluster reformed into one, however, it's sticking on getting out of  
> red.
> 
> {  
> "cluster\_name" : "elasticsearch",  
> "status" : "red",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 4,  
> "number\_of\_data\_nodes" : 4,  
> "active\_primary\_shards" : 234,  
> "active\_shards" : 426,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 8,  
> "unassigned\_shards" : 126  
> }
> 
> It will stay in this mode for a long time and the '\_status' will show a  
> bunch of
> 
> ```
> "failures" : [ {
> "index" : "co0181ca0711",
> "shard" : 1,
> "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1]
> 
> ```
> 
> ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
> nested: IndexMissingException[[co0181ca0711] missing]; "  
> }, {
> 
> which seem to come and go, but not get initialized. With 2 shards and 1  
> replica, it seems that it should be able to recover the missing index from  
> the other shard or the replica, but it sticks at this point until I  
> manually delete what's left of the index. Was this due to the split-brain  
> issue or is this just a limitation of ES? Is there a way to recover the  
> missing index from the replica? How do I find the replicas?
> 
> ...Thanks,  
> ...Ken

---

<div class="post-metadata">

**Author:** ![Kenneth\_Loafman\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kenneth_loafman_2/32/2224_2.png) [@Kenneth\_Loafman\_2](https://discuss.elastic.co/u/Kenneth_Loafman_2)\
**Post date:** [February 21, 2012, 3:30pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/12 "2012-02-21T15:30:00Z")

</div>

I'd like to understand what could have caused this loss of shards. What  
data do you need in order to track it down?

On Mon, Feb 20, 2012 at 9:34 AM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:

> But data will be removed if a shard relocates, right? So what happens in  
> a split-brain situation when the brain is put back together?
> 
> On Mon, Feb 20, 2012 at 6:56 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Thats strange… . In 0.19, we have a better storage system for local  
> > gateway, where state is stored within each index/shard, instead of globally  
> > on the node level. I am still not sure what caused the data to be removed  
> > though…, elasticsearch does not remove data on its own unless instructed  
> > to.
> > 
> > On Saturday, February 18, 2012 at 12:58 AM, Kenneth Loafman wrote:
> > 
> > No data was deleted from the filesystem. What I found when I looked was  
> > an empty directory where the shard should have been.
> > 
> > On Fri, Feb 17, 2012 at 4:35 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > What I meant by data deleted is that some data was deleted from the file  
> > system by any chance? I suggest you start to delete the problematic indexes  
> > that hold the problematic shards when the cluster is up.
> > 
> > On Friday, February 17, 2012 at 11:32 PM, Kenneth Loafman wrote:
> > 
> > Any ideas? I've shut down again, forced fsck on next boot, rebooted and  
> > restarted. No real problems found, so we can rule that out. The logs are  
> > too big to gist. What would you need from them if I could find it?
> > 
> > On Fri, Feb 17, 2012 at 1:55 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:
> > 
> > Hmm, something else is going on. Two of the indexes that were OK  
> > originally are now showing IndexShardMissingException.
> > 
> > ES is indeed hungry!
> > 
> > On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman [kenneth@loafman.com](mailto:kenneth@loafman.com)wrote:
> > 
> > Nothing was deleted manually or through curl. Is this recoverable at  
> > all? What happened? Was this because of the split-cluster condition?
> > 
> > On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > This means that the shard was supposed to exist on that node, but it  
> > can't be found, are you sure nothing was deleted?
> > 
> > On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> > 
> > Yes, a whole bunch of messages repeating like this:
> > 
> > [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot]  
> > received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ],  
> > [P], s[INITIALIZING], reason [Failed to start shard, message  
> > [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for  
> > local recovery (post api), should exists, but doesn't]]]
> > 
> > On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > Do you see anything in the logs? It seems like there are 8 initializing  
> > shards.
> > 
> > On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> > 
> > Hi,
> > 
> > We're on ES 18.6 driving a 4-node cluster on RackSpace.
> > 
> > Last night we had a nework outage on two of the nodes and our 4-node  
> > cluster morphed into 2 2-node clusters. I think that's what happened  
> > anyway. We shut all 4 nodes down cleanly, brought them up one at a time  
> > and the cluster reformed into one, however, it's sticking on getting out of  
> > red.
> > 
> > {  
> > "cluster\_name" : "elasticsearch",  
> > "status" : "red",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 4,  
> > "number\_of\_data\_nodes" : 4,  
> > "active\_primary\_shards" : 234,  
> > "active\_shards" : 426,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 8,  
> > "unassigned\_shards" : 126  
> > }
> > 
> > It will stay in this mode for a long time and the '\_status' will show a  
> > bunch of
> > 
> > ```
> > "failures" : [ {
> > "index" : "co0181ca0711",
> > "shard" : 1,
> > "reason" :
> > 
> > ```
> > 
> > "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested:  
> > RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]];  
> > nested: IndexMissingException[[co0181ca0711] missing]; "  
> > }, {
> > 
> > which seem to come and go, but not get initialized. With 2 shards and 1  
> > replica, it seems that it should be able to recover the missing index from  
> > the other shard or the replica, but it sticks at this point until I  
> > manually delete what's left of the index. Was this due to the split-brain  
> > issue or is this just a limitation of ES? Is there a way to recover the  
> > missing index from the replica? How do I find the replicas?
> > 
> > ...Thanks,  
> > ...Ken

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 21, 2012, 4:03pm UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/13 "2012-02-21T16:03:37Z")

</div>

Maybe we can start with logs from the time that you first restarted the cluster?

On Tuesday, February 21, 2012 at 5:30 PM, Kenneth Loafman wrote:

> I'd like to understand what could have caused this loss of shards. What data do you need in order to track it down?
> 
> On Mon, Feb 20, 2012 at 9:34 AM, Kenneth Loafman \<[kenneth@loafman.com](mailto:kenneth@loafman.com) ([mailto:kenneth@loafman.com](mailto:kenneth@loafman.com))\> wrote:
> 
> > But data will be removed if a shard relocates, right? So what happens in a split-brain situation when the brain is put back together?
> > 
> > On Mon, Feb 20, 2012 at 6:56 AM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > 
> > > Thats strange… . In 0.19, we have a better storage system for local gateway, where state is stored within each index/shard, instead of globally on the node level. I am still not sure what caused the data to be removed though…, elasticsearch does not remove data on its own unless instructed to.
> > > 
> > > On Saturday, February 18, 2012 at 12:58 AM, Kenneth Loafman wrote:
> > > 
> > > > No data was deleted from the filesystem. What I found when I looked was an empty directory where the shard should have been.
> > > > 
> > > > On Fri, Feb 17, 2012 at 4:35 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > > > 
> > > > > What I meant by data deleted is that some data was deleted from the file system by any chance? I suggest you start to delete the problematic indexes that hold the problematic shards when the cluster is up.
> > > > > 
> > > > > On Friday, February 17, 2012 at 11:32 PM, Kenneth Loafman wrote:
> > > > > 
> > > > > > Any ideas? I've shut down again, forced fsck on next boot, rebooted and restarted. No real problems found, so we can rule that out. The logs are too big to gist. What would you need from them if I could find it?
> > > > > > 
> > > > > > On Fri, Feb 17, 2012 at 1:55 PM, Kenneth Loafman \<[kenneth@loafman.com](mailto:kenneth@loafman.com) ([mailto:kenneth@loafman.com](mailto:kenneth@loafman.com))\> wrote:
> > > > > > 
> > > > > > > Hmm, something else is going on. Two of the indexes that were OK originally are now showing IndexShardMissingException.
> > > > > > > 
> > > > > > > ES is indeed hungry!
> > > > > > > 
> > > > > > > On Fri, Feb 17, 2012 at 12:58 PM, Kenneth Loafman \<[kenneth@loafman.com](mailto:kenneth@loafman.com) ([mailto:kenneth@loafman.com](mailto:kenneth@loafman.com))\> wrote:
> > > > > > > 
> > > > > > > > Nothing was deleted manually or through curl. Is this recoverable at all? What happened? Was this because of the split-cluster condition?
> > > > > > > > 
> > > > > > > > On Fri, Feb 17, 2012 at 12:54 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > > > > > > > 
> > > > > > > > > This means that the shard was supposed to exist on that node, but it can't be found, are you sure nothing was deleted?
> > > > > > > > > 
> > > > > > > > > On Friday, February 17, 2012 at 8:49 PM, Kenneth Loafman wrote:
> > > > > > > > > 
> > > > > > > > > > Yes, a whole bunch of messages repeating like this:
> > > > > > > > > > 
> > > > > > > > > > [2012-02-17 18:48:00,586][WARN][cluster.action.shard] [Blindspot] received shard failed for [co0198ca0694][1], node[o9nkBCsISKGt7P6acyshHQ], [P], s[INITIALIZING], reason [Failed to start shard, message [IndexShardGatewayRecoveryException[[co0198ca0694][1] shard allocated for local recovery (post api), should exists, but doesn't]]]
> > > > > > > > > > 
> > > > > > > > > > On Fri, Feb 17, 2012 at 12:45 PM, Shay Banon \<[kimchy@gmail.com](mailto:kimchy@gmail.com) ([mailto:kimchy@gmail.com](mailto:kimchy@gmail.com))\> wrote:
> > > > > > > > > > 
> > > > > > > > > > > Do you see anything in the logs? It seems like there are 8 initializing shards.
> > > > > > > > > > > 
> > > > > > > > > > > On Friday, February 17, 2012 at 8:34 PM, Kenneth Loafman wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > Hi,
> > > > > > > > > > > > 
> > > > > > > > > > > > We're on ES 18.6 driving a 4-node cluster on RackSpace.
> > > > > > > > > > > > 
> > > > > > > > > > > > Last night we had a nework outage on two of the nodes and our 4-node cluster morphed into 2 2-node clusters. I think that's what happened anyway. We shut all 4 nodes down cleanly, brought them up one at a time and the cluster reformed into one, however, it's sticking on getting out of red.
> > > > > > > > > > > > 
> > > > > > > > > > > > {  
> > > > > > > > > > > > "cluster\_name" : "elasticsearch",  
> > > > > > > > > > > > "status" : "red",  
> > > > > > > > > > > > "timed\_out" : false,  
> > > > > > > > > > > > "number\_of\_nodes" : 4,  
> > > > > > > > > > > > "number\_of\_data\_nodes" : 4,  
> > > > > > > > > > > > "active\_primary\_shards" : 234,  
> > > > > > > > > > > > "active\_shards" : 426,  
> > > > > > > > > > > > "relocating\_shards" : 0,  
> > > > > > > > > > > > "initializing\_shards" : 8,  
> > > > > > > > > > > > "unassigned\_shards" : 126  
> > > > > > > > > > > > }
> > > > > > > > > > > > 
> > > > > > > > > > > > It will stay in this mode for a long time and the '\_status' will show a bunch of
> > > > > > > > > > > > 
> > > > > > > > > > > > ```
> > > > > > > > > > > > "failures" : [ {  
> > > > > > > > > > > > "index" : "co0181ca0711",
> > > > > > > > > > > > "shard" : 1,
> > > > > > > > > > > > "reason" : "BroadcastShardOperationFailedException[[co0181ca0711][1] ]; nested: RemoteTransportException[[Whiteout][inet[/10.177.166.64:9300]][indices/status/shard]]; nested: IndexMissingException[[co0181ca0711] missing]; "
> > > > > > > > > > > > }, {
> > > > > > > > > > > > 
> > > > > > > > > > > > ```
> > > > > > > > > > > > 
> > > > > > > > > > > > which seem to come and go, but not get initialized. With 2 shards and 1 replica, it seems that it should be able to recover the missing index from the other shard or the replica, but it sticks at this point until I manually delete what's left of the index. Was this due to the split-brain issue or is this just a limitation of ES? Is there a way to recover the missing index from the replica? How do I find the replicas?
> > > > > > > > > > > > 
> > > > > > > > > > > > ...Thanks,  
> > > > > > > > > > > > ...Ken

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:38am UTC](https://discuss.elastic.co/t/es-ate-my-shards-indexes/6738/14 "2017-07-06T03:38:42Z")

</div>


