# Index inconsistency on 0.16.2 after a network partition

**URL:** <https://discuss.elastic.co/t/index-inconsistency-on-0-16-2-after-a-network-partition/5334>\
**Category:** Elasticsearch\
**Created:** [September 8, 2011, 4:42pm UTC](https://discuss.elastic.co/t/index-inconsistency-on-0-16-2-after-a-network-partition/5334 "2011-09-08T16:42:56Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [September 8, 2011, 4:42pm UTC](https://discuss.elastic.co/t/index-inconsistency-on-0-16-2-after-a-network-partition/5334/1 "2011-09-08T16:42:56Z")

</div>

Hey,  
We've been running 16.2 for a while with no problems (other than the  
memory leak that is fixed in 0.16.5 and we're in the process of moving  
to that release).

However, a couple of nights ago we an incident with our core router  
that caused a network partition. We have four nodes and it appears  
that node 1 was disconnected from nodes 2,3,4.

The 2,3,4 cluster went into a yellow state and began recovering data  
from each other and the single node cluster(1) went to the red state.  
I am not sure if anything corruption occurred at this point.

In order to rectify, I took the following steps:

- Shutdown the single node cluster
- Started back up the node that had been orphaned and it rejoined the  
main 2,3,4 cluster
- Replayed transactions that had occurred while the cluster was split

This should have restored the cluster back to a correct state, but  
something appears to have gone wrong when node 1 was disconnected or  
rejoined the cluster. Three of the larger indexes ended up in an  
inconsistent state. Depending which node was queried, different counts  
would come back. These were the indexes effected:

idol-ft\_20110513220131  
idol-nab\_20110513220132  
idol-reports1\_20110513220132

For example, when here are the counts I get when I hit idol-  
ft\_20110513220131 from all 4 nodes:  
1 - 1154320  
2 - 1079486  
3 - 1080016  
4 - 1228060 - This is the correct count

Log files, cluster state and config file can be viewed here:

> **[Dropbox - Disabled link](https://www.dropbox.com/s/kwrcofdfvrbpomb/ESLogs)**
>
> Dropbox is a free service that lets you bring your photos, docs, and videos anywhere and share them easily. Never email yourself a file again!

In order to address, I needed to rebuild the indexes from our backend  
storage and swapped aliases to make the new indices live.

I still have the inconsistent indices available if there are any  
details you want from them.

I had done extensive testing around this scenario on 0.16.2 and had  
never reproduced this. This seems different than some of the index  
corruption issues that occurred with 0.14.2 that were fixed in 0.16,  
as the destruction that occurred in 0.14 was much more severe and  
resulted in indices getting completely wiped vs inconsistent with all  
of the data being available sometimes.

Please let me know what I can do to help on this. I'll lend whatever  
support necessary to help address.

Thanks,  
Paul

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [September 8, 2011, 5:42pm UTC](https://discuss.elastic.co/t/index-inconsistency-on-0-16-2-after-a-network-partition/5334/2 "2011-09-08T17:42:00Z")

</div>

Hey,

Once you would have restarted node1, then it should have sync'ed with the  
other nodes and provide consistent view of the data. There were  
some enhancements done to this process in 0.17, but its strange that you ran  
into it... .

Btw, in 0.17, you can use the discovery.zen.minimum\_master\_nodes setting  
which would have solved this case of network partitioning (for example, if  
you have it set at 2).

-shay.banon

On Thu, Sep 8, 2011 at 7:42 PM, ppearcy [ppearcy@gmail.com](mailto:ppearcy@gmail.com) wrote:

> Hey,  
> We've been running 16.2 for a while with no problems (other than the  
> memory leak that is fixed in 0.16.5 and we're in the process of moving  
> to that release).
> 
> However, a couple of nights ago we an incident with our core router  
> that caused a network partition. We have four nodes and it appears  
> that node 1 was disconnected from nodes 2,3,4.
> 
> The 2,3,4 cluster went into a yellow state and began recovering data  
> from each other and the single node cluster(1) went to the red state.  
> I am not sure if anything corruption occurred at this point.
> 
> In order to rectify, I took the following steps:
> 
> - Shutdown the single node cluster
> - Started back up the node that had been orphaned and it rejoined the  
> main 2,3,4 cluster
> - Replayed transactions that had occurred while the cluster was split
> 
> This should have restored the cluster back to a correct state, but  
> something appears to have gone wrong when node 1 was disconnected or  
> rejoined the cluster. Three of the larger indexes ended up in an  
> inconsistent state. Depending which node was queried, different counts  
> would come back. These were the indexes effected:
> 
> idol-ft\_20110513220131  
> idol-nab\_20110513220132  
> idol-reports1\_20110513220132
> 
> For example, when here are the counts I get when I hit idol-  
> ft\_20110513220131 from all 4 nodes:  
> 1 - 1154320  
> 2 - 1079486  
> 3 - 1080016  
> 4 - 1228060 - This is the correct count
> 
> Log files, cluster state and config file can be viewed here:  
> [Dropbox - Link Disabled - Simplify your life](http://db.tt/EXenpZW)
> 
> In order to address, I needed to rebuild the indexes from our backend  
> storage and swapped aliases to make the new indices live.
> 
> I still have the inconsistent indices available if there are any  
> details you want from them.
> 
> I had done extensive testing around this scenario on 0.16.2 and had  
> never reproduced this. This seems different than some of the index  
> corruption issues that occurred with 0.14.2 that were fixed in 0.16,  
> as the destruction that occurred in 0.14 was much more severe and  
> resulted in indices getting completely wiped vs inconsistent with all  
> of the data being available sometimes.
> 
> Please let me know what I can do to help on this. I'll lend whatever  
> support necessary to help address.
> 
> Thanks,  
> Paul

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [September 12, 2011, 8:18pm UTC](https://discuss.elastic.co/t/index-inconsistency-on-0-16-2-after-a-network-partition/5334/3 "2011-09-12T20:18:16Z")

</div>

Hey,  
Thanks for the details. A couple of follow up questions, if you have  
a few more moments:

- I am thinking about increasing the discovery.zen.ping\_timeout to  
pretty large number. Something between 5-10 minutes. My understanding  
and what I've seen from testing, leads me to believe this should cause  
any new operations to block until the timeout is reached or the  
network is restored. If a node is brought down cleanly, this shouldn't  
come into play. Is that correct and are there any negatives I haven't  
mentioned?

- I am not sure that the discovery.zen.minimum\_master\_nodes would have  
made a difference in this case. I had a split with 1 vs 3 nodes and  
the inconsistency started when either they were split or when they  
were rejoined. It seems the split/rejoin would have occurred even with  
that setting.

- Is there an easy way to tell which replica of a shard is giving  
which counts? I believe this may be possible with Luke, but not very  
easy to do. If I had known which shards had gotten into a bad state, I  
should have been able to restore them by stopping the node with the  
bad shard and wiping the data for that index, forcing recovery from  
master.

Thanks for all the great work that goes into ES and I appreciate all  
the time spent answering questions like this.

Best Regards,  
Paul

On Sep 8, 11:42 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Hey,
> 
> Once you would have restarted node1, then it should have sync'ed with the  
> other nodes and provide consistent view of the data. There were  
> some enhancements done to this process in 0.17, but its strange that you ran  
> into it... .
> 
> Btw, in 0.17, you can use the discovery.zen.minimum\_master\_nodes setting  
> which would have solved this case of network partitioning (for example, if  
> you have it set at 2).
> 
> -shay.banon
> 
> On Thu, Sep 8, 2011 at 7:42 PM, ppearcy [ppea...@gmail.com](mailto:ppea...@gmail.com) wrote:
> 
> > Hey,  
> > We've been running 16.2 for a while with no problems (other than the  
> > memory leak that is fixed in 0.16.5 and we're in the process of moving  
> > to that release).
> 
> > However, a couple of nights ago we an incident with our core router  
> > that caused a network partition. We have four nodes and it appears  
> > that node 1 was disconnected from nodes 2,3,4.
> 
> > The 2,3,4 cluster went into a yellow state and began recovering data  
> > from each other and the single node cluster(1) went to the red state.  
> > I am not sure if anything corruption occurred at this point.
> 
> > In order to rectify, I took the following steps:
> > 
> > - Shutdown the single node cluster
> > - Started back up the node that had been orphaned and it rejoined the  
> > main 2,3,4 cluster
> > - Replayed transactions that had occurred while the cluster was split
> 
> > This should have restored the cluster back to a correct state, but  
> > something appears to have gone wrong when node 1 was disconnected or  
> > rejoined the cluster. Three of the larger indexes ended up in an  
> > inconsistent state. Depending which node was queried, different counts  
> > would come back. These were the indexes effected:
> 
> > idol-ft\_20110513220131  
> > idol-nab\_20110513220132  
> > idol-reports1\_20110513220132
> 
> > For example, when here are the counts I get when I hit idol-  
> > ft\_20110513220131 from all 4 nodes:  
> > 1 - 1154320  
> > 2 - 1079486  
> > 3 - 1080016  
> > 4 - 1228060 - This is the correct count
> 
> > Log files, cluster state and config file can be viewed here:  
> > [Dropbox - Link Disabled - Simplify your life](http://db.tt/EXenpZW)
> 
> > In order to address, I needed to rebuild the indexes from our backend  
> > storage and swapped aliases to make the new indices live.
> 
> > I still have the inconsistent indices available if there are any  
> > details you want from them.
> 
> > I had done extensive testing around this scenario on 0.16.2 and had  
> > never reproduced this. This seems different than some of the index  
> > corruption issues that occurred with 0.14.2 that were fixed in 0.16,  
> > as the destruction that occurred in 0.14 was much more severe and  
> > resulted in indices getting completely wiped vs inconsistent with all  
> > of the data being available sometimes.
> 
> > Please let me know what I can do to help on this. I'll lend whatever  
> > support necessary to help address.
> 
> > Thanks,  
> > Paul

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [September 12, 2011, 10:52pm UTC](https://discuss.elastic.co/t/index-inconsistency-on-0-16-2-after-a-network-partition/5334/4 "2011-09-12T22:52:17Z")

</div>

On Mon, Sep 12, 2011 at 11:18 PM, ppearcy [ppearcy@gmail.com](mailto:ppearcy@gmail.com) wrote:

> Hey,  
> Thanks for the details. A couple of follow up questions, if you have  
> a few more moments:
> 
> - I am thinking about increasing the discovery.zen.ping\_timeout to  
> pretty large number. Something between 5-10 minutes. My understanding  
> and what I've seen from testing, leads me to believe this should cause  
> any new operations to block until the timeout is reached or the  
> network is restored. If a node is brought down cleanly, this shouldn't  
> come into play. Is that correct and are there any negatives I haven't  
> mentioned?

You mean the fault detection timeout between nodes? Those settings are  
here:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/discovery/zen.html).  
When there is a clean shutdown of a node, or actually socket failure, then  
disconnection will be detected extremely fast. When sending a request and  
not getting a response, then those fd timeouts gets into play.

I do not recommend setting those to too high value, since using the  
minimum\_master\_nodes setting will be affected by it (the nodes will not be  
identified as disconnected).

> - I am not sure that the discovery.zen.minimum\_master\_nodes would have  
> made a difference in this case. I had a split with 1 vs 3 nodes and  
> the inconsistency started when either they were split or when they  
> were rejoined. It seems the split/rejoin would have occurred even with  
> that setting.

It should have helped. What would have happened is that the 1 node  
disconnected would have gotten into a state where it would not have shards  
allocated on it, and it will go into a rejoin process automatically. Your  
logic is sound, as in, the fact that you shutdown the node and reconnected  
it should have resolved any potential conflicts, but I believe what you saw  
is solved in 0.17 (at least based on my tests).

> - Is there an easy way to tell which replica of a shard is giving  
> which counts? I believe this may be possible with Luke, but not very  
> easy to do. If I had known which shards had gotten into a bad state, I  
> should have been able to restore them by stopping the node with the  
> bad shard and wiping the data for that index, forcing recovery from  
> master.

Yes, when you execute a search and set the explain flag to true, it will  
also return the shard and node each hit comes back from.

> Thanks for all the great work that goes into ES and I appreciate all  
> the time spent answering questions like this.
> 
> Best Regards,  
> Paul
> 
> On Sep 8, 11:42 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > Hey,
> > 
> > Once you would have restarted node1, then it should have sync'ed with  
> > the  
> > other nodes and provide consistent view of the data. There were  
> > some enhancements done to this process in 0.17, but its strange that you  
> > ran  
> > into it... .
> > 
> > Btw, in 0.17, you can use the discovery.zen.minimum\_master\_nodes  
> > setting  
> > which would have solved this case of network partitioning (for example,  
> > if  
> > you have it set at 2).
> > 
> > -shay.banon
> > 
> > On Thu, Sep 8, 2011 at 7:42 PM, ppearcy [ppea...@gmail.com](mailto:ppea...@gmail.com) wrote:
> > 
> > > Hey,  
> > > We've been running 16.2 for a while with no problems (other than the  
> > > memory leak that is fixed in 0.16.5 and we're in the process of moving  
> > > to that release).
> > 
> > > However, a couple of nights ago we an incident with our core router  
> > > that caused a network partition. We have four nodes and it appears  
> > > that node 1 was disconnected from nodes 2,3,4.
> > 
> > > The 2,3,4 cluster went into a yellow state and began recovering data  
> > > from each other and the single node cluster(1) went to the red state.  
> > > I am not sure if anything corruption occurred at this point.
> > 
> > > In order to rectify, I took the following steps:
> > > 
> > > - Shutdown the single node cluster
> > > - Started back up the node that had been orphaned and it rejoined the  
> > > main 2,3,4 cluster
> > > - Replayed transactions that had occurred while the cluster was split
> > 
> > > This should have restored the cluster back to a correct state, but  
> > > something appears to have gone wrong when node 1 was disconnected or  
> > > rejoined the cluster. Three of the larger indexes ended up in an  
> > > inconsistent state. Depending which node was queried, different counts  
> > > would come back. These were the indexes effected:
> > 
> > > idol-ft\_20110513220131  
> > > idol-nab\_20110513220132  
> > > idol-reports1\_20110513220132
> > 
> > > For example, when here are the counts I get when I hit idol-  
> > > ft\_20110513220131 from all 4 nodes:  
> > > 1 - 1154320  
> > > 2 - 1079486  
> > > 3 - 1080016  
> > > 4 - 1228060 - This is the correct count
> > 
> > > Log files, cluster state and config file can be viewed here:  
> > > [Dropbox - Link Disabled - Simplify your life](http://db.tt/EXenpZW)
> > 
> > > In order to address, I needed to rebuild the indexes from our backend  
> > > storage and swapped aliases to make the new indices live.
> > 
> > > I still have the inconsistent indices available if there are any  
> > > details you want from them.
> > 
> > > I had done extensive testing around this scenario on 0.16.2 and had  
> > > never reproduced this. This seems different than some of the index  
> > > corruption issues that occurred with 0.14.2 that were fixed in 0.16,  
> > > as the destruction that occurred in 0.14 was much more severe and  
> > > resulted in indices getting completely wiped vs inconsistent with all  
> > > of the data being available sometimes.
> > 
> > > Please let me know what I can do to help on this. I'll lend whatever  
> > > support necessary to help address.
> > 
> > > Thanks,  
> > > Paul

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:54am UTC](https://discuss.elastic.co/t/index-inconsistency-on-0-16-2-after-a-network-partition/5334/5 "2017-07-06T03:54:39Z")

</div>


