# Zen ping timeout causes nodes to lose master permanently

**URL:** <https://discuss.elastic.co/t/zen-ping-timeout-causes-nodes-to-lose-master-permanently/3246>\
**Category:** Elasticsearch\
**Created:** [August 23, 2010, 11:03pm UTC](https://discuss.elastic.co/t/zen-ping-timeout-causes-nodes-to-lose-master-permanently/3246 "2010-08-23T23:03:25Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Grant\_Rodgers](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/grant_rodgers/32/3204_2.png) [@Grant\_Rodgers](https://discuss.elastic.co/u/Grant_Rodgers)\
**Post date:** [August 23, 2010, 11:03pm UTC](https://discuss.elastic.co/t/zen-ping-timeout-causes-nodes-to-lose-master-permanently/3246/1 "2010-08-23T23:03:25Z")

</div>

Summary: I am having problems with nodes losing the master from their  
node list. This is causing health to be reported differently on  
different nodes, and it appears that shard replicas are getting into  
divergent states when clients write to different nodes.

Long version:

We have a backup process on one node that runs nightly. Apparently  
that causes GC pressure making gc collections take longer than normal:

/var/log/elasticsearch/production.log.2010-08-21:[11:16:23,000][WARN]  
[monitor.jvm] [Nocturne] Long GC collection occurred,  
took [42.5s], breached threshold [10s]  
/var/log/elasticsearch/production.log.2010-08-21:[11:37:22,382][WARN]  
[monitor.jvm] [Nocturne] Long GC collection occurred,  
took [55.1s], breached threshold [10s]

Which is fine, our load is very low at that time. The problem is that  
this machine happens to be the master node. On our 2 other nodes we  
get these log lines:

/var/log/elasticsearch/production.log.2010-08-21:[11:16:20,263][WARN]  
[transport] [Devil-Slayer] Transport response handler  
timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
[b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]  
/var/log/elasticsearch/production.log.2010-08-21:[11:37:22,371][WARN]  
[transport] [Devil-Slayer] Transport response handler  
timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
[b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]

/var/log/elasticsearch/production.log.2010-08-21:[11:16:21,757][WARN]  
[transport] [Noh-Varr] Transport response handler  
timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
[b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]

/var/log/elasticsearch/production.log.2010-08-21:[11:37:22,376][WARN]  
[transport] [Noh-Varr] Transport response handler  
timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
[b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]

So the ping times out while the master has high load. But now the non-  
master nodes are in a weird state and don't seem to recover from it.  
They don't list the master in their nodes list, and the cluster health  
reports only 2 active nodes and only knows about the shards on those  
nodes. The master node still knows about all 3 nodes, and reports  
health as if all shards from all 3 nodes are available.

Is elasticsearch built to handle ping timeouts like this? I'm bringing  
it up here because I suspect it is, and this might be a bug.

Also, after running in this state for a day or two, some of the shards  
had document counts that are different by a few dozen, presumably  
because they are not replicating changes to the shards they aren't  
aware of. We were able to fix this issue by doing a rolling restart of  
the cluster, no downtime required. That was pretty cool.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 23, 2010, 11:13pm UTC](https://discuss.elastic.co/t/zen-ping-timeout-causes-nodes-to-lose-master-permanently/3246/2 "2010-08-23T23:13:24Z")

</div>

Hi,

Yes, what happened here is that the cluster got into a split brain.  
Currently, the way elasticsearch handles a split brain is by not trying to  
join these two cluster back into a single cluster. The main problem with  
split brain is when you still have clients working against these two  
different clusters.

I plan to be able to handle better split brain scenarios. One option is  
to have the two different clusters try and join once such a scenario  
happens. Another option is to have the smaller cluster kill itself once  
something like this happens.

Resolving this is currently a manual process that you did, which is  
basically a restart of the smaller cluster.

-shay.banon

On Tue, Aug 24, 2010 at 2:03 AM, Grant Rodgers [grantr@gmail.com](mailto:grantr@gmail.com) wrote:

> Summary: I am having problems with nodes losing the master from their  
> node list. This is causing health to be reported differently on  
> different nodes, and it appears that shard replicas are getting into  
> divergent states when clients write to different nodes.
> 
> Long version:
> 
> We have a backup process on one node that runs nightly. Apparently  
> that causes GC pressure making gc collections take longer than normal:
> 
> /var/log/elasticsearch/production.log.2010-08-21:[11:16:23,000][WARN]  
> [monitor.jvm] [Nocturne] Long GC collection occurred,  
> took [42.5s], breached threshold [10s]  
> /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,382][WARN]  
> [monitor.jvm] [Nocturne] Long GC collection occurred,  
> took [55.1s], breached threshold [10s]
> 
> Which is fine, our load is very low at that time. The problem is that  
> this machine happens to be the master node. On our 2 other nodes we  
> get these log lines:
> 
> /var/log/elasticsearch/production.log.2010-08-21:[11:16:20,263][WARN]  
> [transport] [Devil-Slayer] Transport response handler  
> timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]  
> /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,371][WARN]  
> [transport] [Devil-Slayer] Transport response handler  
> timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> 
> /var/log/elasticsearch/production.log.2010-08-21:[11:16:21,757][WARN]  
> [transport] [Noh-Varr] Transport response handler  
> timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> 
> /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,376][WARN]  
> [transport] [Noh-Varr] Transport response handler  
> timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> 
> So the ping times out while the master has high load. But now the non-  
> master nodes are in a weird state and don't seem to recover from it.  
> They don't list the master in their nodes list, and the cluster health  
> reports only 2 active nodes and only knows about the shards on those  
> nodes. The master node still knows about all 3 nodes, and reports  
> health as if all shards from all 3 nodes are available.
> 
> Is elasticsearch built to handle ping timeouts like this? I'm bringing  
> it up here because I suspect it is, and this might be a bug.
> 
> Also, after running in this state for a day or two, some of the shards  
> had document counts that are different by a few dozen, presumably  
> because they are not replicating changes to the shards they aren't  
> aware of. We were able to fix this issue by doing a rolling restart of  
> the cluster, no downtime required. That was pretty cool.

---

<div class="post-metadata">

**Author:** ![Grant\_Rodgers](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/grant_rodgers/32/3204_2.png) [@Grant\_Rodgers](https://discuss.elastic.co/u/Grant_Rodgers)\
**Post date:** [August 23, 2010, 11:49pm UTC](https://discuss.elastic.co/t/zen-ping-timeout-causes-nodes-to-lose-master-permanently/3246/3 "2010-08-23T23:49:43Z")

</div>

Right, I guess network partitions are pretty hard to detect and repair  
in clustered systems. We'll just do the rolling restart if this  
happens again, it was pretty painless.

On Aug 23, 4:13 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Hi,
> 
> Yes, what happened here is that the cluster got into a split brain.  
> Currently, the way elasticsearch handles a split brain is by not trying to  
> join these two cluster back into a single cluster. The main problem with  
> split brain is when you still have clients working against these two  
> different clusters.
> 
> I plan to be able to handle better split brain scenarios. One option is  
> to have the two different clusters try and join once such a scenario  
> happens. Another option is to have the smaller cluster kill itself once  
> something like this happens.
> 
> Resolving this is currently a manual process that you did, which is  
> basically a restart of the smaller cluster.
> 
> -shay.banon
> 
> On Tue, Aug 24, 2010 at 2:03 AM, Grant Rodgers [gra...@gmail.com](mailto:gra...@gmail.com) wrote:
> 
> > Summary: I am having problems with nodes losing the master from their  
> > node list. This is causing health to be reported differently on  
> > different nodes, and it appears that shard replicas are getting into  
> > divergent states when clients write to different nodes.
> 
> > Long version:
> 
> > We have a backup process on one node that runs nightly. Apparently  
> > that causes GC pressure making gc collections take longer than normal:
> 
> > /var/log/elasticsearch/production.log.2010-08-21:[11:16:23,000][WARN]  
> > [monitor.jvm] [Nocturne] Long GC collection occurred,  
> > took [42.5s], breached threshold [10s]  
> > /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,382][WARN]  
> > [monitor.jvm] [Nocturne] Long GC collection occurred,  
> > took [55.1s], breached threshold [10s]
> 
> > Which is fine, our load is very low at that time. The problem is that  
> > this machine happens to be the master node. On our 2 other nodes we  
> > get these log lines:
> 
> > /var/log/elasticsearch/production.log.2010-08-21:[11:16:20,263][WARN]  
> > [transport] [Devil-Slayer] Transport response handler  
> > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]  
> > /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,371][WARN]  
> > [transport] [Devil-Slayer] Transport response handler  
> > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> 
> > /var/log/elasticsearch/production.log.2010-08-21:[11:16:21,757][WARN]  
> > [transport] [Noh-Varr] Transport response handler  
> > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> 
> > /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,376][WARN]  
> > [transport] [Noh-Varr] Transport response handler  
> > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> 
> > So the ping times out while the master has high load. But now the non-  
> > master nodes are in a weird state and don't seem to recover from it.  
> > They don't list the master in their nodes list, and the cluster health  
> > reports only 2 active nodes and only knows about the shards on those  
> > nodes. The master node still knows about all 3 nodes, and reports  
> > health as if all shards from all 3 nodes are available.
> 
> > Is elasticsearch built to handle ping timeouts like this? I'm bringing  
> > it up here because I suspect it is, and this might be a bug.
> 
> > Also, after running in this state for a day or two, some of the shards  
> > had document counts that are different by a few dozen, presumably  
> > because they are not replicating changes to the shards they aren't  
> > aware of. We were able to fix this issue by doing a rolling restart of  
> > the cluster, no downtime required. That was pretty cool.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 24, 2010, 12:21am UTC](https://discuss.elastic.co/t/zen-ping-timeout-causes-nodes-to-lose-master-permanently/3246/4 "2010-08-24T00:21:34Z")

</div>

There are different ways to handle network partitioning, none of them is  
really good. The most popular solution, which is the dynamo model (though in  
it, eventual consistent should be called: eventual consistent with a chance  
of loosing data) is not really applicable to how elasticsearch works (or  
search engines for that matter). It can be implemented, it will just take so  
many resources out of the system in order to try and implement it, that it  
will render it useless.

There are other ways. For example, requiring a quorum of a known cluster in  
order to accepts "writes", and then automatically rejoining a cluster  
network partitioning has been resolved.

If I remember correctly, you are running on ec2, so I suggest, in any case,  
increasing the failure detector timeouts (currently implementing simple  
polling FD, other implementations will come later):

discovery.zen.fd.ping\_timeout: 1m  
discovery.zen.fd.ping\_retries: 5

This will configure a 5 minutes timeout.

-shay.banon

On Tue, Aug 24, 2010 at 2:49 AM, Grant Rodgers [grantr@gmail.com](mailto:grantr@gmail.com) wrote:

> Right, I guess network partitions are pretty hard to detect and repair  
> in clustered systems. We'll just do the rolling restart if this  
> happens again, it was pretty painless.
> 
> On Aug 23, 4:13 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Hi,
> > 
> > Yes, what happened here is that the cluster got into a split brain.  
> > Currently, the way elasticsearch handles a split brain is by not trying  
> > to  
> > join these two cluster back into a single cluster. The main problem with  
> > split brain is when you still have clients working against these two  
> > different clusters.
> > 
> > I plan to be able to handle better split brain scenarios. One option  
> > is  
> > to have the two different clusters try and join once such a scenario  
> > happens. Another option is to have the smaller cluster kill itself once  
> > something like this happens.
> > 
> > Resolving this is currently a manual process that you did, which is  
> > basically a restart of the smaller cluster.
> > 
> > -shay.banon
> > 
> > On Tue, Aug 24, 2010 at 2:03 AM, Grant Rodgers [gra...@gmail.com](mailto:gra...@gmail.com) wrote:
> > 
> > > Summary: I am having problems with nodes losing the master from their  
> > > node list. This is causing health to be reported differently on  
> > > different nodes, and it appears that shard replicas are getting into  
> > > divergent states when clients write to different nodes.
> > 
> > > Long version:
> > 
> > > We have a backup process on one node that runs nightly. Apparently  
> > > that causes GC pressure making gc collections take longer than normal:
> > 
> > > /var/log/elasticsearch/production.log.2010-08-21:[11:16:23,000][WARN]  
> > > [monitor.jvm] [Nocturne] Long GC collection occurred,  
> > > took [42.5s], breached threshold [10s]  
> > > /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,382][WARN]  
> > > [monitor.jvm] [Nocturne] Long GC collection occurred,  
> > > took [55.1s], breached threshold [10s]
> > 
> > > Which is fine, our load is very low at that time. The problem is that  
> > > this machine happens to be the master node. On our 2 other nodes we  
> > > get these log lines:
> > 
> > > /var/log/elasticsearch/production.log.2010-08-21:[11:16:20,263][WARN]  
> > > [transport] [Devil-Slayer] Transport response handler  
> > > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]  
> > > /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,371][WARN]  
> > > [transport] [Devil-Slayer] Transport response handler  
> > > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> > 
> > > /var/log/elasticsearch/production.log.2010-08-21:[11:16:21,757][WARN]  
> > > [transport] [Noh-Varr] Transport response handler  
> > > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> > 
> > > /var/log/elasticsearch/production.log.2010-08-21:[11:37:22,376][WARN]  
> > > [transport] [Noh-Varr] Transport response handler  
> > > timed out, action [discovery/zen/fd/masterPing], node [[Nocturne]  
> > > [b2d9e7a8-6f6f-497d-903c-b564116e85d5][inet[/10.102.43.160:9300]]]
> > 
> > > So the ping times out while the master has high load. But now the non-  
> > > master nodes are in a weird state and don't seem to recover from it.  
> > > They don't list the master in their nodes list, and the cluster health  
> > > reports only 2 active nodes and only knows about the shards on those  
> > > nodes. The master node still knows about all 3 nodes, and reports  
> > > health as if all shards from all 3 nodes are available.
> > 
> > > Is elasticsearch built to handle ping timeouts like this? I'm bringing  
> > > it up here because I suspect it is, and this might be a bug.
> > 
> > > Also, after running in this state for a day or two, some of the shards  
> > > had document counts that are different by a few dozen, presumably  
> > > because they are not replicating changes to the shards they aren't  
> > > aware of. We were able to fix this issue by doing a rolling restart of  
> > > the cluster, no downtime required. That was pretty cool.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:20am UTC](https://discuss.elastic.co/t/zen-ping-timeout-causes-nodes-to-lose-master-permanently/3246/5 "2017-07-06T04:20:25Z")

</div>


