# Master\_left and no other node elected to become master

**URL:** <https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577>\
**Category:** Elasticsearch\
**Created:** [November 16, 2010, 5:38pm UTC](https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577 "2010-11-16T17:38:56Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [November 16, 2010, 5:38pm UTC](https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577/1 "2010-11-16T17:38:56Z")

</div>

Hey,  
Running master from yesterday morning. Have never hit this issue  
before and was previously running master from a week before that. I  
have two data nodes running:  
dm-adsearchd102-ElasticSearch  
dm-adsearchd103-ElasticSearch

I shutdown the d103 node via the service stop command and saw this in  
the logs for the d102 node:

[2010-11-16 16:25:17,764][INFO][discovery.zen] [DM-  
ADSEARCHD102.dev.local-ElasticSearch] master\_left [[dm-  
adsearchd103.dev.local-ElasticSearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
10.2.20.164:9300]]], reason [shut\_down]  
[2010-11-16 16:25:17,775][INFO][cluster.service] [DM-  
ADSEARCHD102.dev.local-ElasticSearch] master {new [DM-  
ADSEARCHD102.dev.local-ElasticSearch][7yu0RHxETUKROtrrGxBhIw][inet[/  
10.2.20.160:9300]], previous [dm-adsearchd103.dev.local-ElasticSearch]  
[l3Q3CX84SpGfcq-WzPDaUg][inet[/10.2.20.164:9300]]}, removed {[dm-  
adsearchd103.dev.local-ElasticSearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
10.2.20.164:9300]],}, reason: zen-disco-master\_failed ([dm-  
adsearchd103.dev.local-ElasticSearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
10.2.20.164:9300]])

This looks correct, however, we were running two applications with  
Node based clients and they logged nothing at the time of the initial  
disconnect, although, I did see a search fail for  
AlreadyClosedException. A few minutes later, I see:

2010-11-16 16:28:15,879 WARN \> [dm-adsearchd102.dev.local-  
essearcherserver] master\_left and no other node elected to become  
master, current nodes: [[WKSPPEARCYW7.wsod.local-ESIndexer]  
[L2QpS2IfTLqsZRf1OB76rw][inet[/10.2.123.16:9300]]{client=true,  
data=false}, [dm-adsearchd102.dev.local-ESIndexer]  
[qrSmIjycSA2pQ0OtVcdDng][inet[/10.2.20.160:9301]]{client=true,  
data=false}, [dm-adsearchd102.dev.local-essearcherserver]  
[7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
data=false}] (Log4jESLogger.internalWarn:87)(elasticsearch[dm-  
adsearchd102.dev.local-essearcherserver]clusterService#updateTask-  
pool-5-thread-1)

Afterwards, the node clients failed all search and index requests.  
Didn't notice immediately, 20 minutes later noticed, and then  
restarted both apps using node clients and they both came up cleanly.

I'm not sure if somehow there was a split brain condition, more like  
different nodes had a disagreement on who was in the cluster. I do see  
the node client attempting to re-establish communications with the one  
active node, but it is failing:

2010-11-16 16:28:19,746 WARN \> [dm-adsearchd102.dev.local-  
essearcherserver] failed to send ping to [[#zen\_unicast\_1#][inet[DM-  
ADSEARCHD102.dev.local/10.2.20.160:9300]]] (Log4jESLogger.internalWarn:  
91)(elasticsearch[dm-adsearchd102.dev.local-essearcherserver][tp]-  
pool-1-thread-23)  
org.elasticsearch.transport.ReceiveTimeoutTransportException: []  
[inet[DM-ADSEARCHD102.dev.local/10.2.20.160:9300]][discovery/zen/  
unicast]  
at org.elasticsearch.transport.TransportService  
$TimeoutTimerTask.run(TransportService.java:316)  
at org.elasticsearch.timer.TimerService$ThreadedTimerTask  
$1.run(TimerService.java:113)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.runTask(ThreadPoolExecutor.java:886)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.run(ThreadPoolExecutor.java:908)  
at java.lang.Thread.run(Thread.java:619)

On the one good data node, it seems to be ignoring join requests:  
[2010-11-16 16:31:42,880][WARN][discovery.zen] [DM-  
ADSEARCHD102.dev.local-ElasticSearch] received a join request for an  
existing node [[dm-adsearchd102.dev.local-essearcherserver]  
[7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
data=false}]

I can make all the logs available, if that would help.

In my app, I believe if I had closed the node client and re-opened, it  
probably would have addressed.

Please let me know if you need any other details.

Thanks,  
Paul

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 17, 2010, 9:01am UTC](https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577/2 "2010-11-17T09:01:21Z")

</div>

Hey Paul,

Can you mail me the logs? I will have a look.

-shay.banon

On Tue, Nov 16, 2010 at 7:38 PM, Paul [ppearcy@gmail.com](mailto:ppearcy@gmail.com) wrote:

> Hey,  
> Running master from yesterday morning. Have never hit this issue  
> before and was previously running master from a week before that. I  
> have two data nodes running:  
> dm-adsearchd102-Elasticsearch  
> dm-adsearchd103-Elasticsearch
> 
> I shutdown the d103 node via the service stop command and saw this in  
> the logs for the d102 node:
> 
> [2010-11-16 16:25:17,764][INFO][discovery.zen] [DM-  
> ADSEARCHD102.dev.local-Elasticsearch] master\_left [[dm-  
> adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> 10.2.20.164:9300]]], reason [shut\_down]  
> [2010-11-16 16:25:17,775][INFO][cluster.service] [DM-  
> ADSEARCHD102.dev.local-Elasticsearch] master {new [DM-  
> ADSEARCHD102.dev.local-Elasticsearch][7yu0RHxETUKROtrrGxBhIw][inet[/  
> 10.2.20.160:9300]], previous [dm-adsearchd103.dev.local-Elasticsearch]  
> [l3Q3CX84SpGfcq-WzPDaUg][inet[/10.2.20.164:9300]]}, removed {[dm-  
> adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> 10.2.20.164:9300]],}, reason: zen-disco-master\_failed ([dm-  
> adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> 10.2.20.164:9300]])
> 
> This looks correct, however, we were running two applications with  
> Node based clients and they logged nothing at the time of the initial  
> disconnect, although, I did see a search fail for  
> AlreadyClosedException. A few minutes later, I see:
> 
> 2010-11-16 16:28:15,879 WARN \> [dm-adsearchd102.dev.local-  
> essearcherserver] master\_left and no other node elected to become  
> master, current nodes: [[WKSPPEARCYW7.wsod.local-ESIndexer]  
> [L2QpS2IfTLqsZRf1OB76rw][inet[/10.2.123.16:9300]]{client=true,  
> data=false}, [dm-adsearchd102.dev.local-ESIndexer]  
> [qrSmIjycSA2pQ0OtVcdDng][inet[/10.2.20.160:9301]]{client=true,  
> data=false}, [dm-adsearchd102.dev.local-essearcherserver]  
> [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
> data=false}] (Log4jESLogger.internalWarn:87)(elasticsearch[dm-  
> adsearchd102.dev.local-essearcherserver]clusterService#updateTask-  
> pool-5-thread-1)
> 
> Afterwards, the node clients failed all search and index requests.  
> Didn't notice immediately, 20 minutes later noticed, and then  
> restarted both apps using node clients and they both came up cleanly.
> 
> I'm not sure if somehow there was a split brain condition, more like  
> different nodes had a disagreement on who was in the cluster. I do see  
> the node client attempting to re-establish communications with the one  
> active node, but it is failing:
> 
> 2010-11-16 16:28:19,746 WARN \> [dm-adsearchd102.dev.local-  
> essearcherserver] failed to send ping to [[#zen\_unicast\_1#][inet[DM-  
> ADSEARCHD102.dev.local/10.2.20.160:9300]]] (Log4jESLogger.internalWarn:  
> 91)(elasticsearch[dm-adsearchd102.dev.local-essearcherserver][tp]-  
> pool-1-thread-23)  
> org.elasticsearch.transport.ReceiveTimeoutTransportException:   
> [inet[DM-ADSEARCHD102.dev.local/10.2.20.160:9300]][discovery/zen/  
> unicast]  
> at org.elasticsearch.transport.TransportService  
> $TimeoutTimerTask.run(TransportService.java:316)  
> at org.elasticsearch.timer.TimerService$ThreadedTimerTask  
> $1.run(TimerService.java:113)  
> at java.util.concurrent.ThreadPoolExecutor  
> $Worker.runTask(ThreadPoolExecutor.java:886)  
> at java.util.concurrent.ThreadPoolExecutor  
> $Worker.run(ThreadPoolExecutor.java:908)  
> at java.lang.Thread.run(Thread.java:619)
> 
> On the one good data node, it seems to be ignoring join requests:  
> [2010-11-16 16:31:42,880][WARN][discovery.zen] [DM-  
> ADSEARCHD102.dev.local-Elasticsearch] received a join request for an  
> existing node [[dm-adsearchd102.dev.local-essearcherserver]  
> [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
> data=false}]
> 
> I can make all the logs available, if that would help.
> 
> In my app, I believe if I had closed the node client and re-opened, it  
> probably would have addressed.
> 
> Please let me know if you need any other details.
> 
> Thanks,  
> Paul

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 17, 2010, 9:36pm UTC](https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577/3 "2010-11-17T21:36:16Z")

</div>

I think I know why you got this failure, is there a chance that you used  
different 0.13 versions between the client and the server?

-shay.banon

On Wed, Nov 17, 2010 at 11:01 AM, Shay Banon  
[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Hey Paul,
> 
> Can you mail me the logs? I will have a look.
> 
> -shay.banon
> 
> On Tue, Nov 16, 2010 at 7:38 PM, Paul [ppearcy@gmail.com](mailto:ppearcy@gmail.com) wrote:
> 
> > Hey,  
> > Running master from yesterday morning. Have never hit this issue  
> > before and was previously running master from a week before that. I  
> > have two data nodes running:  
> > dm-adsearchd102-Elasticsearch  
> > dm-adsearchd103-Elasticsearch
> > 
> > I shutdown the d103 node via the service stop command and saw this in  
> > the logs for the d102 node:
> > 
> > [2010-11-16 16:25:17,764][INFO][discovery.zen] [DM-  
> > ADSEARCHD102.dev.local-Elasticsearch] master\_left [[dm-  
> > adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> > 10.2.20.164:9300]]], reason [shut\_down]  
> > [2010-11-16 16:25:17,775][INFO][cluster.service] [DM-  
> > ADSEARCHD102.dev.local-Elasticsearch] master {new [DM-  
> > ADSEARCHD102.dev.local-Elasticsearch][7yu0RHxETUKROtrrGxBhIw][inet[/  
> > 10.2.20.160:9300]], previous [dm-adsearchd103.dev.local-Elasticsearch]  
> > [l3Q3CX84SpGfcq-WzPDaUg][inet[/10.2.20.164:9300]]}, removed {[dm-  
> > adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> > 10.2.20.164:9300]],}, reason: zen-disco-master\_failed ([dm-  
> > adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> > 10.2.20.164:9300]])
> > 
> > This looks correct, however, we were running two applications with  
> > Node based clients and they logged nothing at the time of the initial  
> > disconnect, although, I did see a search fail for  
> > AlreadyClosedException. A few minutes later, I see:
> > 
> > 2010-11-16 16:28:15,879 WARN \> [dm-adsearchd102.dev.local-  
> > essearcherserver] master\_left and no other node elected to become  
> > master, current nodes: [[WKSPPEARCYW7.wsod.local-ESIndexer]  
> > [L2QpS2IfTLqsZRf1OB76rw][inet[/10.2.123.16:9300]]{client=true,  
> > data=false}, [dm-adsearchd102.dev.local-ESIndexer]  
> > [qrSmIjycSA2pQ0OtVcdDng][inet[/10.2.20.160:9301]]{client=true,  
> > data=false}, [dm-adsearchd102.dev.local-essearcherserver]  
> > [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
> > data=false}] (Log4jESLogger.internalWarn:87)(elasticsearch[dm-  
> > adsearchd102.dev.local-essearcherserver]clusterService#updateTask-  
> > pool-5-thread-1)
> > 
> > Afterwards, the node clients failed all search and index requests.  
> > Didn't notice immediately, 20 minutes later noticed, and then  
> > restarted both apps using node clients and they both came up cleanly.
> > 
> > I'm not sure if somehow there was a split brain condition, more like  
> > different nodes had a disagreement on who was in the cluster. I do see  
> > the node client attempting to re-establish communications with the one  
> > active node, but it is failing:
> > 
> > 2010-11-16 16:28:19,746 WARN \> [dm-adsearchd102.dev.local-  
> > essearcherserver] failed to send ping to [[#zen\_unicast\_1#][inet[DM-  
> > ADSEARCHD102.dev.local/10.2.20.160:9300]]] (Log4jESLogger.internalWarn:  
> > 91)(elasticsearch[dm-adsearchd102.dev.local-essearcherserver][tp]-  
> > pool-1-thread-23)  
> > org.elasticsearch.transport.ReceiveTimeoutTransportException:   
> > [inet[DM-ADSEARCHD102.dev.local/10.2.20.160:9300]][discovery/zen/  
> > unicast]  
> > at org.elasticsearch.transport.TransportService  
> > $TimeoutTimerTask.run(TransportService.java:316)  
> > at org.elasticsearch.timer.TimerService$ThreadedTimerTask  
> > $1.run(TimerService.java:113)  
> > at java.util.concurrent.ThreadPoolExecutor  
> > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > at java.util.concurrent.ThreadPoolExecutor  
> > $Worker.run(ThreadPoolExecutor.java:908)  
> > at java.lang.Thread.run(Thread.java:619)
> > 
> > On the one good data node, it seems to be ignoring join requests:  
> > [2010-11-16 16:31:42,880][WARN][discovery.zen] [DM-  
> > ADSEARCHD102.dev.local-Elasticsearch] received a join request for an  
> > existing node [[dm-adsearchd102.dev.local-essearcherserver]  
> > [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
> > data=false}]
> > 
> > I can make all the logs available, if that would help.
> > 
> > In my app, I believe if I had closed the node client and re-opened, it  
> > probably would have addressed.
> > 
> > Please let me know if you need any other details.
> > 
> > Thanks,  
> > Paul

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [November 19, 2010, 8:20pm UTC](https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577/4 "2010-11-19T20:20:46Z")

</div>

Hey, apologies for th delayed response. We've tried to be careful with  
mixed versions, but it is possible and we were upgrading at the time.

On 0.13.0 and haven't been able to reproduce. Will post logs if I can  
hit this again.

Thanks!

On Nov 17, 2:36 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> I think I know why you got this failure, is there a chance that you used  
> different 0.13 versions between the client and the server?
> 
> -shay.banon
> 
> On Wed, Nov 17, 2010 at 11:01 AM, Shay Banon  
> [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com)wrote:
> 
> > Hey Paul,
> 
> > Can you mail me the logs? I will have a look.
> 
> > -shay.banon
> 
> > On Tue, Nov 16, 2010 at 7:38 PM, Paul [ppea...@gmail.com](mailto:ppea...@gmail.com) wrote:
> 
> > > Hey,  
> > > Running master from yesterday morning. Have never hit this issue  
> > > before and was previously running master from a week before that. I  
> > > have two data nodes running:  
> > > dm-adsearchd102-Elasticsearch  
> > > dm-adsearchd103-Elasticsearch
> 
> > > I shutdown the d103 node via the service stop command and saw this in  
> > > the logs for the d102 node:
> 
> > > [2010-11-16 16:25:17,764][INFO][discovery.zen] [DM-  
> > > ADSEARCHD102.dev.local-Elasticsearch] master\_left [[dm-  
> > > adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> > > 10.2.20.164:9300]]], reason [shut\_down]  
> > > [2010-11-16 16:25:17,775][INFO][cluster.service] [DM-  
> > > ADSEARCHD102.dev.local-Elasticsearch] master {new [DM-  
> > > ADSEARCHD102.dev.local-Elasticsearch][7yu0RHxETUKROtrrGxBhIw][inet[/  
> > > 10.2.20.160:9300]], previous [dm-adsearchd103.dev.local-Elasticsearch]  
> > > [l3Q3CX84SpGfcq-WzPDaUg][inet[/10.2.20.164:9300]]}, removed {[dm-  
> > > adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> > > 10.2.20.164:9300]],}, reason: zen-disco-master\_failed ([dm-  
> > > adsearchd103.dev.local-Elasticsearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/  
> > > 10.2.20.164:9300]])
> 
> > > This looks correct, however, we were running two applications with  
> > > Node based clients and they logged nothing at the time of the initial  
> > > disconnect, although, I did see a search fail for  
> > > AlreadyClosedException. A few minutes later, I see:
> 
> > > 2010-11-16 16:28:15,879 WARN \> [dm-adsearchd102.dev.local-  
> > > essearcherserver] master\_left and no other node elected to become  
> > > master, current nodes: [[WKSPPEARCYW7.wsod.local-ESIndexer]  
> > > [L2QpS2IfTLqsZRf1OB76rw][inet[/10.2.123.16:9300]]{client=true,  
> > > data=false}, [dm-adsearchd102.dev.local-ESIndexer]  
> > > [qrSmIjycSA2pQ0OtVcdDng][inet[/10.2.20.160:9301]]{client=true,  
> > > data=false}, [dm-adsearchd102.dev.local-essearcherserver]  
> > > [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
> > > data=false}] (Log4jESLogger.internalWarn:87)(elasticsearch[dm-  
> > > adsearchd102.dev.local-essearcherserver]clusterService#updateTask-  
> > > pool-5-thread-1)
> 
> > > Afterwards, the node clients failed all search and index requests.  
> > > Didn't notice immediately, 20 minutes later noticed, and then  
> > > restarted both apps using node clients and they both came up cleanly.
> 
> > > I'm not sure if somehow there was a split brain condition, more like  
> > > different nodes had a disagreement on who was in the cluster. I do see  
> > > the node client attempting to re-establish communications with the one  
> > > active node, but it is failing:
> 
> > > 2010-11-16 16:28:19,746 WARN \> [dm-adsearchd102.dev.local-  
> > > essearcherserver] failed to send ping to [[#zen\_unicast\_1#][inet[DM-  
> > > ADSEARCHD102.dev.local/10.2.20.160:9300]]] (Log4jESLogger.internalWarn:  
> > > 91)(elasticsearch[dm-adsearchd102.dev.local-essearcherserver][tp]-  
> > > pool-1-thread-23)  
> > > org.elasticsearch.transport.ReceiveTimeoutTransportException:   
> > > [inet[DM-ADSEARCHD102.dev.local/10.2.20.160:9300]][discovery/zen/  
> > > unicast]  
> > > at org.elasticsearch.transport.TransportService  
> > > $TimeoutTimerTask.run(TransportService.java:316)  
> > > at org.elasticsearch.timer.TimerService$ThreadedTimerTask  
> > > $1.run(TimerService.java:113)  
> > > at java.util.concurrent.ThreadPoolExecutor  
> > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > at java.util.concurrent.ThreadPoolExecutor  
> > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > at java.lang.Thread.run(Thread.java:619)
> 
> > > On the one good data node, it seems to be ignoring join requests:  
> > > [2010-11-16 16:31:42,880][WARN][discovery.zen] [DM-  
> > > ADSEARCHD102.dev.local-Elasticsearch] received a join request for an  
> > > existing node [[dm-adsearchd102.dev.local-essearcherserver]  
> > > [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true,  
> > > data=false}]
> 
> > > I can make all the logs available, if that would help.
> 
> > > In my app, I believe if I had closed the node client and re-opened, it  
> > > probably would have addressed.
> 
> > > Please let me know if you need any other details.
> 
> > > Thanks,  
> > > Paul

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 19, 2010, 8:23pm UTC](https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577/5 "2010-11-19T20:23:13Z")

</div>

```
    Hey,
    
    Â Â Someone was using snapshot versions of 0.13 and got this behavior as well, and it was because of mixed 0.13 versions (the cluster state has changed in between). This completely explains the behavior...Â Â ping if it happens...cheers,-shay.banon
	
	
    On Friday, November 19, 2010 at 10:20 PM, Paul wrote:
    
        Hey, apologies for th delayed response. We've tried to be careful withmixed versions, but it is possible and we were upgrading at the time.On 0.13.0 and haven't been able to reproduce. Will post logs if I canhit this again.Thanks!On Nov 17, 2:36Â pm, Shay Banon <shay.ba...@elasticsearch.com> wrote: I think I know why you got this failure, is there a chance that you used different 0.13 versions between the client and the server? -shay.banon On Wed, Nov 17, 2010 at 11:01 AM, Shay Banon wrote: > Hey Paul, > Â Can you mail me the logs? I will have a look. > -shay.banon > On Tue, Nov 16, 2010 at 7:38 PM, Paul wrote: >> Hey, >> Â Running master from yesterday morning. Have never hit this issue >> before and was previously running master from a week before that. I >> have two data nodes running: >> dm-adsearchd102-ElasticSearch >> dm-adsearchd103-ElasticSearch >> I shutdown the d103 node via the service stop command and saw this in >> the logs for the d102 node: >> [2010-11-16

```

16:25:17,764][INFO][discovery.zen Â Â Â Â Â Â] [DM- \>\> ADSEARCHD102.dev.local-ElasticSearch] master\_left [[dm- \>\> adsearchd103.dev.local-ElasticSearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/ \>\> 10.2.20.164:9300]]], reason [shut\_down] \>\> [2010-11-16 16:25:17,775][INFO][cluster.service Â Â Â Â Â] [DM- \>\> ADSEARCHD102.dev.local-ElasticSearch] master {new [DM- \>\> ADSEARCHD102.dev.local-ElasticSearch][7yu0RHxETUKROtrrGxBhIw][inet[/ \>\> 10.2.20.160:9300]], previous [dm-adsearchd103.dev.local-ElasticSearch] \>\> [l3Q3CX84SpGfcq-WzPDaUg][inet[/10.2.20.164:9300]]}, removed {[dm- \>\> adsearchd103.dev.local-ElasticSearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/ \>\> 10.2.20.164:9300]],}, reason: zen-disco-master\_failed ([dm- \>\> adsearchd103.dev.local-ElasticSearch][l3Q3CX84SpGfcq-WzPDaUg][inet[/ \>\> 10.2.20.164:9300]]) \>\> This looks correct, however, we were running two applications with \>\> Node based clients and they logged nothing at the time of the initial \>\> disconnect, although, I did see a search f  
ail for \>\> AlreadyClosedException. A few minutes later, I see: \>\> 2010-11-16 16:28:15,879 Â WARN \> [dm-adsearchd102.dev.local- \>\> essearcherserver] master\_left and no other node elected to become \>\> master, current nodes: [[WKSPPEARCYW7.wsod.local-ESIndexer] \>\> [L2QpS2IfTLqsZRf1OB76rw][inet[/10.2.123.16:9300]]{client=true, \>\> data=false}, [dm-adsearchd102.dev.local-ESIndexer] \>\> [qrSmIjycSA2pQ0OtVcdDng][inet[/10.2.20.160:9301]]{client=true, \>\> data=false}, [dm-adsearchd102.dev.local-essearcherserver] \>\> [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true, \>\> data=false}] (Log4jESLogger.internalWarn:87)(elasticsearch[dm- \>\> adsearchd102.dev.local-essearcherserver]clusterService#updateTask- \>\> pool-5-thread-1) \>\> Afterwards, the node clients failed all search and index requests. \>\> Didn't notice immediately, 20 minutes later noticed, and then \>\> restarted both apps using node clients and they both came up cleanly. \>\> I'm not sure if somehow there was a split brain condition,  
more like \>\> different nodes had a disagreement on who was in the cluster. I do see \>\> the node client attempting to re-establish communications with the one \>\> active node, but it is failing: \>\> 2010-11-16 16:28:19,746 Â WARN \> [dm-adsearchd102.dev.local- \>\> essearcherserver] failed to send ping to [[#zen\_unicast\_1#][inet[DM- \>\> ADSEARCHD102.dev.local/10.2.20.160:9300]]] (Log4jESLogger.internalWarn: \>\> 91)(elasticsearch[dm-adsearchd102.dev.local-essearcherserver][tp]- \>\> pool-1-thread-23) \>\> org.elasticsearch.transport.ReceiveTimeoutTransportException: [] \>\> [inet[DM-ADSEARCHD102.dev.local/10.2.20.160:9300]][discovery/zen/ \>\> unicast] \>\> Â Â Â Â at org.elasticsearch.transport.TransportService \>\> $TimeoutTimerTask.run(TransportService.java:316) \>\> Â Â Â Â at org.elasticsearch.timer.TimerService$ThreadedTimerTask \>\> $1.run(TimerService.java:113) \>\> Â Â Â Â at java.util.concurrent.ThreadPoolExecutor \>\> $Worker.runTask(ThreadPoolExecutor.java:886) \>\> Â Â Â Â at java.util.con  
current.ThreadPoolExecutor \>\> $Worker.run(ThreadPoolExecutor.java:908) \>\> Â Â Â Â at java.lang.Thread.run(Thread.java:619) \>\> On the one good data node, it seems to be ignoring join requests: \>\> [2010-11-16 16:31:42,880][WARN][discovery.zen Â Â Â Â Â Â] [DM- \>\> ADSEARCHD102.dev.local-ElasticSearch] received a join request for an \>\> existing node [[dm-adsearchd102.dev.local-essearcherserver] \>\> [7G6qabvhRKuAuxRBN92t6Q][inet[/10.2.20.160:9302]]{client=true, \>\> data=false}] \>\> I can make all the logs available, if that would help. \>\> In my app, I believe if I had closed the node client and re-opened, it \>\> probably would have addressed. \>\> Please let me know if you need any other details. \>\> Thanks, \>\> [Paul...@gmail.com](mailto:Paul...@gmail.com)\>[.ba...@elasticsearch.com](mailto:.ba...@elasticsearch.com)\>

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:16am UTC](https://discuss.elastic.co/t/master-left-and-no-other-node-elected-to-become-master/3577/6 "2017-07-06T04:16:16Z")

</div>


