# Master election takes too long

**URL:** <https://discuss.elastic.co/t/master-election-takes-too-long/177305>\
**Category:** Elasticsearch\
**Created:** [April 17, 2019, 1:43pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305 "2019-04-17T13:43:40Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 17, 2019, 1:43pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/1 "2019-04-17T13:43:40Z")

</div>

I am using 5.4.1 version.  
I see in the logs

```
[2019-04-16T21:42:40,996][INFO][o.e.d.z.ZenDiscovery] [node-2] master_left [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}], reason [shut_down]

```

After 3 seconds

```
[2019-04-16T21:42:44,000][WARN][o.e.d.z.ZenDiscovery] [node-2] failed to connect to master [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}], retrying...
org.elasticsearch.transport.ConnectTransportException: [node-1][172.20.67.236:9300] connect_timeout[1s]

```

and after 3 seconds

```
[2019-04-16T21:42:47,024][INFO][o.e.c.s.ClusterService] [node-2] detected_master {node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}, reason: zen-disco-receive(from master [master {node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300} committed version [27]])
[2019-04-16T21:42:47,025][WARN][o.e.c.NodeConnectionsService] [node-2] failed to connect to node {node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300} (tried [1] times)

```

My configuration is:

```
discovery.zen.commit_timeout: 2s
discovery.zen.publish_timeout: 2s
discovery.zen.fd.ping_timeout: 1s
transport.tcp.connect_timeout: 1s

```

What can I change in order to reduce the 6-7 seconds time of failover?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 17, 2019, 3:12pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/2 "2019-04-17T15:12:11Z")

</div>

> [@Alex\_Davidovich](#):
>
> My configuration is:
> 
> ```auto
> discovery.zen.commit_timeout: 2s
> discovery.zen.publish_timeout: 2s
> discovery.zen.fd.ping_timeout: 1s
> transport.tcp.connect_timeout: 1s
> 
> ```

These timeouts are dangerously short. I would expect this cluster to become unstable under any kind of realistic load.

> [@Alex\_Davidovich](#):
>
> What can I change in order to reduce the 6-7 seconds time of failover?

7.0.0 has a new, faster, election mechanism, and in 7.1.0 we also expect to avoid synchronous connections to failed nodes thanks to [#39629](https://github.com/elastic/elasticsearch/pull/39629). I recommend upgrading.

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 17, 2019, 3:15pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/3 "2019-04-17T15:15:45Z")

</div>

Thank you for the quick reply.  
Let's say I will not use any of this configurations.  
How much time should I expect ES to failover and elect a new master in 5.4.1?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 17, 2019, 3:23pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/4 "2019-04-17T15:23:19Z")

</div>

It depends on exactly how the master fails. If you simply shut it down then, with default settings, I'd expect a new master to be elected within a little over 4 seconds:

- 1 second to discover the master has disconnected
- 3 seconds to elect the new master
- a few network round-trips to complete the process and re-establish the cluster

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 17, 2019, 3:24pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/5 "2019-04-17T15:24:56Z")

</div>

In my case master was shut down. It took 6-7 seconds with the configurations shown above.  
Does this make sense? I need to configure a logical timeout in my application. If 5 seconds is safe (the 4 you said + 1 buffer), I will use it.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 17, 2019, 3:27pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/6 "2019-04-17T15:27:54Z")

</div>

> [@Alex\_Davidovich](#):
>
> It took 6-7 seconds with the configurations shown above. Does this make sense?

I find it a little surprising that it took 7 seconds if the master were cleanly shut down. Can you share a more comprehensive set of logs from all of the nodes?

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 17, 2019, 5:49pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/7 "2019-04-17T17:49:27Z")

</div>

Sure.

node-1 is the master and stopped cleanly

```
[2019-04-16T21:38:58,779][INFO][o.e.c.s.ClusterService] [node-1] added {{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300},}, reason: zen-disco-node-join[{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}]
[2019-04-16T21:39:00,444][INFO][o.e.c.r.a.AllocationService] [node-1] Cluster health status changed from [YELLOW] to [GREEN] (reason: [shards started [[events_1555318236265][0]] ...]).
[2019-04-16T21:42:40,979][INFO][o.e.n.Node] [node-1] stopping ...
[2019-04-16T21:42:41,154][INFO][o.e.n.Node] [node-1] stopped
[2019-04-16T21:42:41,154][INFO][o.e.n.Node] [node-1] closing ...
[2019-04-16T21:42:41,160][INFO][o.e.n.Node] [node-1] closed

```

node-2

```
[2019-04-16T21:42:40,996][INFO][o.e.d.z.ZenDiscovery] [node-2] master_left [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.
20.67.236}{172.20.67.236:9300}], reason [shut_down]
[2019-04-16T21:42:40,997][WARN][o.e.d.z.ZenDiscovery] [node-2] master left (reason = shut_down), current nodes: nodes: 
{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ} {172.20.67.236}{172.20.67.236:9300}, master
{node-2}{SiCuliEBS8WzO0mBISJNnA}{DzVFJK5FRj-5FRnBxEtjsg}{172.20.67.238}{172.20.67.238:9300}, local
{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}

[2019-04-16T21:42:44,000][WARN][o.e.d.z.ZenDiscovery] [node-2] failed to connect to master [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZ
USfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}], retrying...
org.elasticsearch.transport.ConnectTransportException: [node-1][172.20.67.236:9300] connect_timeout[1s]

```

Took node-2 4 seconds from the time master left to the time it failed to connect... why?

```
[2019-04-16T21:42:47,024][INFO][o.e.c.s.ClusterService] [node-2] detected_master {node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}, reason: zen-disco-receive(from master [master {node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300} committed version [27]])
[2019-04-16T21:42:47,025][WARN][o.e.c.NodeConnectionsService] [node-2] failed to connect to node {node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300} (tried [1] times)

```

Took 3 more seconds to detect the master and still trying to connect the old one. why?

node-3

```
[2019-04-16T21:42:40,994][INFO][o.e.d.z.ZenDiscovery] [node-3] master_left [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.
20.67.236}{172.20.67.236:9300}], reason [shut_down]
[2019-04-16T21:42:40,997][WARN][o.e.d.z.ZenDiscovery] [node-3] master left (reason = shut_down), current nodes: nodes: 
{node-2}{SiCuliEBS8WzO0mBISJNnA}{DzVFJK5FRj-5FRnBxEtjsg}{172.20.67.238}{172.20.67.238:9300}
{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}, master
{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}, local

[2019-04-16T21:42:41,000][WARN][o.e.c.NodeConnectionsService] [node-3] failed to connect to node {node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZ
USfspSh5PQ}{172.20.67.236}{172.20.67.236:9300} (tried [1] times)

[2019-04-16T21:42:47,015][INFO][o.e.c.s.ClusterService] [node-3] new_master {node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}, reason: zen-disco-elected-as-master ([1] nodes joined)[{node-2}{SiCuliEBS8WzO0mBISJNnA}{DzVFJK5FRj-5FRnBxEtjsg}{172.20.67.238}{172.20.67.238:9300}]

```

Trying to perform actions via PreBuiltTransportClient from the java application got these errors when retrying 3 times at a period of 4 seconds

```
WARN 2019-04-16 21:42:43,106 [EventPublisher~i31c] : ElasticManagementClient(ElasticManagementClient.performAction:242) - Failed to perform action, retries left : 2, got exception
org.elasticsearch.cluster.block.ClusterBlockException: blocked by: [SERVICE_UNAVAILABLE/2/no master];

WARN 2019-04-16 21:42:45,111 [EventPublisher~i31c] : ElasticManagementClient(ElasticManagementClient.performAction:242) - Failed to perform action, retries left : 1, got exception
org.elasticsearch.cluster.block.ClusterBlockException: blocked by: [SERVICE_UNAVAILABLE/2/no master];

WARN 2019-04-16 21:42:47,113 [EventPublisher~i31c] : ElasticManagementClient(ElasticManagementClient.performAction:242) - Failed to perform action, retries left : 0, got exception
org.elasticsearch.cluster.block.ClusterBlockException: blocked by: [SERVICE_UNAVAILABLE/2/no master];
```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 17, 2019, 6:52pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/8 "2019-04-17T18:52:07Z")

</div>

> [@Alex\_Davidovich](#):
>
> ```plaintext
> [2019-04-16T21:42:44,000][WARN][o.e.d.z.ZenDiscovery] [node-2] failed to connect to master [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZ
> USfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}], retrying...
> org.elasticsearch.transport.ConnectTransportException: [node-1][172.20.67.236:9300] connect_timeout[1s]
> 
> ```

This is a little odd as it indicates that `node-2` still thinks that `node-1` is the master. Could you share the whole stack trace from this exception, including any `Caused by` inner exceptions?

Also, can you tell us a bit more about your environment? In particular, are you running in Docker?

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 18, 2019, 6:14am UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/9 "2019-04-18T06:14:08Z")

</div>

I agree, it is odd...  
This is the stack

```
[2019-04-16T21:42:44,000][WARN][o.e.d.z.ZenDiscovery] [node-2] failed to connect to master [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}], retrying...
org.elasticsearch.transport.ConnectTransportException: [node-1][172.20.67.236:9300] connect_timeout[1s]
    at org.elasticsearch.transport.netty4.Netty4Transport.connectToChannels(Netty4Transport.java:360) ~[?:?]
    at org.elasticsearch.transport.TcpTransport.openConnection(TcpTransport.java:534) ~[elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.transport.TcpTransport.connectToNode(TcpTransport.java:473) ~[elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.transport.TransportService.connectToNode(TransportService.java:315) ~[elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.transport.TransportService.connectToNode(TransportService.java:302) ~[elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.discovery.zen.ZenDiscovery.joinElectedMaster(ZenDiscovery.java:468) [elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.discovery.zen.ZenDiscovery.innerJoinCluster(ZenDiscovery.java:420) [elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.discovery.zen.ZenDiscovery.access$4100(ZenDiscovery.java:83) [elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.discovery.zen.ZenDiscovery$JoinThreadControl$1.run(ZenDiscovery.java:1197) [elasticsearch-5.4.1.jar:5.4.1]
    at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingRunnable.run(ThreadContext.java:569) [elasticsearch-5.4.1.jar:5.4.1]
    at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) [?:1.8.0_181]
    at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) [?:1.8.0_181]
    at java.lang.Thread.run(Thread.java:748) [?:1.8.0_181]
Caused by: io.netty.channel.AbstractChannel$AnnotatedConnectException: Connection refused: 172.20.67.236/172.20.67.236:9300
    at sun.nio.ch.SocketChannelImpl.checkConnect(Native Method) ~[?:?]
    at sun.nio.ch.SocketChannelImpl.finishConnect(SocketChannelImpl.java:717) ~[?:?]
    at io.netty.channel.socket.nio.NioSocketChannel.doFinishConnect(NioSocketChannel.java:352) ~[?:?]
    at io.netty.channel.nio.AbstractNioChannel$AbstractNioUnsafe.finishConnect(AbstractNioChannel.java:340) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:632) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.processSelectedKeysPlain(NioEventLoop.java:544) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:498) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:458) ~[?:?]
    at io.netty.util.concurrent.SingleThreadEventExecutor$5.run(SingleThreadEventExecutor.java:858) ~[?:?]
    ... 1 more
Caused by: java.net.ConnectException: Connection refused
    at sun.nio.ch.SocketChannelImpl.checkConnect(Native Method) ~[?:?]
    at sun.nio.ch.SocketChannelImpl.finishConnect(SocketChannelImpl.java:717) ~[?:?]
    at io.netty.channel.socket.nio.NioSocketChannel.doFinishConnect(NioSocketChannel.java:352) ~[?:?]
    at io.netty.channel.nio.AbstractNioChannel$AbstractNioUnsafe.finishConnect(AbstractNioChannel.java:340) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:632) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.processSelectedKeysPlain(NioEventLoop.java:544) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:498) ~[?:?]
    at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:458) ~[?:?]
    at io.netty.util.concurrent.SingleThreadEventExecutor$5.run(SingleThreadEventExecutor.java:858) ~[?:?]
    ... 1 more

```

I am not using docker and the nodes are **not** virtual.  
All nodes configured to be data and master eligible.  
Any more information is needed?  
Again, thank you for your help.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 18, 2019, 6:57am UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/10 "2019-04-18T06:57:04Z")

</div>

Ok, I think I can reproduce this by shutting the master down very soon after another node has joined the cluster. Is that what's happening here?

When a node starts up it tries to discover the rest of the cluster, and it keeps hold of the information it discovers for a short while (6 seconds) afterwards. If it needs to perform another election it might use stale information and end up electing a master that's no longer there.

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 18, 2019, 7:22am UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/11 "2019-04-18T07:22:43Z")

</div>

> [@DavidTurner](#):
>
> When a node starts up it tries to discover the rest of the cluster, and it keeps hold of the information it discovers for a short while (6 seconds) afterwards. If it needs to perform another election it might use stale information and end up electing a master that's no longer there.

I think that the gap between the join in and the shutting down of the other node was 4 minutes.

Adding the logs:

node-2

```
[2019-04-16T21:38:58,786][INFO][o.e.c.s.ClusterService] [node-2] added {{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300},}, reason: zen-disco-receive(from master [master {node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300} committed version [24]])
[2019-04-16T21:42:40,996][INFO][o.e.d.z.ZenDiscovery] [node-2] master_left [{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}], reason [shut_down]
[2019-04-16T21:42:40,997][WARN][o.e.d.z.ZenDiscovery] [node-2] master left (reason = shut_down), current nodes: nodes: 
{node-1}{Nuzkh0bnRw6LPnqkO-KUDg}{wxOLyucfTDqZUSfspSh5PQ}{172.20.67.236}{172.20.67.236:9300}, master
{node-2}{SiCuliEBS8WzO0mBISJNnA}{DzVFJK5FRj-5FRnBxEtjsg}{172.20.67.238}{172.20.67.238:9300}, local
{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}

```

node-3

```
[2019-04-17T10:17:35,138][INFO][o.e.h.n.Netty4HttpServerTransport] [node-3] publish_address {172.20.67.240:9200}, bound_addresses {172.20.67.240:9200}
[2019-04-17T10:17:35,140][INFO][o.e.n.Node] [node-3] started
[2019-04-17T10:17:35,408][INFO][o.e.g.GatewayService] [node-3] recovered [1] indices into cluster_state
[2019-04-17T10:17:35,676][INFO][o.e.c.r.a.AllocationService] [node-3] Cluster health status changed from [RED] to [YELLOW] (reason: [shards started [[events_1555318236265][0]] ...]).
[2019-04-17T10:17:37,549][INFO][o.e.c.r.a.AllocationService] [node-3] Cluster health status changed from [YELLOW] to [GREEN] (reason: [shards started [[events_1555318236265][0]] ...]).

```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 18, 2019, 7:34am UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/12 "2019-04-18T07:34:21Z")

</div>

It's quite hard to see clearly all that is going on from these small log excerpts, so I'm having to do a certain amount of speculation. I note, for instance, that `node-3` joined the cluster here too, just before `node-1` was shut down:

> [@Alex\_Davidovich](#):
>
> ```plaintext
> [2019-04-16T21:38:58,779][INFO][o.e.c.s.ClusterService] [node-1] added {{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300},}, reason: zen-disco-node-join[{node-3}{3DbhJ-TaSoGvAmVBYOE3aw}{FL_uo65LS7yGmjxsUVPmXQ}{172.20.67.240}{172.20.67.240:9300}]
> 
> ```

Can you reproduce what you're seeing with `logger.org.elasticsearch.discovery: TRACE` and share the complete logs from all of the nodes?

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 18, 2019, 7:35am UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/13 "2019-04-18T07:35:31Z")

</div>

I will try, it will take at least a day. Thank you

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 18, 2019, 1:15pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/14 "2019-04-18T13:15:57Z")

</div>

Reproduced, I think. node-3 was shutting down.  
Can't put here the logs since there is a limitation.  
Adding links if that's ok:

[node-1](https://www.evernote.com/shard/s571/sh/5d36475e-aa9d-46da-9c1f-d95070c69b23/bac29a97497e22d6b595b094fba1a8c3)

[node-2](https://www.evernote.com/Home.action?login=true#n=1d200e9d-fd2d-41e5-a576-4c64d8e5aae0&s=s571&ses=4&sh=2&sds=5&)

[node-3](https://www.evernote.com/shard/s571/sh/6f543b50-51b2-43e0-9699-28684f261994/d6ef10eb8ff2362235e8f8a37a463c5a)

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 18, 2019, 1:54pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/15 "2019-04-18T13:54:43Z")

</div>

The `node-2` link doesn't work, could you double-check it?

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 18, 2019, 1:59pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/16 "2019-04-18T13:59:48Z")

</div>

Trying again  
[node-2](https://www.evernote.com/shard/s571/sh/1d200e9d-fd2d-41e5-a576-4c64d8e5aae0/2aaf5c3e91dedef0febd3ed984b06750)

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 18, 2019, 2:10pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/17 "2019-04-18T14:10:19Z")

</div>

Ok, I see. `node-3` takes around 200ms to stop:

```auto
[2019-04-18T12:53:55,648][INFO][o.e.n.Node] [node-3] stopping ...
[2019-04-18T12:53:55,823][INFO][o.e.n.Node] [node-3] stopped

```

One of the first things it does is announce that it's shutting down, triggering another election:

```auto
[2019-04-18T12:53:55,659][INFO][o.e.d.z.ZenDiscovery] [node-1] master_left [{node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}], reason [shut_down]
[2019-04-18T12:53:55,659][WARN][o.e.d.z.ZenDiscovery] [node-1] master left (reason = shut_down), current nodes: nodes:
   {node-2}{nS8tF5PhTmau1Jc8mJhx_A}{Xw7nuJ5IS4uEgUulQEb3YQ}{172.20.71.33}{172.20.71.33:9300}
   {node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}, master
   {node-1}{qtQ9zxHGS0yu8VyS_GDCuw}{a3ufzEfyQAubZvHWX9GXZg}{172.20.71.43}{172.20.71.43:9300}, local

[2019-04-18T12:53:55,659][DEBUG][o.e.d.z.MasterFaultDetection] [node-1] [master] stopping fault detection against master [{node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}], reason [master left (reason = shut_down)]
[2019-04-18T12:53:55,659][TRACE][o.e.d.z.NodeJoinController] [node-1] starting an election context, will accumulate joins

```

However, `node-1` reacts very quickly to the news, starts its discovery phase, and discovers `node-3` within under 20ms:

```auto
[2019-04-18T12:53:55,659][TRACE][o.e.d.z.ZenDiscovery] [node-1] starting to ping
[2019-04-18T12:53:55,660][TRACE][o.e.d.z.UnicastZenPing] [node-1] resolved host [172.20.71.43] to [172.20.71.43:9300]
[2019-04-18T12:53:55,660][TRACE][o.e.d.z.UnicastZenPing] [node-1] resolved host [172.20.71.33] to [172.20.71.33:9300]
[2019-04-18T12:53:55,660][TRACE][o.e.d.z.UnicastZenPing] [node-1] resolved host [172.20.71.55] to [172.20.71.55:9300]
[2019-04-18T12:53:55,660][TRACE][o.e.d.z.UnicastZenPing] [node-1] [3] sending to {node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}
[2019-04-18T12:53:55,660][TRACE][o.e.d.z.UnicastZenPing] [node-1] [3] sending to {node-1}{qtQ9zxHGS0yu8VyS_GDCuw}{a3ufzEfyQAubZvHWX9GXZg}{172.20.71.43}{172.20.71.43:9300}
[2019-04-18T12:53:55,660][TRACE][o.e.d.z.UnicastZenPing] [node-1] [3] sending to {node-2}{nS8tF5PhTmau1Jc8mJhx_A}{Xw7nuJ5IS4uEgUulQEb3YQ}{172.20.71.33}{172.20.71.33:9300}
[2019-04-18T12:53:55,661][TRACE][o.e.d.z.UnicastZenPing] [node-1] [3] received response from {node-1}{qtQ9zxHGS0yu8VyS_GDCuw}{a3ufzEfyQAubZvHWX9GXZg}{172.20.71.43}{172.20.71.43:9300}: [ping_response{node [{node-1}{qtQ9zxHGS0yu8VyS_GDCuw}{a3ufzEfyQAubZvHWX9GXZg}{172.20.71.43}{172.20.71.43:9300}], id[24], master [null],cluster_state_version [230], cluster_name[elastic-infinibox]}, ping_response{node [{node-1}{qtQ9zxHGS0yu8VyS_GDCuw}{a3ufzEfyQAubZvHWX9GXZg}{172.20.71.43}{172.20.71.43:9300}], id[25], master [null],cluster_state_version [230], cluster_name[elastic-infinibox]}]
[2019-04-18T12:53:55,661][TRACE][o.e.d.z.UnicastZenPing] [node-1] [3] received response from {node-2}{nS8tF5PhTmau1Jc8mJhx_A}{Xw7nuJ5IS4uEgUulQEb3YQ}{172.20.71.33}{172.20.71.33:9300}: [ping_response{node [{node-1}{qtQ9zxHGS0yu8VyS_GDCuw}{a3ufzEfyQAubZvHWX9GXZg}{172.20.71.43}{172.20.71.43:9300}], id[24], master [null],cluster_state_version [230], cluster_name[elastic-infinibox]}, ping_response{node [{node-2}{nS8tF5PhTmau1Jc8mJhx_A}{Xw7nuJ5IS4uEgUulQEb3YQ}{172.20.71.33}{172.20.71.33:9300}], id[8], master [{node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}],cluster_state_version [230], cluster_name[elastic-infinibox]}]
[2019-04-18T12:53:55,661][TRACE][o.e.d.z.UnicastZenPing] [node-1] [3] received response from {node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}: [ping_response{node [{node-1}{qtQ9zxHGS0yu8VyS_GDCuw}{a3ufzEfyQAubZvHWX9GXZg}{172.20.71.43}{172.20.71.43:9300}], id[24], master [null],cluster_state_version [230], cluster_name[elastic-infinibox]}, ping_response{node [{node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}], id[27], master [{node-3}{8v1o1noDS5e_e7TAgI9K7A}{WCrG1STSQMWu6MCMRiHIqQ}{172.20.71.55}{172.20.71.55:9300}],cluster_state_version [230], cluster_name[elastic-infinibox]}]

```

I think we could call this a bug: `node-3` arguably shouldn't respond to pings after it's announced that it's shutting down. The only maintained version with this bug is now 6.7 (it doesn't occur in 7.0 or later) and I don't think it's critical enough to warrant a fix there.

This means that what I said earlier isn't the case - it now looks like it can take ~7 seconds to elect a new master because sometimes there will be two elections, each taking 3 seconds. TIL.

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 18, 2019, 2:13pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/18 "2019-04-18T14:13:03Z")

</div>

> [@DavidTurner](#):
>
> he only maintained version with this bug is now 6.7 (it doesn't occur in 7.0 or later) and I don't think it's critical enough to warrant a fix there.
> 
> This means that what I said earlier isn't the case - it now looks like it can take ~7 seconds to elect a new master because sometimes there will be two elections, each taking 3 seconds. TIL

Thank you very much for the help. We will increase our application's timeouts to give master election longer grace period. 15 seconds would be enough in your opinion?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 18, 2019, 2:14pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/19 "2019-04-18T14:14:17Z")

</div>

> [@Alex\_Davidovich](#):
>
> 15 seconds would be enough in your opinion?

Yes, that should be ample time to discover a master after shutting the master down.

---

<div class="post-metadata">

**Author:** ![Alex\_Davidovich](https://avatars.discourse-cdn.com/v4/letter/a/b9e5f3/32.png) [@Alex\_Davidovich](https://discuss.elastic.co/u/Alex_Davidovich)\
**Post date:** [April 18, 2019, 2:15pm UTC](https://discuss.elastic.co/t/master-election-takes-too-long/177305/20 "2019-04-18T14:15:21Z")

</div>

Thank you!

[Next page](https://discuss.elastic.co/t/master-election-takes-too-long/177305.md?page=2)
