# 5 node cluster breaks when master is shut down

**URL:** <https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027>\
**Category:** Elasticsearch\
**Created:** [January 12, 2016, 7:34pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027 "2016-01-12T19:34:49Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Stuart\_Cracraft](https://avatars.discourse-cdn.com/v4/letter/s/d6d6ee/32.png) [@Stuart\_Cracraft](https://discuss.elastic.co/u/Stuart_Cracraft)\
**Post date:** [January 12, 2016, 7:34pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/1 "2016-01-12T19:34:49Z")

</div>

The elasticsearch.yml on the master is:

[cluster.name](http://cluster.name): elasticsearchlogstashkibana  
[node.name](http://node.name): "elksrv1"  
node.master: true  
node.data: true  
index.number\_of\_replicas: 2  
index.number\_of\_shards: 5  
indices.recovery.compress: false

The 4 other nodes are as above but:

[cluster.name](http://cluster.name): elasticsearchlogstashkibana  
[node.name](http://node.name): "elksrv[2-4]"  
node.master: false  
node.data: true  
index.number\_of\_replicas: 2  
index.number\_of\_shards: 5  
indices.recovery.compress: false

What are the settings to permit resiliency of master so that when it crashes or service is taken down the cluster survives?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [January 12, 2016, 7:38pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/2 "2016-01-12T19:38:32Z")

</div>

Don't have a single master. For a five node cluster it's most likely unnecessary and wasteful to have a dedicated master, especially since it becomes a single point of failure, so just make all data nodes master-eligible and drop the current dedicated master. If you insist on dedicated masters you need three of them.

---

<div class="post-metadata">

**Author:** ![Stuart\_Cracraft](https://avatars.discourse-cdn.com/v4/letter/s/d6d6ee/32.png) [@Stuart\_Cracraft](https://discuss.elastic.co/u/Stuart_Cracraft)\
**Post date:** [January 12, 2016, 7:39pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/3 "2016-01-12T19:39:23Z")

</div>

So the other nodes should have node.master: true instead of the present node.master: false? I.e. all nodes have node.master: true?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [January 12, 2016, 7:40pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/4 "2016-01-12T19:40:16Z")

</div>

Oh, and make sure you set discovery.zen.minimum\_master\_nodes to N/2+1, i.e. 2 for three node clusters and 3 for four och five node clusters.

---

<div class="post-metadata">

**Author:** ![Stuart\_Cracraft](https://avatars.discourse-cdn.com/v4/letter/s/d6d6ee/32.png) [@Stuart\_Cracraft](https://discuss.elastic.co/u/Stuart_Cracraft)\
**Post date:** [January 12, 2016, 7:41pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/5 "2016-01-12T19:41:41Z")

</div>

So I am hearing this:

[cluster.name](http://cluster.name): elasticsearchlogstashkibana  
[node.name](http://node.name): "elksrv[1-5]"  
node.master: true  
node.data: true  
index.number\_of\_replicas: 2  
index.number\_of\_shards: 5  
indices.recovery.compress: false  
discovery.zen.minimum\_master\_nodes: 3

for all.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [January 12, 2016, 7:50pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/6 "2016-01-12T19:50:44Z")

</div>

Yes. And you'll probably want to have a five node cluster since a four node cluster won't be able to survive two nodes being down (which I guess is the point of having two replicas?). If your current dedicated master isn't powerful enough to be a data node you can keep it as a pure master node, but that doesn't mean it'll actually be elected master.

---

<div class="post-metadata">

**Author:** ![Stuart\_Cracraft](https://avatars.discourse-cdn.com/v4/letter/s/d6d6ee/32.png) [@Stuart\_Cracraft](https://discuss.elastic.co/u/Stuart_Cracraft)\
**Post date:** [January 12, 2016, 9:46pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/7 "2016-01-12T21:46:41Z")

</div>

Should discovery.zen.ping.unicast.hosts: have the list of all the nodes, nothing, or something else?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [January 13, 2016, 4:31am UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/8 "2016-01-13T04:31:05Z")

</div>

You don't have to list all the nodes, but you need to list enough nodes so that the cluster will be able to form even if some of the nodes are down. In your case you'll want to list at least three nodes since two can be out of service. However, since you should be managing your config files with a configuration management tool that can generate files based on templating it might be just as easy to list all nodes.

---

<div class="post-metadata">

**Author:** ![Stuart\_Cracraft](https://avatars.discourse-cdn.com/v4/letter/s/d6d6ee/32.png) [@Stuart\_Cracraft](https://discuss.elastic.co/u/Stuart_Cracraft)\
**Post date:** [January 13, 2016, 11:02pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/9 "2016-01-13T23:02:14Z")

</div>

Okay - this seems to work, mostly, for each of the 5 node's elasticsearch.yml  
(but see the concern below the config):  
[cluster.name](http://cluster.name): elasticsearchlogstashkibana  
[node.name](http://node.name): "elksrv5"  
node.master: true  
node.data: true  
index.number\_of\_shards: 5  
index.number\_of\_replicas: 2  
indices.recovery.compress: false  
discovery.zen.minimum\_master\_nodes: 3  
discovery.zen.ping.multicast.enabled: false  
discovery.zen.ping.unicast.hosts: ["[elksrv1.channel-corp.com](http://elksrv1.channel-corp.com)","[elksrv2.channel-corp.com](http://elksrv2.channel-corp.com)","[elksrv3.channel-corp.com](http://elksrv3.channel-corp.com)","[elksrv4.channel-corp.com](http://elksrv4.channel-corp.com)","[elksrv5.channel-corp.com](http://elksrv5.channel-corp.com)"]  
script.engine.groovy.inline.update: on

When I test this by taking the non-master's down, no problem. We go yellow then  
green after it automatically reassigns a moderate number of shards. The process takes perhaps a minute or two.

When I take this by taking the master down, problem. We go to red and stay red with unassigned shard count (small usually - tested on two nodes) When I turn the now non-master (election to a new master does take place), the small unassigned shard count goes back to zero after a minute and it goes from red to yellow to green.

But the cluster does not recover to yellow or green from a down master, unless I am missing something.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [January 14, 2016, 7:23am UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/10 "2016-01-14T07:23:28Z")

</div>

> When I take this by taking the master down, problem. We go to red and stay red with unassigned shard count (small usually - tested on two nodes) When I turn the now non-master (election to a new master does take place), the small unassigned shard count goes back to zero after a minute and it goes from red to yellow to green.

So you're saying that shards stay unassigned even though a new master is elected and there is at least one replica of the shards in question? That's unexpected. Are there any clues in the ES logs?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 14, 2016, 8:08am UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/11 "2016-01-14T08:08:29Z")

</div>

Which version of Elasticsearch are you using? Are all nodes in the cluster the same version?

---

<div class="post-metadata">

**Author:** ![Stuart\_Cracraft](https://avatars.discourse-cdn.com/v4/letter/s/d6d6ee/32.png) [@Stuart\_Cracraft](https://discuss.elastic.co/u/Stuart_Cracraft)\
**Post date:** [January 15, 2016, 6:31pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/12 "2016-01-15T18:31:11Z")

</div>

1 is running 1.2.2.

The others are 1.6.0

Is this a problem?

How do I upgrade 1.2.2 to 1.6.0.

Thanks ahead.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 15, 2016, 6:39pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/13 "2016-01-15T18:39:52Z")

</div>

Yes, that is a problem. Different versions use different Lucene versions, so when a shard has been upgraded on one of the newer instances, it can no longer be reallocated to the older node. There may also be other issues depending on which versions are use, so all nodes in a cluster should always be the same version.

Instructions for upgrading can be found [here](https://www.elastic.co/guide/en/elasticsearch/reference/1.6/setup-upgrade.html).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:24pm UTC](https://discuss.elastic.co/t/5-node-cluster-breaks-when-master-is-shut-down/39027/14 "2017-07-05T23:24:02Z")

</div>


