# Avoiding the Split Brain

**URL:** <https://discuss.elastic.co/t/avoiding-the-split-brain/128746>\
**Category:** Elasticsearch\
**Created:** [April 19, 2018, 4:19pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746 "2018-04-19T16:19:29Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![aslamy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aslamy1/32/26607_2.png) [@aslamy1](https://discuss.elastic.co/u/aslamy1)\
**Post date:** [April 19, 2018, 4:19pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/1 "2018-04-19T16:19:29Z")

</div>

Hi !  
I have a cluster with two nodes.  
Can I avoid split brain if I always send create/update/delete requests to one of this nodes ?  
(I know there are other solutions)

---

<div class="post-metadata">

**Author:** ![Ant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ant/32/33267_2.png) [@Ant](https://discuss.elastic.co/u/Ant)\
**Post date:** [April 19, 2018, 4:33pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/2 "2018-04-19T16:33:04Z")

</div>

I don't think that would help as if they lost connectivity then you would still have the issue that each decided they were the master of their own single node cluster the one you update would be different to the one you didn't so they would still have 2 different views of the same indexes when they came back together.

If they live on the same swtich then the only way they get partitioned is if the switch dies in which case you can't talk to it anyway.

the only way I can see you avoiding split brain though is to set

`discovery.zen.minimum_master_nodes: 2`  
in  
`elasticsearch.yml`

that way unless both are online they won't do anything, the down side of that is if one node dies you lost the cluster regardless but youcan then always set the value to 1 and restart the node that is still alive.

---

<div class="post-metadata">

**Author:** ![JKhondhu](https://avatars.discourse-cdn.com/v4/letter/j/ed655f/32.png) [@JKhondhu](https://discuss.elastic.co/u/JKhondhu)\
**Post date:** [April 19, 2018, 4:36pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/3 "2018-04-19T16:36:27Z")

</div>

+1 to the above but also, elasticsearch does request routing on your behalf, so you don't need to think about where you send the create/update/delete req to. It will be routed to the correct index and its shards as per your commands.

---

<div class="post-metadata">

**Author:** ![aslamy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aslamy1/32/26607_2.png) [@aslamy1](https://discuss.elastic.co/u/aslamy1)\
**Post date:** [April 19, 2018, 4:45pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/4 "2018-04-19T16:45:50Z")

</div>

Thank you

---

<div class="post-metadata">

**Author:** ![aslamy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aslamy1/32/26607_2.png) [@aslamy1](https://discuss.elastic.co/u/aslamy1)\
**Post date:** [April 20, 2018, 7:13am UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/5 "2018-04-20T07:13:08Z")

</div>

The problem is not that a server goes down completely. but when the servers can not communicate with each other. Both can be master.

If I add discovery.zen.ping\_timeout to 60s, the chance of Split Brain will be very low. If they can't communicate for more than 60 seconds, probably one has completely gone down.

---

<div class="post-metadata">

**Author:** ![Bernt\_Rostad](https://avatars.discourse-cdn.com/v4/letter/b/3ab097/32.png) [@Bernt\_Rostad](https://discuss.elastic.co/u/Bernt_Rostad)\
**Post date:** [April 20, 2018, 9:38am UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/6 "2018-04-20T09:38:56Z")

</div>

If you only have two nodes in your cluster I think a better solution would be to elect one of them to always be master by setting **node.master: false** in the in elasticsearch.yml for the other node. That's the only way you can guarantee no split brains in your 2-node cluster.

---

<div class="post-metadata">

**Author:** ![aslamy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aslamy1/32/26607_2.png) [@aslamy1](https://discuss.elastic.co/u/aslamy1)\
**Post date:** [April 20, 2018, 10:39am UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/7 "2018-04-20T10:39:16Z")

</div>

Thank for your answer Bernt\_Rostad.  
In that case I lose high availability advantage. If Master node goes down then nothing will work.

My nodes live in the same network. If they do not communicate with each other for more than 1 minute then there is a high probability that one has gone down. And the second node can take over and become a master.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 20, 2018, 11:06am UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/8 "2018-04-20T11:06:04Z")

</div>

With just 2 nodes you can not achieve true high availability. If one node goes down or the two nodes loses connectivity, the cluster should not be able to elect a master as a node can not determine whether the other node is down, doing long GC or is just disconnected. I would recommend adding a third small dedicated master node so you can reach a majority even if one of the nodes go down or is disconnected.

---

<div class="post-metadata">

**Author:** ![aslamy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aslamy1/32/26607_2.png) [@aslamy1](https://discuss.elastic.co/u/aslamy1)\
**Post date:** [April 20, 2018, 11:41am UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/9 "2018-04-20T11:41:20Z")

</div>

I understand that 3 nodes is best solution. But if the master node is so busy that it can't communicate during (1-2) minutes with other nodes then even 3 nodes will not help to have high availability.

For our search solution it is ok that the system does not respond for 1-2 minutes. But we want if master goes down after 1-2 minutes a backup system take control (second node).

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 20, 2018, 11:47am UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/10 "2018-04-20T11:47:24Z")

</div>

> [@aslamy1](#):
>
> I understand that 3 nodes is best solution. But if the master node is so busy that it can't communicate during (1-2) minutes with other nodes then even 3 nodes will not help to have high availability.

If the current master node is not able to respond, the remaining nodes can automatically elect another master which will make the cluster available again, so it does help.

> [@aslamy1](#):
>
> For our search solution it is ok that the system does not respond for 1-2 minutes. But we want if master goes down after 1-2 minutes a backup system take control (second node).

You can not have this happen automatically with just 2 nodes without risking a split brain scenario. This would require manual intervention.

---

<div class="post-metadata">

**Author:** ![aslamy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aslamy1/32/26607_2.png) [@aslamy1](https://discuss.elastic.co/u/aslamy1)\
**Post date:** [April 20, 2018, 1:05pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/11 "2018-04-20T13:05:22Z")

</div>

> If the current master node is not able to respond, the remaining nodes can automatically elect another master which will make the cluster available again, so it does help.

What I do not understand is way it's ok two nodes elect another master but not only one node. What will happen if he remaining nodes automatically elect another master and our lost master come back? Now we have again two master nodes.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 20, 2018, 1:12pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/12 "2018-04-20T13:12:51Z")

</div>

In Elasticsearch it requires a majority of master-eligible nodes to elect a master. This means that 2 nodes are required to be present in both cases. If the master goes away or is partitioned off in a three node cluster, it will no longer be a master as it does not have a majority of nodes behind him. The two nodes that are still up and in contact can form a majority and elect a new master. When the former master node comes back it will join as a non-master node.

If you only have two nodes, no master can be elected as long as both nodes are not present. You can still read, but not write in order to prevent data loss.

---

<div class="post-metadata">

**Author:** ![aslamy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aslamy1/32/26607_2.png) [@aslamy1](https://discuss.elastic.co/u/aslamy1)\
**Post date:** [April 20, 2018, 1:15pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/13 "2018-04-20T13:15:14Z")

</div>

Thank you . Now I understand how it works.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 18, 2018, 1:15pm UTC](https://discuss.elastic.co/t/avoiding-the-split-brain/128746/14 "2018-05-18T13:15:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
