# Split brain?

**URL:** <https://discuss.elastic.co/t/split-brain/6256>\
**Category:** Elasticsearch\
**Created:** [December 30, 2011, 3:22am UTC](https://discuss.elastic.co/t/split-brain/6256 "2011-12-30T03:22:05Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Darron\_Froese](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/darron_froese/32/2984_2.png) [@Darron\_Froese](https://discuss.elastic.co/u/Darron_Froese)\
**Post date:** [December 30, 2011, 3:22am UTC](https://discuss.elastic.co/t/split-brain/6256/1 "2011-12-30T03:22:05Z")

</div>

I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5 for  
the last 3 weeks. It's been working great and we've had several  
million records inserted and deleted in that time.

Here is the config:

> **[path: work: /var/lib/elasticsearch/work logs: /var/log/elasticsearch ...](http://d.pr/WbU7)**
>
> path:
> work: /var/lib/elasticsearch/work
> logs: /var/log/elasticsearch
> 
> cluster:
> name: overherd
> 
> discovery:
> zen:
> ping:
> multicast:
> enabled: true
> 
> gateway:
> type: local
> re...

Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
boxes to log to syslog as well) - it appears that something didn't  
work so well during the upgrade and I was left with 2 boxes both  
thinking that they're masters.

Here are the logs from the boxes during the upgrade:

> **[\[18:09:12,503\]\[INFO \]\[cluster.service \] \[Speed Demon\] removed {\[Molten...](http://d.pr/HHbC)**
>
> \[18:09:12,503\]\[INFO \]\[cluster.service \] \[Speed Demon\] removed {\[Molten Man\]\[\_aUSMmtGT5WyOZA4akvo5g\]\[inet\[/50.57.133.231:9300\]\],}, reason: zen-disco-node\_left(\[Molten Man\]\[\_aUSMmtGT5WyOZA4...

  

> **[\[18:09:05,530\]\[INFO \]\[node \] \[Molten Man\] {0.18.5}\[28391\]:...](http://d.pr/n3jk)**
>
> \[18:09:05,530\]\[INFO \]\[node \] \[Molten Man\] {0.18.5}\[28391\]: stopping ...
> \[18:09:14,335\]\[INFO \]\[node \] \[Molten Man\] {0.18.5}\[28391\]: stopped
> \[18:09:14,335\]\[IN...

And then from the next day:

> **[\[03:08:04,390\]\[INFO \]\[node \] \[Glob\] {0.18.6}\[25632\]:...](http://d.pr/QEkU)**
>
> \[03:08:04,390\]\[INFO \]\[node \] \[Glob\] {0.18.6}\[25632\]: stopping ...
> \[03:08:04,669\]\[INFO \]\[node \] \[Glob\] {0.18.6}\[25632\]: stopped
> \[03:08:04,670\]\[INFO \]\[node ...

  

> **[\[03:08:05,189\]\[INFO \]\[cluster.service \] \[Stranger\] removed...](http://d.pr/I7fo)**
>
> \[03:08:05,189\]\[INFO \]\[cluster.service \] \[Stranger\] removed {\[Glob\]\[rn5aKxrLQSOLTWviMqP6Sw\]\[inet\[/50.57.133.226:9300\]\],}, reason: zen-disco-node\_left(\[Glob\]\[rn5aKxrLQSOLTWviMqP6Sw\]\[inet\[/5...

I tried to get them to re-connect, but I couldn't get anything to work  
correctly - they were both completely separate.

I now have a single box with the correct index:

> **[Screen Shot 2011-12-29 at 8.02.12 PM.jpg • Droplr™](http://d.pr/4TTq)**
>
> Shared with Droplr

I have a chef recipe that builds new elasticsearch boxes so I spun up  
a new box and tried to get it to join the cluster, but no dice - it's  
like none of the other boxes exist. I've also tried to go back down to  
0.18.5 - no dice.

Is there a way I can point a new box at that master directly and say:  
"Hey you're a slave, there is the master."

I'm a bit of a loss here - and don't want to admit defeat - but I'm a  
little lost here.

It's a system that we're building and I CAN lose the data, but I  
really want to understand:

1. Why this happened.
2. How I can recover from this.
3. How I can prevent this from happening in the future.

Thanks in advance if anybody can point to something I will greatly  
appreciate it.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 30, 2011, 11:32am UTC](https://discuss.elastic.co/t/split-brain/6256/2 "2011-12-30T11:32:46Z")

</div>

Heya:

First, regarding the split brain, it can obviously happen, especially with  
2 servers. If the network gets disconnected between the two for example. In  
this case, you will have two separate one node cluster and you will need to  
resolve it yourself by restarting one of them. You should see in the logs  
the fact that one node got disconnected from the other. If you had a larger  
cluster you could have defined "minimum\_master\_nodes" parameter to reduce  
the chances of it happening.

I am not sure why when you restart the node its not finding the other node.  
I am assuming you are using unicast discovery, are you sure its configured  
properly? You can set discovery: TRACE in the logging.yml file to see which  
nodes it tries to ping and what the status of that is.

-shay.banon

On Fri, Dec 30, 2011 at 5:22 AM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca) wrote:

> I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5 for  
> the last 3 weeks. It's been working great and we've had several  
> million records inserted and deleted in that time.
> 
> Here is the config:
> 
> [http://d.pr/WbU7](http://d.pr/WbU7)
> 
> Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
> boxes to log to syslog as well) - it appears that something didn't  
> work so well during the upgrade and I was left with 2 boxes both  
> thinking that they're masters.
> 
> Here are the logs from the boxes during the upgrade:
> 
> [http://d.pr/HHbC](http://d.pr/HHbC)  
> [http://d.pr/n3jk](http://d.pr/n3jk)
> 
> And then from the next day:
> 
> [http://d.pr/QEkU](http://d.pr/QEkU)  
> [http://d.pr/I7fo](http://d.pr/I7fo)
> 
> I tried to get them to re-connect, but I couldn't get anything to work  
> correctly - they were both completely separate.
> 
> I now have a single box with the correct index:
> 
> [http://d.pr/4TTq](http://d.pr/4TTq)
> 
> I have a chef recipe that builds new elasticsearch boxes so I spun up  
> a new box and tried to get it to join the cluster, but no dice - it's  
> like none of the other boxes exist. I've also tried to go back down to  
> 0.18.5 - no dice.
> 
> Is there a way I can point a new box at that master directly and say:  
> "Hey you're a slave, there is the master."
> 
> I'm a bit of a loss here - and don't want to admit defeat - but I'm a  
> little lost here.
> 
> It's a system that we're building and I CAN lose the data, but I  
> really want to understand:
> 
> 1. Why this happened.
> 2. How I can recover from this.
> 3. How I can prevent this from happening in the future.
> 
> Thanks in advance if anybody can point to something I will greatly  
> appreciate it.

---

<div class="post-metadata">

**Author:** ![Darron\_Froese](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/darron_froese/32/2984_2.png) [@Darron\_Froese](https://discuss.elastic.co/u/Darron_Froese)\
**Post date:** [December 30, 2011, 7:39pm UTC](https://discuss.elastic.co/t/split-brain/6256/3 "2011-12-30T19:39:13Z")

</div>

I was using multicast discovery and it was working great before -  
heres the log with extra debugging:

[http://d.pr/9pWC](http://d.pr/9pWC)

It looks like it didn't find anything at all.

So I setup unicast and added the master to the config:

[http://d.pr/NxBE](http://d.pr/NxBE)

And it seemed to work:

[http://d.pr/jxjW](http://d.pr/jxjW)  
[http://d.pr/Ubzi](http://d.pr/Ubzi)

I can switch to using unicast - that's no problem - just not sure why  
that happened - maybe Rackspace made some network changes - not sure  
why it worked for almost a month and suddenly stopped.

A couple questions.

1. Should I put all of the IPs of all of the nodes in there?
2. If I go up to a 3 node cluster - should I put "minimum\_master\_nodes" to 2?

Thanks for the tips Shay - really appreciate it.

On Fri, Dec 30, 2011 at 4:32 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Heya:
> 
> First, regarding the split brain, it can obviously happen, especially with 2  
> servers. If the network gets disconnected between the two for example. In  
> this case, you will have two separate one node cluster and you will need to  
> resolve it yourself by restarting one of them. You should see in the logs  
> the fact that one node got disconnected from the other. If you had a larger  
> cluster you could have defined "minimum\_master\_nodes" parameter to reduce  
> the chances of it happening.
> 
> I am not sure why when you restart the node its not finding the other node.  
> I am assuming you are using unicast discovery, are you sure its configured  
> properly? You can set discovery: TRACE in the logging.yml file to see which  
> nodes it tries to ping and what the status of that is.
> 
> -shay.banon
> 
> On Fri, Dec 30, 2011 at 5:22 AM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca) wrote:
> 
> > I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5 for  
> > the last 3 weeks. It's been working great and we've had several  
> > million records inserted and deleted in that time.
> > 
> > Here is the config:
> > 
> > [http://d.pr/WbU7](http://d.pr/WbU7)
> > 
> > Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
> > boxes to log to syslog as well) - it appears that something didn't  
> > work so well during the upgrade and I was left with 2 boxes both  
> > thinking that they're masters.
> > 
> > Here are the logs from the boxes during the upgrade:
> > 
> > [http://d.pr/HHbC](http://d.pr/HHbC)  
> > [http://d.pr/n3jk](http://d.pr/n3jk)
> > 
> > And then from the next day:
> > 
> > [http://d.pr/QEkU](http://d.pr/QEkU)  
> > [http://d.pr/I7fo](http://d.pr/I7fo)
> > 
> > I tried to get them to re-connect, but I couldn't get anything to work  
> > correctly - they were both completely separate.
> > 
> > I now have a single box with the correct index:
> > 
> > [http://d.pr/4TTq](http://d.pr/4TTq)
> > 
> > I have a chef recipe that builds new elasticsearch boxes so I spun up  
> > a new box and tried to get it to join the cluster, but no dice - it's  
> > like none of the other boxes exist. I've also tried to go back down to  
> > 0.18.5 - no dice.
> > 
> > Is there a way I can point a new box at that master directly and say:  
> > "Hey you're a slave, there is the master."
> > 
> > I'm a bit of a loss here - and don't want to admit defeat - but I'm a  
> > little lost here.
> > 
> > It's a system that we're building and I CAN lose the data, but I  
> > really want to understand:
> > 
> > 1. Why this happened.
> > 2. How I can recover from this.
> > 3. How I can prevent this from happening in the future.
> > 
> > Thanks in advance if anybody can point to something I will greatly  
> > appreciate it.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 30, 2011, 8:40pm UTC](https://discuss.elastic.co/t/split-brain/6256/4 "2011-12-30T20:40:22Z")

</div>

Strange that multicast worked..., as far as I know its not supported in  
Rackspace. Yea, you should put all the IPs in the unicast list, its  
recommended if possible. If you go up to 3 nodes, then I think it make  
sense to have minimum master nodes set to 2, yea.

On Fri, Dec 30, 2011 at 9:39 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca) wrote:

> I was using multicast discovery and it was working great before -  
> heres the log with extra debugging:
> 
> [http://d.pr/9pWC](http://d.pr/9pWC)
> 
> It looks like it didn't find anything at all.
> 
> So I setup unicast and added the master to the config:
> 
> [http://d.pr/NxBE](http://d.pr/NxBE)
> 
> And it seemed to work:
> 
> [http://d.pr/jxjW](http://d.pr/jxjW)  
> [http://d.pr/Ubzi](http://d.pr/Ubzi)
> 
> I can switch to using unicast - that's no problem - just not sure why  
> that happened - maybe Rackspace made some network changes - not sure  
> why it worked for almost a month and suddenly stopped.
> 
> A couple questions.
> 
> 1. Should I put all of the IPs of all of the nodes in there?
> 2. If I go up to a 3 node cluster - should I put "minimum\_master\_nodes" to  
> 2?
> 
> Thanks for the tips Shay - really appreciate it.
> 
> On Fri, Dec 30, 2011 at 4:32 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Heya:
> > 
> > First, regarding the split brain, it can obviously happen, especially  
> > with 2  
> > servers. If the network gets disconnected between the two for example. In  
> > this case, you will have two separate one node cluster and you will need  
> > to  
> > resolve it yourself by restarting one of them. You should see in the logs  
> > the fact that one node got disconnected from the other. If you had a  
> > larger  
> > cluster you could have defined "minimum\_master\_nodes" parameter to reduce  
> > the chances of it happening.
> > 
> > I am not sure why when you restart the node its not finding the other  
> > node.  
> > I am assuming you are using unicast discovery, are you sure its  
> > configured  
> > properly? You can set discovery: TRACE in the logging.yml file to see  
> > which  
> > nodes it tries to ping and what the status of that is.
> > 
> > -shay.banon
> > 
> > On Fri, Dec 30, 2011 at 5:22 AM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)  
> > wrote:
> > 
> > > I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5 for  
> > > the last 3 weeks. It's been working great and we've had several  
> > > million records inserted and deleted in that time.
> > > 
> > > Here is the config:
> > > 
> > > [http://d.pr/WbU7](http://d.pr/WbU7)
> > > 
> > > Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
> > > boxes to log to syslog as well) - it appears that something didn't  
> > > work so well during the upgrade and I was left with 2 boxes both  
> > > thinking that they're masters.
> > > 
> > > Here are the logs from the boxes during the upgrade:
> > > 
> > > [http://d.pr/HHbC](http://d.pr/HHbC)  
> > > [http://d.pr/n3jk](http://d.pr/n3jk)
> > > 
> > > And then from the next day:
> > > 
> > > [http://d.pr/QEkU](http://d.pr/QEkU)  
> > > [http://d.pr/I7fo](http://d.pr/I7fo)
> > > 
> > > I tried to get them to re-connect, but I couldn't get anything to work  
> > > correctly - they were both completely separate.
> > > 
> > > I now have a single box with the correct index:
> > > 
> > > [http://d.pr/4TTq](http://d.pr/4TTq)
> > > 
> > > I have a chef recipe that builds new elasticsearch boxes so I spun up  
> > > a new box and tried to get it to join the cluster, but no dice - it's  
> > > like none of the other boxes exist. I've also tried to go back down to  
> > > 0.18.5 - no dice.
> > > 
> > > Is there a way I can point a new box at that master directly and say:  
> > > "Hey you're a slave, there is the master."
> > > 
> > > I'm a bit of a loss here - and don't want to admit defeat - but I'm a  
> > > little lost here.
> > > 
> > > It's a system that we're building and I CAN lose the data, but I  
> > > really want to understand:
> > > 
> > > 1. Why this happened.
> > > 2. How I can recover from this.
> > > 3. How I can prevent this from happening in the future.
> > > 
> > > Thanks in advance if anybody can point to something I will greatly  
> > > appreciate it.

---

<div class="post-metadata">

**Author:** ![Darron\_Froese](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/darron_froese/32/2984_2.png) [@Darron\_Froese](https://discuss.elastic.co/u/Darron_Froese)\
**Post date:** [December 30, 2011, 9:42pm UTC](https://discuss.elastic.co/t/split-brain/6256/5 "2011-12-30T21:42:23Z")

</div>

Yeah - it was working great - all configs are in a git repo and pushed  
out via chef - been working great for a little over 3 weeks in  
production - and a couple weeks before in testing.

I have a ticket into Rackspace to see if they have changed something -  
but will just switch to unicast now.

Thanks for your help - will be updating my configs now.

On Fri, Dec 30, 2011 at 1:40 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Strange that multicast worked..., as far as I know its not supported in  
> Rackspace. Yea, you should put all the IPs in the unicast list, its  
> recommended if possible. If you go up to 3 nodes, then I think it make sense  
> to have minimum master nodes set to 2, yea.
> 
> On Fri, Dec 30, 2011 at 9:39 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca) wrote:
> 
> > I was using multicast discovery and it was working great before -  
> > heres the log with extra debugging:
> > 
> > [http://d.pr/9pWC](http://d.pr/9pWC)
> > 
> > It looks like it didn't find anything at all.
> > 
> > So I setup unicast and added the master to the config:
> > 
> > [http://d.pr/NxBE](http://d.pr/NxBE)
> > 
> > And it seemed to work:
> > 
> > [http://d.pr/jxjW](http://d.pr/jxjW)  
> > [http://d.pr/Ubzi](http://d.pr/Ubzi)
> > 
> > I can switch to using unicast - that's no problem - just not sure why  
> > that happened - maybe Rackspace made some network changes - not sure  
> > why it worked for almost a month and suddenly stopped.
> > 
> > A couple questions.
> > 
> > 1. Should I put all of the IPs of all of the nodes in there?
> > 2. If I go up to a 3 node cluster - should I put "minimum\_master\_nodes" to  
> > 2?
> > 
> > Thanks for the tips Shay - really appreciate it.
> > 
> > On Fri, Dec 30, 2011 at 4:32 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > > Heya:
> > > 
> > > First, regarding the split brain, it can obviously happen, especially  
> > > with 2  
> > > servers. If the network gets disconnected between the two for example.  
> > > In  
> > > this case, you will have two separate one node cluster and you will need  
> > > to  
> > > resolve it yourself by restarting one of them. You should see in the  
> > > logs  
> > > the fact that one node got disconnected from the other. If you had a  
> > > larger  
> > > cluster you could have defined "minimum\_master\_nodes" parameter to  
> > > reduce  
> > > the chances of it happening.
> > > 
> > > I am not sure why when you restart the node its not finding the other  
> > > node.  
> > > I am assuming you are using unicast discovery, are you sure its  
> > > configured  
> > > properly? You can set discovery: TRACE in the logging.yml file to see  
> > > which  
> > > nodes it tries to ping and what the status of that is.
> > > 
> > > -shay.banon
> > > 
> > > On Fri, Dec 30, 2011 at 5:22 AM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)  
> > > wrote:
> > > 
> > > > I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5 for  
> > > > the last 3 weeks. It's been working great and we've had several  
> > > > million records inserted and deleted in that time.
> > > > 
> > > > Here is the config:
> > > > 
> > > > [http://d.pr/WbU7](http://d.pr/WbU7)
> > > > 
> > > > Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
> > > > boxes to log to syslog as well) - it appears that something didn't  
> > > > work so well during the upgrade and I was left with 2 boxes both  
> > > > thinking that they're masters.
> > > > 
> > > > Here are the logs from the boxes during the upgrade:
> > > > 
> > > > [http://d.pr/HHbC](http://d.pr/HHbC)  
> > > > [http://d.pr/n3jk](http://d.pr/n3jk)
> > > > 
> > > > And then from the next day:
> > > > 
> > > > [http://d.pr/QEkU](http://d.pr/QEkU)  
> > > > [http://d.pr/I7fo](http://d.pr/I7fo)
> > > > 
> > > > I tried to get them to re-connect, but I couldn't get anything to work  
> > > > correctly - they were both completely separate.
> > > > 
> > > > I now have a single box with the correct index:
> > > > 
> > > > [http://d.pr/4TTq](http://d.pr/4TTq)
> > > > 
> > > > I have a chef recipe that builds new elasticsearch boxes so I spun up  
> > > > a new box and tried to get it to join the cluster, but no dice - it's  
> > > > like none of the other boxes exist. I've also tried to go back down to  
> > > > 0.18.5 - no dice.
> > > > 
> > > > Is there a way I can point a new box at that master directly and say:  
> > > > "Hey you're a slave, there is the master."
> > > > 
> > > > I'm a bit of a loss here - and don't want to admit defeat - but I'm a  
> > > > little lost here.
> > > > 
> > > > It's a system that we're building and I CAN lose the data, but I  
> > > > really want to understand:
> > > > 
> > > > 1. Why this happened.
> > > > 2. How I can recover from this.
> > > > 3. How I can prevent this from happening in the future.
> > > > 
> > > > Thanks in advance if anybody can point to something I will greatly  
> > > > appreciate it.

---

<div class="post-metadata">

**Author:** ![Darron\_Froese](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/darron_froese/32/2984_2.png) [@Darron\_Froese](https://discuss.elastic.co/u/Darron_Froese)\
**Post date:** [January 4, 2012, 8:26am UTC](https://discuss.elastic.co/t/split-brain/6256/6 "2012-01-04T08:26:19Z")

</div>

FYI - heard from Rackspace:

"We did recently implement multicast filtering on our new XenServer  
Linux deployments. This was originally the intended design as having  
multicast between all customers in the same huddles can be  
problematic. I apologize that you used this as a feature before it was  
blocked, but I feel it may ultimately the best with multicast  
filtered.

Your timeline corresponds exactly with when I heard that the changes  
were being rolled out."

Oh well - makes sense now.

On Fri, Dec 30, 2011 at 2:42 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca) wrote:

> Yeah - it was working great - all configs are in a git repo and pushed  
> out via chef - been working great for a little over 3 weeks in  
> production - and a couple weeks before in testing.
> 
> I have a ticket into Rackspace to see if they have changed something -  
> but will just switch to unicast now.
> 
> Thanks for your help - will be updating my configs now.
> 
> On Fri, Dec 30, 2011 at 1:40 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Strange that multicast worked..., as far as I know its not supported in  
> > Rackspace. Yea, you should put all the IPs in the unicast list, its  
> > recommended if possible. If you go up to 3 nodes, then I think it make sense  
> > to have minimum master nodes set to 2, yea.
> > 
> > On Fri, Dec 30, 2011 at 9:39 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca) wrote:
> > 
> > > I was using multicast discovery and it was working great before -  
> > > heres the log with extra debugging:
> > > 
> > > [http://d.pr/9pWC](http://d.pr/9pWC)
> > > 
> > > It looks like it didn't find anything at all.
> > > 
> > > So I setup unicast and added the master to the config:
> > > 
> > > [http://d.pr/NxBE](http://d.pr/NxBE)
> > > 
> > > And it seemed to work:
> > > 
> > > [http://d.pr/jxjW](http://d.pr/jxjW)  
> > > [http://d.pr/Ubzi](http://d.pr/Ubzi)
> > > 
> > > I can switch to using unicast - that's no problem - just not sure why  
> > > that happened - maybe Rackspace made some network changes - not sure  
> > > why it worked for almost a month and suddenly stopped.
> > > 
> > > A couple questions.
> > > 
> > > 1. Should I put all of the IPs of all of the nodes in there?
> > > 2. If I go up to a 3 node cluster - should I put "minimum\_master\_nodes" to  
> > > 2?
> > > 
> > > Thanks for the tips Shay - really appreciate it.
> > > 
> > > On Fri, Dec 30, 2011 at 4:32 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > 
> > > > Heya:
> > > > 
> > > > First, regarding the split brain, it can obviously happen, especially  
> > > > with 2  
> > > > servers. If the network gets disconnected between the two for example.  
> > > > In  
> > > > this case, you will have two separate one node cluster and you will need  
> > > > to  
> > > > resolve it yourself by restarting one of them. You should see in the  
> > > > logs  
> > > > the fact that one node got disconnected from the other. If you had a  
> > > > larger  
> > > > cluster you could have defined "minimum\_master\_nodes" parameter to  
> > > > reduce  
> > > > the chances of it happening.
> > > > 
> > > > I am not sure why when you restart the node its not finding the other  
> > > > node.  
> > > > I am assuming you are using unicast discovery, are you sure its  
> > > > configured  
> > > > properly? You can set discovery: TRACE in the logging.yml file to see  
> > > > which  
> > > > nodes it tries to ping and what the status of that is.
> > > > 
> > > > -shay.banon
> > > > 
> > > > On Fri, Dec 30, 2011 at 5:22 AM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)  
> > > > wrote:
> > > > 
> > > > > I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5 for  
> > > > > the last 3 weeks. It's been working great and we've had several  
> > > > > million records inserted and deleted in that time.
> > > > > 
> > > > > Here is the config:
> > > > > 
> > > > > [http://d.pr/WbU7](http://d.pr/WbU7)
> > > > > 
> > > > > Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
> > > > > boxes to log to syslog as well) - it appears that something didn't  
> > > > > work so well during the upgrade and I was left with 2 boxes both  
> > > > > thinking that they're masters.
> > > > > 
> > > > > Here are the logs from the boxes during the upgrade:
> > > > > 
> > > > > [http://d.pr/HHbC](http://d.pr/HHbC)  
> > > > > [http://d.pr/n3jk](http://d.pr/n3jk)
> > > > > 
> > > > > And then from the next day:
> > > > > 
> > > > > [http://d.pr/QEkU](http://d.pr/QEkU)  
> > > > > [http://d.pr/I7fo](http://d.pr/I7fo)
> > > > > 
> > > > > I tried to get them to re-connect, but I couldn't get anything to work  
> > > > > correctly - they were both completely separate.
> > > > > 
> > > > > I now have a single box with the correct index:
> > > > > 
> > > > > [http://d.pr/4TTq](http://d.pr/4TTq)
> > > > > 
> > > > > I have a chef recipe that builds new elasticsearch boxes so I spun up  
> > > > > a new box and tried to get it to join the cluster, but no dice - it's  
> > > > > like none of the other boxes exist. I've also tried to go back down to  
> > > > > 0.18.5 - no dice.
> > > > > 
> > > > > Is there a way I can point a new box at that master directly and say:  
> > > > > "Hey you're a slave, there is the master."
> > > > > 
> > > > > I'm a bit of a loss here - and don't want to admit defeat - but I'm a  
> > > > > little lost here.
> > > > > 
> > > > > It's a system that we're building and I CAN lose the data, but I  
> > > > > really want to understand:
> > > > > 
> > > > > 1. Why this happened.
> > > > > 2. How I can recover from this.
> > > > > 3. How I can prevent this from happening in the future.
> > > > > 
> > > > > Thanks in advance if anybody can point to something I will greatly  
> > > > > appreciate it.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 4, 2012, 11:16am UTC](https://discuss.elastic.co/t/split-brain/6256/7 "2012-01-04T11:16:56Z")

</div>

Interesting!

On Wed, Jan 4, 2012 at 10:26 AM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca) wrote:

> FYI - heard from Rackspace:
> 
> "We did recently implement multicast filtering on our new XenServer  
> Linux deployments. This was originally the intended design as having  
> multicast between all customers in the same huddles can be  
> problematic. I apologize that you used this as a feature before it was  
> blocked, but I feel it may ultimately the best with multicast  
> filtered.
> 
> Your timeline corresponds exactly with when I heard that the changes  
> were being rolled out."
> 
> Oh well - makes sense now.
> 
> On Fri, Dec 30, 2011 at 2:42 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)  
> wrote:
> 
> > Yeah - it was working great - all configs are in a git repo and pushed  
> > out via chef - been working great for a little over 3 weeks in  
> > production - and a couple weeks before in testing.
> > 
> > I have a ticket into Rackspace to see if they have changed something -  
> > but will just switch to unicast now.
> > 
> > Thanks for your help - will be updating my configs now.
> > 
> > On Fri, Dec 30, 2011 at 1:40 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > > Strange that multicast worked..., as far as I know its not supported in  
> > > Rackspace. Yea, you should put all the IPs in the unicast list, its  
> > > recommended if possible. If you go up to 3 nodes, then I think it make  
> > > sense  
> > > to have minimum master nodes set to 2, yea.
> > > 
> > > On Fri, Dec 30, 2011 at 9:39 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)  
> > > wrote:
> > > 
> > > > I was using multicast discovery and it was working great before -  
> > > > heres the log with extra debugging:
> > > > 
> > > > [http://d.pr/9pWC](http://d.pr/9pWC)
> > > > 
> > > > It looks like it didn't find anything at all.
> > > > 
> > > > So I setup unicast and added the master to the config:
> > > > 
> > > > [http://d.pr/NxBE](http://d.pr/NxBE)
> > > > 
> > > > And it seemed to work:
> > > > 
> > > > [http://d.pr/jxjW](http://d.pr/jxjW)  
> > > > [http://d.pr/Ubzi](http://d.pr/Ubzi)
> > > > 
> > > > I can switch to using unicast - that's no problem - just not sure why  
> > > > that happened - maybe Rackspace made some network changes - not sure  
> > > > why it worked for almost a month and suddenly stopped.
> > > > 
> > > > A couple questions.
> > > > 
> > > > 1. Should I put all of the IPs of all of the nodes in there?
> > > > 2. If I go up to a 3 node cluster - should I put  
> > > > "minimum\_master\_nodes" to  
> > > > 2?
> > > > 
> > > > Thanks for the tips Shay - really appreciate it.
> > > > 
> > > > On Fri, Dec 30, 2011 at 4:32 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > > 
> > > > > Heya:
> > > > > 
> > > > > First, regarding the split brain, it can obviously happen, especially  
> > > > > with 2  
> > > > > servers. If the network gets disconnected between the two for  
> > > > > example.  
> > > > > In  
> > > > > this case, you will have two separate one node cluster and you will  
> > > > > need  
> > > > > to  
> > > > > resolve it yourself by restarting one of them. You should see in the  
> > > > > logs  
> > > > > the fact that one node got disconnected from the other. If you had a  
> > > > > larger  
> > > > > cluster you could have defined "minimum\_master\_nodes" parameter to  
> > > > > reduce  
> > > > > the chances of it happening.
> > > > > 
> > > > > I am not sure why when you restart the node its not finding the other  
> > > > > node.  
> > > > > I am assuming you are using unicast discovery, are you sure its  
> > > > > configured  
> > > > > properly? You can set discovery: TRACE in the logging.yml file to see  
> > > > > which  
> > > > > nodes it tries to ping and what the status of that is.
> > > > > 
> > > > > -shay.banon
> > > > > 
> > > > > On Fri, Dec 30, 2011 at 5:22 AM, Darron Froese \<darron@nonfiction.ca
> > 
> > > > > wrote:
> > > > > 
> > > > > > I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5  
> > > > > > for  
> > > > > > the last 3 weeks. It's been working great and we've had several  
> > > > > > million records inserted and deleted in that time.
> > > > > > 
> > > > > > Here is the config:
> > > > > > 
> > > > > > [http://d.pr/WbU7](http://d.pr/WbU7)
> > > > > > 
> > > > > > Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
> > > > > > boxes to log to syslog as well) - it appears that something didn't  
> > > > > > work so well during the upgrade and I was left with 2 boxes both  
> > > > > > thinking that they're masters.
> > > > > > 
> > > > > > Here are the logs from the boxes during the upgrade:
> > > > > > 
> > > > > > [http://d.pr/HHbC](http://d.pr/HHbC)  
> > > > > > [http://d.pr/n3jk](http://d.pr/n3jk)
> > > > > > 
> > > > > > And then from the next day:
> > > > > > 
> > > > > > [http://d.pr/QEkU](http://d.pr/QEkU)  
> > > > > > [http://d.pr/I7fo](http://d.pr/I7fo)
> > > > > > 
> > > > > > I tried to get them to re-connect, but I couldn't get anything to  
> > > > > > work  
> > > > > > correctly - they were both completely separate.
> > > > > > 
> > > > > > I now have a single box with the correct index:
> > > > > > 
> > > > > > [http://d.pr/4TTq](http://d.pr/4TTq)
> > > > > > 
> > > > > > I have a chef recipe that builds new elasticsearch boxes so I spun  
> > > > > > up  
> > > > > > a new box and tried to get it to join the cluster, but no dice -  
> > > > > > it's  
> > > > > > like none of the other boxes exist. I've also tried to go back down  
> > > > > > to  
> > > > > > 0.18.5 - no dice.
> > > > > > 
> > > > > > Is there a way I can point a new box at that master directly and  
> > > > > > say:  
> > > > > > "Hey you're a slave, there is the master."
> > > > > > 
> > > > > > I'm a bit of a loss here - and don't want to admit defeat - but I'm  
> > > > > > a  
> > > > > > little lost here.
> > > > > > 
> > > > > > It's a system that we're building and I CAN lose the data, but I  
> > > > > > really want to understand:
> > > > > > 
> > > > > > 1. Why this happened.
> > > > > > 2. How I can recover from this.
> > > > > > 3. How I can prevent this from happening in the future.
> > > > > > 
> > > > > > Thanks in advance if anybody can point to something I will greatly  
> > > > > > appreciate it.

---

<div class="post-metadata">

**Author:** ![Stanislas\_Polu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stanislas_polu/32/2873_2.png) [@Stanislas\_Polu](https://discuss.elastic.co/u/Stanislas_Polu)\
**Post date:** [January 4, 2012, 11:31am UTC](https://discuss.elastic.co/t/split-brain/6256/8 "2012-01-04T11:31:00Z")

</div>

Very interesting! Thanks!

-stan

--  
Stanislas Polu  
Mo: +33 6 83 71 90 04 | Tw: @spolu | [http://teleportd.com](http://teleportd.com) | Realtime Photo  
Search

On Wed, Jan 4, 2012 at 12:16 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Interesting!
> 
> On Wed, Jan 4, 2012 at 10:26 AM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)wrote:
> 
> > FYI - heard from Rackspace:
> > 
> > "We did recently implement multicast filtering on our new XenServer  
> > Linux deployments. This was originally the intended design as having  
> > multicast between all customers in the same huddles can be  
> > problematic. I apologize that you used this as a feature before it was  
> > blocked, but I feel it may ultimately the best with multicast  
> > filtered.
> > 
> > Your timeline corresponds exactly with when I heard that the changes  
> > were being rolled out."
> > 
> > Oh well - makes sense now.
> > 
> > On Fri, Dec 30, 2011 at 2:42 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)  
> > wrote:
> > 
> > > Yeah - it was working great - all configs are in a git repo and pushed  
> > > out via chef - been working great for a little over 3 weeks in  
> > > production - and a couple weeks before in testing.
> > > 
> > > I have a ticket into Rackspace to see if they have changed something -  
> > > but will just switch to unicast now.
> > > 
> > > Thanks for your help - will be updating my configs now.
> > > 
> > > On Fri, Dec 30, 2011 at 1:40 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > 
> > > > Strange that multicast worked..., as far as I know its not supported in  
> > > > Rackspace. Yea, you should put all the IPs in the unicast list, its  
> > > > recommended if possible. If you go up to 3 nodes, then I think it make  
> > > > sense  
> > > > to have minimum master nodes set to 2, yea.
> > > > 
> > > > On Fri, Dec 30, 2011 at 9:39 PM, Darron Froese [darron@nonfiction.ca](mailto:darron@nonfiction.ca)  
> > > > wrote:
> > > > 
> > > > > I was using multicast discovery and it was working great before -  
> > > > > heres the log with extra debugging:
> > > > > 
> > > > > [http://d.pr/9pWC](http://d.pr/9pWC)
> > > > > 
> > > > > It looks like it didn't find anything at all.
> > > > > 
> > > > > So I setup unicast and added the master to the config:
> > > > > 
> > > > > [http://d.pr/NxBE](http://d.pr/NxBE)
> > > > > 
> > > > > And it seemed to work:
> > > > > 
> > > > > [http://d.pr/jxjW](http://d.pr/jxjW)  
> > > > > [http://d.pr/Ubzi](http://d.pr/Ubzi)
> > > > > 
> > > > > I can switch to using unicast - that's no problem - just not sure why  
> > > > > that happened - maybe Rackspace made some network changes - not sure  
> > > > > why it worked for almost a month and suddenly stopped.
> > > > > 
> > > > > A couple questions.
> > > > > 
> > > > > 1. Should I put all of the IPs of all of the nodes in there?
> > > > > 2. If I go up to a 3 node cluster - should I put  
> > > > > "minimum\_master\_nodes" to  
> > > > > 2?
> > > > > 
> > > > > Thanks for the tips Shay - really appreciate it.
> > > > > 
> > > > > On Fri, Dec 30, 2011 at 4:32 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > > > 
> > > > > > Heya:
> > > > > > 
> > > > > > First, regarding the split brain, it can obviously happen,  
> > > > > > especially  
> > > > > > with 2  
> > > > > > servers. If the network gets disconnected between the two for  
> > > > > > example.  
> > > > > > In  
> > > > > > this case, you will have two separate one node cluster and you will  
> > > > > > need  
> > > > > > to  
> > > > > > resolve it yourself by restarting one of them. You should see in the  
> > > > > > logs  
> > > > > > the fact that one node got disconnected from the other. If you had a  
> > > > > > larger  
> > > > > > cluster you could have defined "minimum\_master\_nodes" parameter to  
> > > > > > reduce  
> > > > > > the chances of it happening.
> > > > > > 
> > > > > > I am not sure why when you restart the node its not finding the  
> > > > > > other  
> > > > > > node.  
> > > > > > I am assuming you are using unicast discovery, are you sure its  
> > > > > > configured  
> > > > > > properly? You can set discovery: TRACE in the logging.yml file to  
> > > > > > see  
> > > > > > which  
> > > > > > nodes it tries to ping and what the status of that is.
> > > > > > 
> > > > > > -shay.banon
> > > > > > 
> > > > > > On Fri, Dec 30, 2011 at 5:22 AM, Darron Froese \<  
> > > > > > darron@nonfiction.ca\>  
> > > > > > wrote:
> > > > > > 
> > > > > > > I had a cluster of 2 x 2GB Rackspace cloud boxes running ES 0.18.5  
> > > > > > > for  
> > > > > > > the last 3 weeks. It's been working great and we've had several  
> > > > > > > million records inserted and deleted in that time.
> > > > > > > 
> > > > > > > Here is the config:
> > > > > > > 
> > > > > > > [http://d.pr/WbU7](http://d.pr/WbU7)
> > > > > > > 
> > > > > > > Yesterday I updated the boxes to 0.18.6 (and was trying to get the  
> > > > > > > boxes to log to syslog as well) - it appears that something didn't  
> > > > > > > work so well during the upgrade and I was left with 2 boxes both  
> > > > > > > thinking that they're masters.
> > > > > > > 
> > > > > > > Here are the logs from the boxes during the upgrade:
> > > > > > > 
> > > > > > > [http://d.pr/HHbC](http://d.pr/HHbC)  
> > > > > > > [http://d.pr/n3jk](http://d.pr/n3jk)
> > > > > > > 
> > > > > > > And then from the next day:
> > > > > > > 
> > > > > > > [http://d.pr/QEkU](http://d.pr/QEkU)  
> > > > > > > [http://d.pr/I7fo](http://d.pr/I7fo)
> > > > > > > 
> > > > > > > I tried to get them to re-connect, but I couldn't get anything to  
> > > > > > > work  
> > > > > > > correctly - they were both completely separate.
> > > > > > > 
> > > > > > > I now have a single box with the correct index:
> > > > > > > 
> > > > > > > [http://d.pr/4TTq](http://d.pr/4TTq)
> > > > > > > 
> > > > > > > I have a chef recipe that builds new elasticsearch boxes so I spun  
> > > > > > > up  
> > > > > > > a new box and tried to get it to join the cluster, but no dice -  
> > > > > > > it's  
> > > > > > > like none of the other boxes exist. I've also tried to go back  
> > > > > > > down to  
> > > > > > > 0.18.5 - no dice.
> > > > > > > 
> > > > > > > Is there a way I can point a new box at that master directly and  
> > > > > > > say:  
> > > > > > > "Hey you're a slave, there is the master."
> > > > > > > 
> > > > > > > I'm a bit of a loss here - and don't want to admit defeat - but  
> > > > > > > I'm a  
> > > > > > > little lost here.
> > > > > > > 
> > > > > > > It's a system that we're building and I CAN lose the data, but I  
> > > > > > > really want to understand:
> > > > > > > 
> > > > > > > 1. Why this happened.
> > > > > > > 2. How I can recover from this.
> > > > > > > 3. How I can prevent this from happening in the future.
> > > > > > > 
> > > > > > > Thanks in advance if anybody can point to something I will greatly  
> > > > > > > appreciate it.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:43am UTC](https://discuss.elastic.co/t/split-brain/6256/9 "2017-07-06T03:43:53Z")

</div>


