# Cluster nodes doesn't reconnect

**URL:** <https://discuss.elastic.co/t/cluster-nodes-doesnt-reconnect/12691>\
**Category:** Elasticsearch\
**Created:** [July 8, 2013, 8:31am UTC](https://discuss.elastic.co/t/cluster-nodes-doesnt-reconnect/12691 "2013-07-08T08:31:30Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![planckiii](https://avatars.discourse-cdn.com/v4/letter/p/5fc32e/32.png) [@planckiii](https://discuss.elastic.co/u/planckiii)\
**Post date:** [July 8, 2013, 8:31am UTC](https://discuss.elastic.co/t/cluster-nodes-doesnt-reconnect/12691/1 "2013-07-08T08:31:30Z")

</div>

Hi,  
I have setup of elasticsearch 0.90.0 with two nodes, each one on different  
data center. From time to time cluster status goes "yellow":  
{  
"cluster\_name" : "my\_cluster",  
"status" : "yellow",  
"timed\_out" : false,  
"number\_of\_nodes" : 1,  
"number\_of\_data\_nodes" : 1,  
"active\_primary\_shards" : 10,  
"active\_shards" : 10,  
"relocating\_shards" : 0,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 10  
}

Probably it's some kind of short network freeze between es nodes, because  
test which i run (at intervals of 15s):

# nc -z -v -w 2 second\_node 9300

Fri Jul 5 21:31:07 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
succeeded!  
Fri Jul 5 21:31:22 CEST 2013 nc: connect to second\_node port 9300 (tcp)  
timed out: Operation now in progress  
Fri Jul 5 21:31:39 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
succeeded!

What is strange for me, es nodes couldn't reconnect and i have that kind of  
errors:  
first\_node: [https://gist.github.com/planckiii/5947058](https://gist.github.com/planckiii/5947058)  
second\_node: [https://gist.github.com/planckiii/5947068](https://gist.github.com/planckiii/5947068)

My Transport and Discover configurations on both nodes:  
################################## Transport  
##################################  
transport.tcp.connect.timeout: 5s  
################################## Discovery  
##################################  
discovery.zen.ping\_timeout: 5s  
discovery.zen.fd.ping\_timeout: 30s  
discovery.zen.fd.ping\_interval: 2s  
discovery.zen.fd.ping\_retries: 10  
discovery.zen.ping.multicast.enabled: false

After one node reset everything goes OK and cluster is properly balanced:  
{  
"cluster\_name" : "my\_cluster",  
"status" : "green",  
"timed\_out" : false,  
"number\_of\_nodes" : 2,  
"number\_of\_data\_nodes" : 2,  
"active\_primary\_shards" : 10,  
"active\_shards" : 20,  
"relocating\_shards" : 0,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 0  
}

Any ideas what could be wrong with my setup ?

Regards

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![imdhmd](https://avatars.discourse-cdn.com/v4/letter/i/d07c76/32.png) [@imdhmd](https://discuss.elastic.co/u/imdhmd)\
**Post date:** [July 8, 2013, 3:06pm UTC](https://discuss.elastic.co/t/cluster-nodes-doesnt-reconnect/12691/2 "2013-07-08T15:06:56Z")

</div>

Hey planckiii,  
\*  
\*  
\*\>\> \*it's some kind of short network freeze between es nodes  
Based on your above statement, i assume that connection between the two  
nodes is weak/unreliable.  
In that case you should see if it helps increasing ping timeout, retries  
values. Also, do you have any specific reason to disable multicast?

Also, if you have head plugin installed on both the nodes and when this  
happens again, could you bring up the head site pages of both the nodes and  
see if they are both becoming master and hence are not able to form back  
into one cluster. This would be a case of split-brain problem.  
To resolve this, you have two choices:

- Exclude one of the nodes from becoming master, using the setting  
`node.master: false`

- Use zookeeper plugin to externalize master election  
([GitHub - sonian/elasticsearch-zookeeper](https://github.com/sonian/elasticsearch-zookeeper))

- Imdad

On Monday, July 8, 2013 2:01:30 PM UTC+5:30, planckiii wrote:

> Hi,  
> I have setup of elasticsearch 0.90.0 with two nodes, each one on different  
> data center. From time to time cluster status goes "yellow":  
> {  
> "cluster\_name" : "my\_cluster",  
> "status" : "yellow",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 1,  
> "number\_of\_data\_nodes" : 1,  
> "active\_primary\_shards" : 10,  
> "active\_shards" : 10,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 0,  
> "unassigned\_shards" : 10  
> }
> 
> Probably it's some kind of short network freeze between es nodes, because  
> test which i run (at intervals of 15s):
> 
> # nc -z -v -w 2 second\_node 9300
> 
> Fri Jul 5 21:31:07 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
> succeeded!  
> Fri Jul 5 21:31:22 CEST 2013 nc: connect to second\_node port 9300 (tcp)  
> timed out: Operation now in progress  
> Fri Jul 5 21:31:39 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
> succeeded!
> 
> What is strange for me, es nodes couldn't reconnect and i have that kind  
> of errors:  
> first\_node: [first\_node, primary data center: · GitHub](https://gist.github.com/planckiii/5947058)  
> second\_node: [second\_node, secondary data center: · GitHub](https://gist.github.com/planckiii/5947068)
> 
> My Transport and Discover configurations on both nodes:  
> ################################## Transport  
> ##################################  
> transport.tcp.connect.timeout: 5s  
> ################################## Discovery  
> ##################################  
> discovery.zen.ping\_timeout: 5s  
> discovery.zen.fd.ping\_timeout: 30s  
> discovery.zen.fd.ping\_interval: 2s  
> discovery.zen.fd.ping\_retries: 10  
> discovery.zen.ping.multicast.enabled: false
> 
> After one node reset everything goes OK and cluster is properly balanced:  
> {  
> "cluster\_name" : "my\_cluster",  
> "status" : "green",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 2,  
> "number\_of\_data\_nodes" : 2,  
> "active\_primary\_shards" : 10,  
> "active\_shards" : 20,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 0,  
> "unassigned\_shards" : 0  
> }
> 
> Any ideas what could be wrong with my setup ?
> 
> Regards

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![planckiii](https://avatars.discourse-cdn.com/v4/letter/p/5fc32e/32.png) [@planckiii](https://discuss.elastic.co/u/planckiii)\
**Post date:** [July 9, 2013, 8:58pm UTC](https://discuss.elastic.co/t/cluster-nodes-doesnt-reconnect/12691/3 "2013-07-09T20:58:57Z")

</div>

W dniu poniedziałek, 8 lipca 2013 17:06:56 UTC+2 użytkownik Imdad Ahmed  
napisał:

> Hey planckiii,

Hi, thanks for quick rep 🙂

> - 
> - 
> 
> \*\>\> \*it's some kind of short network freeze between es nodes  
> Based on your above statement, i assume that connection between the two  
> nodes is weak/unreliable.  
> In that case you should see if it helps increasing ping timeout, retries  
> values. Also, do you have any specific reason to disable multicast?

I tried with that:  
discovery.zen.ping\_timeout: 5s  
discovery.zen.fd.ping\_timeout: 30s  
discovery.zen.fd.ping\_interval: 2s  
discovery.zen.fd.ping\_retries: 10  
Network freezes shouldn't be higher than few seconds so in theory that  
should bo OK. About multicast - that are VMs behind internal NAT, multicast  
couldn't work outside that NAT ☹

> Also, if you have head plugin installed on both the nodes and when this  
> happens again, could you bring up the head site pages of both the nodes and  
> see if they are both becoming master and hence are not able to form back  
> into one cluster. This would be a case of split-brain problem.  
> To resolve this, you have two choices:
> 
> - Exclude one of the nodes from becoming master, using the setting  
> `node.master: false`
> - Use zookeeper plugin to externalize master election (  
> [GitHub - sonian/elasticsearch-zookeeper](https://github.com/sonian/elasticsearch-zookeeper))

I agree that probably it's split-brain problem after disconnect - but there  
isn't any info in log-s about that ☹ I will check that on next failure.  
Thank's for advice - i will update status of that problem.

> - Imdad
> 
> On Monday, July 8, 2013 2:01:30 PM UTC+5:30, planckiii wrote:
> 
> > Hi,  
> > I have setup of elasticsearch 0.90.0 with two nodes, each one on  
> > different data center. From time to time cluster status goes "yellow":  
> > {  
> > "cluster\_name" : "my\_cluster",  
> > "status" : "yellow",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 1,  
> > "number\_of\_data\_nodes" : 1,  
> > "active\_primary\_shards" : 10,  
> > "active\_shards" : 10,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 0,  
> > "unassigned\_shards" : 10  
> > }
> > 
> > Probably it's some kind of short network freeze between es nodes, because  
> > test which i run (at intervals of 15s):
> > 
> > # nc -z -v -w 2 second\_node 9300
> > 
> > Fri Jul 5 21:31:07 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
> > succeeded!  
> > Fri Jul 5 21:31:22 CEST 2013 nc: connect to second\_node port 9300 (tcp)  
> > timed out: Operation now in progress  
> > Fri Jul 5 21:31:39 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
> > succeeded!
> > 
> > What is strange for me, es nodes couldn't reconnect and i have that kind  
> > of errors:  
> > first\_node: [first\_node, primary data center: · GitHub](https://gist.github.com/planckiii/5947058)  
> > second\_node: [second\_node, secondary data center: · GitHub](https://gist.github.com/planckiii/5947068)
> > 
> > My Transport and Discover configurations on both nodes:  
> > ################################## Transport  
> > ##################################  
> > transport.tcp.connect.timeout: 5s  
> > ################################## Discovery  
> > ##################################  
> > discovery.zen.ping\_timeout: 5s  
> > discovery.zen.fd.ping\_timeout: 30s  
> > discovery.zen.fd.ping\_interval: 2s  
> > discovery.zen.fd.ping\_retries: 10  
> > discovery.zen.ping.multicast.enabled: false
> > 
> > After one node reset everything goes OK and cluster is properly balanced:  
> > {  
> > "cluster\_name" : "my\_cluster",  
> > "status" : "green",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 2,  
> > "number\_of\_data\_nodes" : 2,  
> > "active\_primary\_shards" : 10,  
> > "active\_shards" : 20,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 0,  
> > "unassigned\_shards" : 0  
> > }
> > 
> > Any ideas what could be wrong with my setup ?
> > 
> > Regards

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [July 10, 2013, 8:45am UTC](https://discuss.elastic.co/t/cluster-nodes-doesnt-reconnect/12691/4 "2013-07-10T08:45:08Z")

</div>

Hey,

running a cluster in a cross-data-center setup is generally not a good  
idea. For example if you are using replicas, every indexing operation goes  
to both data centers and returns only when both are finished. This will  
introduce high latency to your system. The same is true for searches going  
to several shards, which a shared across both data centers. If you can, try  
to build a different sync mechanism than this kind of high-risk setup  
(writing data to both systems, which are an independent cluster for itself,  
maybe?).

--Alex

On Tue, Jul 9, 2013 at 10:58 PM, planckiii [planckiii@gmail.com](mailto:planckiii@gmail.com) wrote:

> W dniu poniedziałek, 8 lipca 2013 17:06:56 UTC+2 użytkownik Imdad Ahmed  
> napisał:
> 
> > Hey planckiii,
> 
> Hi, thanks for quick rep 🙂
> 
> > - 
> > - 
> > 
> > \*\>\> \*it's some kind of short network freeze between es nodes  
> > Based on your above statement, i assume that connection between the two  
> > nodes is weak/unreliable.  
> > In that case you should see if it helps increasing ping timeout, retries  
> > values. Also, do you have any specific reason to disable multicast?
> 
> I tried with that:  
> discovery.zen.ping\_timeout: 5s  
> discovery.zen.fd.ping\_timeout: 30s  
> discovery.zen.fd.ping\_interval: 2s  
> discovery.zen.fd.ping\_retries: 10  
> Network freezes shouldn't be higher than few seconds so in theory that  
> should bo OK. About multicast - that are VMs behind internal NAT, multicast  
> couldn't work outside that NAT ☹
> 
> > Also, if you have head plugin installed on both the nodes and when this  
> > happens again, could you bring up the head site pages of both the nodes and  
> > see if they are both becoming master and hence are not able to form back  
> > into one cluster. This would be a case of split-brain problem.  
> > To resolve this, you have two choices:
> > 
> > - Exclude one of the nodes from becoming master, using the setting  
> > `node.master: false`
> > - Use zookeeper plugin to externalize master election (  
> > [https://github.com/sonian/\*\*elasticsearch-zookeeper](https://github.com/sonian/**elasticsearch-zookeeper)[https://github.com/sonian/elasticsearch-zookeeper](https://github.com/sonian/elasticsearch-zookeeper)  
> > )
> 
> I agree that probably it's split-brain problem after disconnect - but  
> there isn't any info in log-s about that ☹ I will check that on next  
> failure. Thank's for advice - i will update status of that problem.
> 
> > - Imdad
> > 
> > On Monday, July 8, 2013 2:01:30 PM UTC+5:30, planckiii wrote:
> > 
> > > Hi,  
> > > I have setup of elasticsearch 0.90.0 with two nodes, each one on  
> > > different data center. From time to time cluster status goes "yellow":  
> > > {  
> > > "cluster\_name" : "my\_cluster",  
> > > "status" : "yellow",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 1,  
> > > "number\_of\_data\_nodes" : 1,  
> > > "active\_primary\_shards" : 10,  
> > > "active\_shards" : 10,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 0,  
> > > "unassigned\_shards" : 10  
> > > }
> > > 
> > > Probably it's some kind of short network freeze between es nodes,  
> > > because test which i run (at intervals of 15s):
> > > 
> > > # nc -z -v -w 2 second\_node 9300
> > > 
> > > Fri Jul 5 21:31:07 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
> > > succeeded!  
> > > Fri Jul 5 21:31:22 CEST 2013 nc: connect to second\_node port 9300 (tcp)  
> > > timed out: Operation now in progress  
> > > Fri Jul 5 21:31:39 CEST 2013 Connection to second\_node 9300 port [tcp/_]  
> > > succeeded!
> > > 
> > > What is strange for me, es nodes couldn't reconnect and i have that kind  
> > > of errors:  
> > > first\_node: [https://gist](https://gist).\*\*[github.com/planckiii/5947058](http://github.com/planckiii/5947058)[https://gist.github.com/planckiii/5947058](https://gist.github.com/planckiii/5947058)  
> > > second\_node: [https://gist](https://gist).\*\*[github.com/planckiii/5947068](http://github.com/planckiii/5947068)[https://gist.github.com/planckiii/5947068](https://gist.github.com/planckiii/5947068)
> > > 
> > > My Transport and Discover configurations on both nodes:  
> > > ##############################**#### Transport  
> > > ##############################**####  
> > > transport.tcp.connect.timeout: 5s  
> > > ##############################**#### Discovery  
> > > ##############################**####  
> > > discovery.zen.ping\_timeout: 5s  
> > > discovery.zen.fd.ping\_timeout: 30s  
> > > discovery.zen.fd.ping\_\*\*interval: 2s  
> > > discovery.zen.fd.ping\_retries: 10  
> > > discovery.zen.ping.multicast.\*\*enabled: false
> > > 
> > > After one node reset everything goes OK and cluster is properly balanced:  
> > > {  
> > > "cluster\_name" : "my\_cluster",  
> > > "status" : "green",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 2,  
> > > "number\_of\_data\_nodes" : 2,  
> > > "active\_primary\_shards" : 10,  
> > > "active\_shards" : 20,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 0,  
> > > "unassigned\_shards" : 0  
> > > }
> > > 
> > > Any ideas what could be wrong with my setup ?
> > > 
> > > Regards
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:27am UTC](https://discuss.elastic.co/t/cluster-nodes-doesnt-reconnect/12691/5 "2017-07-06T02:27:17Z")

</div>


