# Nodes randomly disconnected

**URL:** <https://discuss.elastic.co/t/nodes-randomly-disconnected/16756>\
**Category:** Elasticsearch\
**Created:** [April 2, 2014, 4:40am UTC](https://discuss.elastic.co/t/nodes-randomly-disconnected/16756 "2014-04-02T04:40:10Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Hans\_Krijger](https://avatars.discourse-cdn.com/v4/letter/h/a4c791/32.png) [@Hans\_Krijger](https://discuss.elastic.co/u/Hans_Krijger)\
**Post date:** [April 2, 2014, 4:40am UTC](https://discuss.elastic.co/t/nodes-randomly-disconnected/16756/1 "2014-04-02T04:40:10Z")

</div>

We have a cluster running 1.0.0 in Azure using unicast discovery. Recently  
we started seeing exceptions like these in the logs:

[2014-04-01 21:40:22,720][DEBUG][action.admin.indices.status] [ES2PROD-M01]  
[usg-2014-03-04][4], node[3cCeFKJrTMWaIhE3R6tlZA], [P], s[STARTED]: Failed  
to execute [  
org.elasticsearch.action.admin.indices.status.IndicesStatusRequest@2c06e67[https://github.com/org.elasticsearch.action.admin.indices.status.IndicesStatusRequest/elasticsearch/commit/2c06e675](https://github.com/org.elasticsearch.action.admin.indices.status.IndicesStatusRequest/elasticsearch/commit/2c06e675)  
]  
org.elasticsearch.transport.NodeDisconnectedException:  
[ES2PROD-D07][inet[/10.0.64.68:9300]][indices/status/s] disconnected

In this case D07 is still up and running. After several dozen of these  
exceptions, D07 is disconnected:

[2014-04-01 21:40:24,096][INFO][cluster.service] [ES2PROD-M01] removed  
{[ES2PROD-D07][3cCeFKJrTMWaIhE3R6tlZA][es2prod-d07][inet[/10.0.64.68:9300]]{master=false},},  
reason:  
zen-disco-node\_failed([ES2PROD-D07][3cCeFKJrTMWaIhE3R6tlZA][es2prod-d07][inet[/10.0.64.68:9300]]{master=false}),  
reason transport disconnected (with verified connect)

Four seconds later the same node is added back:

[2014-04-01 21:40:28,712][INFO][cluster.service] [ES2PROD-M01] added  
{[ES2PROD-D07][3cCeFKJrTMWaIhE3R6tlZA][es2prod-d07][inet[/10.0.64.68:9300]]{master=false},},  
reason: zen-disco-receive(join from  
node[[ES2PROD-D07][3cCeFKJrTMWaIhE3R6tlZA][es2prod-d07][inet[/10.0.64.68:9300]]{master=false}])

In the mean time the cluster goes yellow and starts recovery. This does not  
seem like a timeout type of issue since it happens so quickly, and then the  
disconnected node is added right back.

Any ideas how we can get more info on the root cause and avoid this from  
happening?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/df768c5e-5833-42b5-a804-b7d07f51996b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/df768c5e-5833-42b5-a804-b7d07f51996b%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Binh\_Ly\_2](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@Binh\_Ly\_2](https://discuss.elastic.co/u/Binh_Ly_2)\
**Post date:** [April 4, 2014, 12:52pm UTC](https://discuss.elastic.co/t/nodes-randomly-disconnected/16756/2 "2014-04-04T12:52:39Z")

</div>

This could be caused be unreliable network connectivity, or if your nodes  
are somehow overloaded and can't respond to other nodes in a timely manner.  
In a cloud environment, this could happen more often on the lowest tier  
instances. If indeed network connectivity is the cause, you can increase  
ping\_timeout a bit to accommodate.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ee8a8a21-6500-4979-b469-16158772ccd6%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ee8a8a21-6500-4979-b469-16158772ccd6%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![BornGenius](https://avatars.discourse-cdn.com/v4/letter/b/e19b73/32.png) [@BornGenius](https://discuss.elastic.co/u/BornGenius)\
**Post date:** [April 4, 2014, 1:15pm UTC](https://discuss.elastic.co/t/nodes-randomly-disconnected/16756/3 "2014-04-04T13:15:45Z")

</div>

Hi Hans,

We were also facing this issue and the reason was that there were spikes in the network connectivity between nodes due to which the master was not able to discover the data nodes ,during zen discovery and hence removed those nodes and brought them back. You can try to increase the discovery.zen.ping.timeout that defaults to 3 seconds ,this should fix the issue.

All the best..

AryanJ

"Give users what they actually want, not what they say they want. And whatever you do, don’t give them new features just because your competitors have them!!!!!" – Kathy Sierra

---

<div class="post-metadata">

**Author:** ![Anil\_Karaka](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@Anil\_Karaka](https://discuss.elastic.co/u/Anil_Karaka)\
**Post date:** [April 1, 2015, 6:09am UTC](https://discuss.elastic.co/t/nodes-randomly-disconnected/16756/4 "2015-04-01T06:09:10Z")

</div>

Hi AryanJ

What value did you have for ping.timeout, Our cluster is on AWS and we are  
facing this problem for a long time. Each node leaves the cluster at least  
once except for master in a day.  
I am keeping it for 10secs. Can I set it to even bigger value?

Thanks.

On Friday, April 4, 2014 at 6:45:46 PM UTC+5:30, AryanJ wrote:

> Hi Hans,
> 
> We were also facing this issue and the reason was that there were spikes  
> in  
> the network connectivity between nodes due to which the master was not  
> able  
> to discover the data nodes ,during zen discovery and hence removed those  
> nodes and brought them back. You can try to increase the  
> discovery.zen.ping.timeout that defaults to 3 seconds ,this should fix the  
> issue.
> 
> All the best..
> 
> AryanJ
> 
> "Give users what they actually want, not what they say they want. And  
> whatever you do, don’t give them new features just because your  
> competitors  
> have them!!!!!" – Kathy Sierra
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Nodes-randomly-disconnected-tp4053290p4053514.html](http://elasticsearch-users.115913.n3.nabble.com/Nodes-randomly-disconnected-tp4053290p4053514.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/fd2ce05c-3054-4943-8df0-5eea643db20e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/fd2ce05c-3054-4943-8df0-5eea643db20e%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:22am UTC](https://discuss.elastic.co/t/nodes-randomly-disconnected/16756/5 "2017-07-06T00:22:28Z")

</div>


