# Nodes keep disconnecting from cluster at random

**URL:** <https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592>\
**Category:** Elasticsearch\
**Created:** [November 29, 2017, 12:22pm UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592 "2017-11-29T12:22:03Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Paul\_Zaltsman](https://avatars.discourse-cdn.com/v4/letter/p/b9e5f3/32.png) [@Paul\_Zaltsman](https://discuss.elastic.co/u/Paul_Zaltsman)\
**Post date:** [November 29, 2017, 12:22pm UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/1 "2017-11-29T12:22:04Z")

</div>

Hi, i have an issue with a 5.5.2 ES cluster.  
We see a lot of disconnects between our nodes.  
Even sending a simple \_nodes query causes random disconnections. In this example. Only the node i've asked answered. Rest failed.

We have 10 Data nodes and 3 Masters in a cluster. 40 indexes with 300shards. 2TB size cluster All in the same Google datacenter.  
Os is centos7.3.1611 core.  
No network congestion or any unusual behavior and yet this a result of a GET \_nodes query.  
At this point some nodes disconnect and the cluster becomes yellow with unassigned shards.  
All servers are 5.5.2 jvm1.8.0\_151.

id v  
4tgD 5.5.2  
9UW6 5.5.2  
8scZ 5.5.2  
rFIc 5.5.2  
XhZN 5.5.2  
vNyd 5.5.2  
\_cmX 5.5.2  
UI-s 5.5.2  
eQxy 5.5.2  
hgyP 5.5.2  
xz8x 5.5.2  
1bZd 5.5.2  
cyJg 5.5.2

---

<div class="post-metadata">

**Author:** ![Paul\_Zaltsman](https://avatars.discourse-cdn.com/v4/letter/p/b9e5f3/32.png) [@Paul\_Zaltsman](https://discuss.elastic.co/u/Paul_Zaltsman)\
**Post date:** [November 29, 2017, 12:24pm UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/2 "2017-11-29T12:24:45Z")

</div>

Didn't let me paste the whole node info.  
[https://pastebin.com/xLP99vXM](https://pastebin.com/xLP99vXM)

---

<div class="post-metadata">

**Author:** ![Paul\_Zaltsman](https://avatars.discourse-cdn.com/v4/letter/p/b9e5f3/32.png) [@Paul\_Zaltsman](https://discuss.elastic.co/u/Paul_Zaltsman)\
**Post date:** [November 29, 2017, 12:27pm UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/3 "2017-11-29T12:27:49Z")

</div>

Stack trace of one of the nodes while running this  
[https://pastebin.com/wXs5SgNp](https://pastebin.com/wXs5SgNp)

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [December 1, 2017, 8:42am UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/4 "2017-12-01T08:42:01Z")

</div>

Hey,

thats an odd one. The exception means, that a serialized data stream could not be read from another node. This should not happen, especially not when you have the same Elasticsearch versions everywhere.

Can you run

```auto
GET _cat/master
GET _cat/nodes?v&h=id,name,version,jdk,node.role,master

```

---

<div class="post-metadata">

**Author:** ![Paul\_Zaltsman](https://avatars.discourse-cdn.com/v4/letter/p/b9e5f3/32.png) [@Paul\_Zaltsman](https://discuss.elastic.co/u/Paul_Zaltsman)\
**Post date:** [December 4, 2017, 10:20am UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/5 "2017-12-04T10:20:33Z")

</div>

Hi. Yes, i'm aware this shouldn't happen. But it still does.  
You may notice the cluster is bigger than i initially said, we are replacing some of the machines, you can ignore that.

\_cat/master

8scZVDYGTvaHA9B4EfCYCA 10.241.0.61 10.241.0.61 prod-es-ma-1

\_cat/nodes?v&h=id,name,version,jdk,node.role,master  
id name version jdk node.role master  
SG00 prod-es-dn-6a 5.5.2 1.8.0\_151 di -  
4tgD prod-es-dn-3 5.5.2 1.8.0\_151 di -  
rFIc prod-es-dn-5 5.5.2 1.8.0\_151 di -  
eoiz prod-es-dn-1a 5.5.2 1.8.0\_151 di -  
xz8x prod-es-dn-6 5.5.2 1.8.0\_151 di -  
XhZN prod-es-dn-10 5.5.2 1.8.0\_151 di -  
eQxy prod-es-dn-8 5.5.2 1.8.0\_151 di -  
8p\_p prod-es-dn-10a 5.5.2 1.8.0\_151 di -  
9UW6 prod-es-dn-9 5.5.2 1.8.0\_151 di -  
xHCC prod-es-dn-11a 5.5.2 1.8.0\_151 di -  
8scZ prod-es-ma-1 5.5.2 1.8.0\_151 m \*  
vNyd prod-es-ma-3 5.5.2 1.8.0\_151 m -  
w6qd prod-es-dn-5a 5.5.2 1.8.0\_151 di -  
QnKW prod-es-dn-2a 5.5.2 1.8.0\_151 di -  
1bZd prod-es-dn-2 5.5.2 1.8.0\_151 di -  
_cmX prod-es-dn-7 5.5.2 1.8.0\_151 di -  
aAA_ prod-es-dn-12a 5.5.2 1.8.0\_151 di -  
XYSt prod-es-dn-4a 5.5.2 1.8.0\_151 di -  
Bcf7 prod-es-dn-3a 5.5.2 1.8.0\_151 di -  
LBBL prod-es-dn-8a 5.5.2 1.8.0\_151 di -  
hJ11 prod-es-dn-9a 5.5.2 1.8.0\_151 di -  
BquO prod-es-dn-7a 5.5.2 1.8.0\_151 di -  
UI-s prod-es-dn-4 5.5.2 1.8.0\_151 di -  
cyJg prod-es-dn-1 5.5.2 1.8.0\_151 di -  
hgyP prod-es-ma-2 5.5.2 1.8.0\_151 m -

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 4, 2017, 10:52am UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/6 "2017-12-04T10:52:11Z")

</div>

Is there anything in the logs on the nodes that are disconnecting, e.g. long GC?

---

<div class="post-metadata">

**Author:** ![Paul\_Zaltsman](https://avatars.discourse-cdn.com/v4/letter/p/b9e5f3/32.png) [@Paul\_Zaltsman](https://discuss.elastic.co/u/Paul_Zaltsman)\
**Post date:** [December 4, 2017, 11:15am UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/7 "2017-12-04T11:15:41Z")

</div>

There is some GC on some of the machines, but this is inconsistent with the nodes that actually disconnect.  
You can see in the log i provided in my original post.

---

<div class="post-metadata">

**Author:** ![Paul\_Zaltsman](https://avatars.discourse-cdn.com/v4/letter/p/b9e5f3/32.png) [@Paul\_Zaltsman](https://discuss.elastic.co/u/Paul_Zaltsman)\
**Post date:** [December 7, 2017, 11:19am UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/8 "2017-12-07T11:19:45Z")

</div>

Bump.. Any idea in which direction to dig into this issue?  
Our cluster still has those errors...

[DEBUG][o.e.a.a.c.n.i.TransportNodesInfoAction] [prod-es-dn-7a] failed to execute on node [SG00OvqdRIaUTz\_wUlHMFg]  
org.elasticsearch.transport.RemoteTransportException: [Failed to deserialize response of type [org.elasticsearch.action.admin.cluster.node.info.NodeInfo]]  
Caused by: org.elasticsearch.transport.TransportSerializationException: Failed to deserialize response of type [org.elasticsearch.action.admin.cluster.node.i  
nfo.NodeInfo]..

[o.e.t.n.Netty4Transport] [prod-es-dn-7a] exception caught on transport layer [[id: 0x8b42e333, L:/10.241.0.76:41400 - R:10  
.241.0.75/10.241.0.75:9300]], closing connection  
java.lang.IllegalStateException: Message not fully read (response) for requestId [70056], handler [org.elasticsearch.transport.TransportService$ContextRestor  
eResponseHandler/org.elasticsearch.action.support.nodes.TransportNodesAction$AsyncAction$1@5caf60c0], error [false]; resetting

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 4, 2018, 11:20am UTC](https://discuss.elastic.co/t/nodes-keep-disconnecting-from-cluster-at-random/109592/9 "2018-01-04T11:20:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
