# Problem communicating within nodes in cluster - send message failed, node gets removed from cluster

**URL:** <https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736>\
**Category:** Elasticsearch\
**Created:** [November 14, 2018, 8:16pm UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736 "2018-11-14T20:16:48Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![elco\_comm1982](https://avatars.discourse-cdn.com/v4/letter/e/0ea827/32.png) [@elco\_comm1982](https://discuss.elastic.co/u/elco_comm1982)\
**Post date:** [November 14, 2018, 8:16pm UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/1 "2018-11-14T20:16:49Z")

</div>

Hi,

We keep seeing this issue intermittently in our cluster.. After the message failed error comes, the node gets removed from the cluster and the cluster state goes into red..

Immediately afterwards, the node gets added back into the cluster.  
This happens a number of times, each hour, though there is no specified frequency. Is there a way we can work around this issue?

[2018-11-14T11:59:16,974][WARN][o.e.x.s.t.n.SecurityNetty4ServerTransport] [es-node1-001] send message failed [channel: NettyTcpChannel{localAddress=/173.37.96.31:9300, remoteAddress=/173.36.39.60:50426}]  
javax.net.ssl.SSLException: SSLEngine closed already  
at io.netty.handler.ssl.SslHandler.wrap(...)(Unknown Source) ~[?:?]  
[2018-11-14T11:59:16,974][WARN][o.e.x.s.t.n.SecurityNetty4ServerTransport] [es-node1-001] send message failed [channel: NettyTcpChannel{localAddress=/173.37.96.31:9300, remoteAddress=/173.36.39.60:50426}]  
javax.net.ssl.SSLException: SSLEngine closed already  
at io.netty.handler.ssl.SslHandler.wrap(...)(Unknown Source) ~[?:?]  
[2018-11-14T12:00:08,117][WARN][o.e.x.s.t.n.SecurityNetty4ServerTransport] [es-node1-001] exception caught on transport layer [NettyTcpChannel{localAddress=/173.37.96.31:9300, remoteAddress=/173.36.39.60:50672}], closing connection  
[2018-11-14T12:07:02,344][WARN][o.e.x.s.t.n.SecurityNetty4ServerTransport] [es-node1-001] send message failed [channel: NettyTcpChannel{localAddress=0.0.0.0/0.0.0.0:9300, remoteAddress=/173.36.39.60:50730}]  
java.nio.channels.ClosedChannelException: null  
at io.netty.channel.AbstractChannel$AbstractUnsafe.write(...)(Unknown Source) ~[?:?]  
[2018-11-14T12:08:31,298][INFO][o.e.c.s.ClusterApplierService] [es-node1-001] removed {{es-node2-001}{V4A7hjvtQcyFW7BlxG-j4w}{3GXiTiZ\_QfCfFtprDhA2og}{es-node2-001}{173.36.39.60:9300}{xpack.installed=true},}, reason: apply cluster state (from master [master {es-node1-002}{44pA9ErPTb-y3zOylW8Z\_Q}{byFkKyl5S1a\_s3RiujXyRA}{es-node1-002}{173.37.96.32:9300}{xpack.installed=true} committed version [148]])  
[2018-11-14T12:08:31,809][DEBUG][o.e.a.a.c.n.s.TransportNodesStatsAction] [es-node1-001] failed to execute on node [V4A7hjvtQcyFW7BlxG-j4w]  
org.elasticsearch.transport.NodeDisconnectedException: [es-node2-001][173.36.39.60:9300][cluster:monitor/nodes/stats[n]] disconnected  
[2018-11-14T12:08:58,991][INFO][o.e.c.s.ClusterApplierService] [es-node1-001] added {{es-node2-001}{V4A7hjvtQcyFW7BlxG-j4w}{3GXiTiZ\_QfCfFtprDhA2og}{es-node2-001}{173.36.39.60:9300}{xpack.installed=true},}, reason: apply cluster state (from master [master {es-node1-002}{44pA9ErPTb-y3zOylW8Z\_Q}{byFkKyl5S1a\_s3RiujXyRA}{es-node1-002}{173.37.96.32:9300}{xpack.installed=true} committed version [152]])

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 15, 2018, 4:25am UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/2 "2018-11-15T04:25:48Z")

</div>

Are those the logs from the node that disappears?

---

<div class="post-metadata">

**Author:** ![elco\_comm1982](https://avatars.discourse-cdn.com/v4/letter/e/0ea827/32.png) [@elco\_comm1982](https://discuss.elastic.co/u/elco_comm1982)\
**Post date:** [November 15, 2018, 6:51am UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/3 "2018-11-15T06:51:19Z")

</div>

Hi,  
Thats right.. these are from the node that keeps popping out of the cluster

---

<div class="post-metadata">

**Author:** ![elco\_comm1982](https://avatars.discourse-cdn.com/v4/letter/e/0ea827/32.png) [@elco\_comm1982](https://discuss.elastic.co/u/elco_comm1982)\
**Post date:** [November 15, 2018, 7:10am UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/4 "2018-11-15T07:10:21Z")

</div>

Something like this too saw it once

[2018-11-14T23:04:48,723][WARN][o.e.t.TransportService] [es-node1-002] Received response for a request that has timed out, sent [82838ms] ago, timed out [52837ms] ago, action [internal:discovery/zen/fd/master\_ping], node [{es-node2-002}{rMYDt\_5ITuCTJbjUPluLJA}{Quk0hNopRJ-CZJyQz81AAg}{es-node2-002}{173.36.39.61:9300}{ml.machine\_memory=67556810752, ml.max\_open\_jobs=20, xpack.installed=true, ml.enabled=true}], id [536]

---

<div class="post-metadata">

**Author:** ![Tek\_Chand](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tek_chand/32/34318_2.png) [@Tek\_Chand](https://discuss.elastic.co/u/Tek_Chand)\
**Post date:** [November 15, 2018, 7:19am UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/5 "2018-11-15T07:19:30Z")

</div>

@elco_comm1982, Can you please let me know how many master node you have in your cluster? And what the value of `discovery.zen.minimum_master_nodes` have you set in your elasticsearch.yml file.

Thanks.

---

<div class="post-metadata">

**Author:** ![elco\_comm1982](https://avatars.discourse-cdn.com/v4/letter/e/0ea827/32.png) [@elco\_comm1982](https://discuss.elastic.co/u/elco_comm1982)\
**Post date:** [November 16, 2018, 7:47pm UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/6 "2018-11-16T19:47:42Z")

</div>

Hi, we have 4 master eligible nodes in the cluster. The cluster is spread across 2 data centers, with each data center having 2 nodes. The value of `discovery.zen.minimum_master_nodes` is 2. After doing a number of things, i found the overridden value of thread\_pool.write.queue\_size, thread\_pool.index.queue\_size was causing the issue.. Once i removed it, the cluster was stable.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 16, 2018, 8:10pm UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/7 "2018-11-16T20:10:35Z")

</div>

That is not good. As per [these guidelines](https://www.elastic.co/guide/en/elasticsearch/reference/6.5/modules-node.html#split-brain), you should have `minimum_master_nodes` set to `3` as you have 4 master-eligible nodes in order to avoid split-brain scenarios.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 16, 2018, 8:34pm UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/8 "2018-11-16T20:34:56Z")

</div>

> [@elco\_comm1982](#):
>
> The cluster is spread across 2 data centers

Are these datacentres spread far apart?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 18, 2018, 9:31am UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/9 "2018-11-18T09:31:36Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> you should have `minimum_master_nodes` set to `2` as you have 4 master-eligible nodes in order to avoid split-brain scenarios.

I think that's a typo, `minimum_master_nodes` must be at least 3 if there are 4 master-eligible nodes. If you leave it at 2 then it's only a matter of time before a brief connectivity loss between your two data centres leads to data loss.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 18, 2018, 10:31am UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/10 "2018-11-18T10:31:19Z")

</div>

Thanks for spotting this. Have corrected it.

---

<div class="post-metadata">

**Author:** ![elco\_comm1982](https://avatars.discourse-cdn.com/v4/letter/e/0ea827/32.png) [@elco\_comm1982](https://discuss.elastic.co/u/elco_comm1982)\
**Post date:** [November 19, 2018, 7:48pm UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/12 "2018-11-19T19:48:39Z")

</div>

Thanks.. Will do this

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 17, 2018, 7:48pm UTC](https://discuss.elastic.co/t/problem-communicating-within-nodes-in-cluster-send-message-failed-node-gets-removed-from-cluster/156736/13 "2018-12-17T19:48:50Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
