# Nodes silently leaving and returning?

**URL:** <https://discuss.elastic.co/t/nodes-silently-leaving-and-returning/69660>\
**Category:** Elasticsearch\
**Created:** [December 21, 2016, 10:03am UTC](https://discuss.elastic.co/t/nodes-silently-leaving-and-returning/69660 "2016-12-21T10:03:25Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![maf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maf/32/77965_2.png) [@maf](https://discuss.elastic.co/u/maf)\
**Post date:** [December 21, 2016, 10:03am UTC](https://discuss.elastic.co/t/nodes-silently-leaving-and-returning/69660/1 "2016-12-21T10:03:25Z")

</div>

We recently upgraded to 2.4.2 and are now seeing nodes mysteriously leaving and coming back. That is the cluster becomes yellow for a while and then goes back to green.

In the log on the master node I see this:  
[2016-12-21 07:20:41,821][WARN][cluster.action.shard] [es-150e.foo.bar] [reference\_2015-04-01\_2][5] received shard failed for target shard [[reference\_2015-04-01\_2][5], node[-wygWN2DQG6FibS7MoW25g], [R], v[183], s[STARTED], a[id=Qpn3IcmGQPWMlotMpyZPYg]], indexUUID [MNw2RvCPSTeEnNNJYoUIxw], message [failed to perform indices:data/write/bulk[s] on replica on node {es-247d.foo.bar}{-wygWN2DQG6FibS7MoW25g}{10.0.69.125}{10.0.69.125:9300}{aws\_availability\_zone=us-east-1d, index\_set=reference\_partitioned, max\_local\_storage\_nodes=1, master=false}], failure [NodeDisconnectedException[[es-247d.foo.bar][10.0.69.125:9300][indices:data/write/bulk[s][r]] disconnected]]  
NodeDisconnectedException[[es-247d.foo.bar][10.0.69.125:9300][indices:data/write/bulk[s][r]] disconnected]  
[2016-12-21 07:20:41,831][INFO][cluster.routing.allocation] [es-150e.foo.bar] Cluster health status changed from [GREEN] to [YELLOW] (reason: [shards failed [[reference\_2015-04-01\_2][5]] ...]).  
[2016-12-21 07:54:50,472][INFO][cluster.routing.allocation] [es-150e.foo.bar] Cluster health status changed from [YELLOW] to [GREEN] (reason: [shards started [[reference\_2015-04-01\_2][5]] ...]).

The strange thing is that I do not see any relevant log entries on the client (in this case e-247d). During the time the cluster was yellow that client only logged a few lines like this:  
[2016-12-21 07:28:20,360][WARN][index.fielddata] [es-247d.foo.bar] [reference\_2015-09-22\_1] failed to find format [compressed] for field [attributes.document\_position], will use default  
This is a separate problem which we will fix.

But can somebody explain to me why the luster went yellow for a while? There were no interruptions in network traffic or spikes in CPU-usage. We have also seen the same behavior at other times then involving other nodes.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 21, 2016, 11:26pm UTC](https://discuss.elastic.co/t/nodes-silently-leaving-and-returning/69660/2 "2016-12-21T23:26:18Z")

</div>

Check the logs on the node that left, there should be something before/during/after the time of drop out.

---

<div class="post-metadata">

**Author:** ![maf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maf/32/77965_2.png) [@maf](https://discuss.elastic.co/u/maf)\
**Post date:** [December 22, 2016, 7:09am UTC](https://discuss.elastic.co/t/nodes-silently-leaving-and-returning/69660/3 "2016-12-22T07:09:15Z")

</div>

As I wrote in my original post there are no relevant log messages on the node that the master considered to be disconnected.

---

<div class="post-metadata">

**Author:** ![maf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maf/32/77965_2.png) [@maf](https://discuss.elastic.co/u/maf)\
**Post date:** [January 16, 2017, 8:39am UTC](https://discuss.elastic.co/t/nodes-silently-leaving-and-returning/69660/4 "2017-01-16T08:39:13Z")

</div>

This turned out to be caused by a network hardware issue on one of the nodes. About 5% of all TCP connections to/from that node failed. The tricky thing that al lt of other nodes also unexpectedly left the es-cluster. But all problems stopped when we decommissioned that one node.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 13, 2017, 8:39am UTC](https://discuss.elastic.co/t/nodes-silently-leaving-and-returning/69660/5 "2017-02-13T08:39:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
