# Single unresponsive node stalls overall cluster performance

**URL:** <https://discuss.elastic.co/t/single-unresponsive-node-stalls-overall-cluster-performance/113978>\
**Category:** Elasticsearch\
**Created:** [January 3, 2018, 8:24pm UTC](https://discuss.elastic.co/t/single-unresponsive-node-stalls-overall-cluster-performance/113978 "2018-01-03T20:24:52Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![drs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/drs/32/8001_2.png) [@drs](https://discuss.elastic.co/u/drs)\
**Post date:** [January 3, 2018, 8:24pm UTC](https://discuss.elastic.co/t/single-unresponsive-node-stalls-overall-cluster-performance/113978/1 "2018-01-03T20:24:52Z")

</div>

At seemingly random times, one of my data nodes becomes unresponsive. The box's disk i/o reads get pegged at about 200 MB/s until the Elasticsearch service is restarted.

 ![es5-data05 diskio](https://us1.discourse-cdn.com/elastic/original/3X/8/e/8ef344be87c2f0f43bfa5fd5ff8ac3826164a3be.png)

(the service was restarted at about 11:30 am).

During this time, searches and indexes to the cluster time out, and the search queue gets backed up on the affected node:

```auto
GET es5-client01:9200/_cat/thread_pool/search | sort
es5-client01 search 0 0 0
es5-client02 search 0 0 0
es5-client03 search 0 0 0
es5-data01 search 0 0 0
es5-data02 search 0 0 0
es5-data03 search 0 0 0
es5-data04 search 0 0 0
es5-data05 search 13 986 6057
es5-data06 search 0 0 0
es5-data07 search 0 0 0
es5-data08 search 0 0 0
es5-data09 search 0 0 0
es5-data10 search 0 0 0
es5-data11 search 0 0 0
es5-data12 search 0 0 0
es5-master01 search 0 0 0
es5-master02 search 0 0 0
es5-master03 search 0 0 0

```

Restarting the service on the affected data node immediately improves performance.

I've captured the hot-threads on the affected node: [https://pastebin.com/np698JcX](https://pastebin.com/np698JcX)

I'm running Elasticsearch v5.5.2 with the search-guard plugin.

I've run into this issue several times over the last month but am at a loss as to what the cause could be. Any advice?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 4, 2018, 10:04am UTC](https://discuss.elastic.co/t/single-unresponsive-node-stalls-overall-cluster-performance/113978/2 "2018-01-04T10:04:57Z")

</div>

Is there anything in the logs around the time the node gets pegged? Have you been able to track down what data is being written?

---

<div class="post-metadata">

**Author:** ![drs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/drs/32/8001_2.png) [@drs](https://discuss.elastic.co/u/drs)\
**Post date:** [January 4, 2018, 1:19pm UTC](https://discuss.elastic.co/t/single-unresponsive-node-stalls-overall-cluster-performance/113978/3 "2018-01-04T13:19:40Z")

</div>

There's a `removed` and `added` in the logs that correlate with the beginning of the high disk i/o, but that might be a coincidence (beginning about 15:15 in the timestamps below):

```auto
[2018-01-03T11:36:09,925][INFO][o.e.m.j.JvmGcMonitorService] [es5-data05] [gc][185906] overhead, spent [430ms] collecting in the last [1.3s]
[2018-01-03T15:14:29,865][INFO][o.e.c.s.ClusterService] [es5-data05] removed {{es5-client01}{JvzC36f5QWCv5wpykOl1gg}{nIZKAKNXSr6IA9k1TZ5Y7g}{10.208.0.137}{10.208.0.137:9300}{ml.max_open_jobs=10, ml.enabled=true},}, reason: zen-disco-receive(from master [master {es5-master01}{n-WLDE5PSC6Z727V_Jx4CQ}{RN6za0HDTjSySu4BcyLaXw}{10.100.4.27}{10.100.4.27:9300}{ml.max_open_jobs=10, ml.enabled=true} committed version [11621]])
[2018-01-03T15:15:04,882][INFO][o.e.c.s.ClusterService] [es5-data05] added {{es5-client01}{JvzC36f5QWCv5wpykOl1gg}{jTMJAX2GTE609ORBkLjMXw}{10.208.0.137}{10.208.0.137:9300}{ml.max_open_jobs=10, ml.enabled=true},}, reason: zen-disco-receive(from master [master {es5-master01}{n-WLDE5PSC6Z727V_Jx4CQ}{RN6za0HDTjSySu4BcyLaXw}{10.100.4.27}{10.100.4.27:9300}{ml.max_open_jobs=10, ml.enabled=true} committed version [11622]])
[2018-01-03T16:30:54,196][INFO][o.e.x.m.j.p.NativeController] Native controller process has stopped - no new native processes can be started
[2018-01-03T16:30:54,341][INFO][o.e.n.Node] [es5-data05] stopping ...

```

Note, there's nothing being written to disk, it's high disk _reads_. When it happens again, I can try to track down what's being read.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 1, 2018, 1:19pm UTC](https://discuss.elastic.co/t/single-unresponsive-node-stalls-overall-cluster-performance/113978/4 "2018-02-01T13:19:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
