# Client nodes stop responding

**URL:** <https://discuss.elastic.co/t/client-nodes-stop-responding/20501>\
**Category:** Elasticsearch\
**Created:** [October 30, 2014, 2:29pm UTC](https://discuss.elastic.co/t/client-nodes-stop-responding/20501 "2014-10-30T14:29:52Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jeff\_Keller](https://avatars.discourse-cdn.com/v4/letter/j/d9b06d/32.png) [@Jeff\_Keller](https://discuss.elastic.co/u/Jeff_Keller)\
**Post date:** [October 30, 2014, 2:29pm UTC](https://discuss.elastic.co/t/client-nodes-stop-responding/20501/1 "2014-10-30T14:29:52Z")

</div>

We are running into an issue where our client nodes will stop responding to  
requests that require checking with other nodes. The following is our setup:

1 dedicated master node

2 dedicated data nodes

3 client nodes (master: false; data: false)

Everything is fine and happy for a while, and then after 30-45 minutes, one  
of the client nodes will stop sending responses to queries that require  
talking with other nodes. We are using the HTTP REST API. When things go  
badly, the following will hang:

curl -XGET ‘[http://localhost:9200/\_search?size=1’](http://localhost:9200/_search?size=1%E2%80%99)

curl -XGET ‘[http://localhost:9200/\_cat/thread\_pool?v’](http://localhost:9200/_cat/thread_pool?v%E2%80%99)

But the following will succeed (as it can just use metadata on the node  
itself):

curl -XGET ‘[http://localhost:9200/\_cluster/health?pretty=1’](http://localhost:9200/_cluster/health?pretty=1%E2%80%99)

The problem node doesn’t seem to have any CPU or IO load. We don’t seem to  
be running into heap issues. netstat doesn’t report any connections in  
TIME\_WAIT on any of the nodes. If we run queries from the problem client  
node at the command prompt directly at the data node, everything works. So,  
if we instead run:

curl -XGET ‘[http://data.node.ip:9200/\_search?size=1](http://data.node.ip:9200/_search?size=1)  
[http://localhost:9200/\_search?size=1](http://localhost:9200/_search?size=1)’

It works as expected. This tells me there isn’t a socket exhaustion issue  
since we can make new connections from the problem node to other nodes.

We turned logged all the way up (“ALL”) on one of the client nodes until it  
started failing, but there was nothing in there of interest. The last few  
minutes just had messages about the idle connection reaper running every  
minute.

We tried increasing the various connections\_per\_node values to:

transport.connections\_per\_node.bulk =\> 6

transport.connections\_per\_node.reg =\> 12

transport.connections\_per\_node.state =\> 2

transport.connections\_per\_node.ping =\> 2

This made no noticeable difference.

When one of the client nodes has started having problems, the cluster still  
sees the node as part of the cluster. When we kill the ES process on that  
node, all the other nodes then notice it went away as expected. When we  
restart ES on the problem node, it comes back up and everything works great  
for another 30-45 minutes.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b8febb38-cd43-4102-b4fb-6dcdd9749aa5%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b8febb38-cd43-4102-b4fb-6dcdd9749aa5%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:52am UTC](https://discuss.elastic.co/t/client-nodes-stop-responding/20501/2 "2017-07-06T00:52:58Z")

</div>


