# Elasticsearch cluster timeout when node dies

**URL:** <https://discuss.elastic.co/t/elasticsearch-cluster-timeout-when-node-dies/14892>\
**Category:** Elasticsearch\
**Created:** [December 16, 2013, 4:39pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-timeout-when-node-dies/14892 "2013-12-16T16:39:38Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jjheinon](https://avatars.discourse-cdn.com/v4/letter/j/f04885/32.png) [@jjheinon](https://discuss.elastic.co/u/jjheinon)\
**Post date:** [December 16, 2013, 4:39pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-timeout-when-node-dies/14892/1 "2013-12-16T16:39:38Z")

</div>

I have an Elasticsearch cluster with three servers (testnode00, testnode01,  
testnode02), with two Elasticsearch instances running on each server (ports  
9300 and 9301). Total 6 instances.  
The instances have been configured with  
cluster.routing.allocation.awareness.attributes=zone,tag setting so that  
instances running on the same server can both die and the cluster still  
works properly.

Config file in [https://gist.github.com/jjheinon/7989423](https://gist.github.com/jjheinon/7989423)

This works in real life too, I can shut down both instances on the same  
server and everything still works.

Everything works fine, until I actually shut down one of the servers (i.e.  
testnode01)

Then the whole cluster will become unresponsive.

The basic status requests do work:

curl '[http://testnode00:9200/](http://testnode00:9200/)'  
-\>  
{  
"ok" : true,  
"status" : 200,  
"name" : "testnode00\_ebs",  
"version" : {  
"number" : "0.90.5",  
"build\_hash" : "c8714e8e0620b62638f660f6144831792b9dedee",  
"build\_timestamp" : "2013-09-17T13:09:46Z",  
"build\_snapshot" : false,  
"lucene\_version" : "4.4"  
},  
"tagline" : "You Know, for Search"  
}

Cluster health request also works:

curl '[http://testnode00:9200/\_cluster/health](http://testnode00:9200/_cluster/health)'  
-\>  
{  
"active\_primary\_shards":120,"active\_shards":240,"cluster\_name":  
"test\_cluster",  
"initializing\_shards":0,"number\_of\_data\_nodes":6,"number\_of\_nodes":6,"  
relocating\_shards":2,"status":  
"green",  
"timed\_out":false,"unassigned\_shards":0}

but node status request times out:  
curl '[http://testnode00:9200/\_nodes/stats](http://testnode00:9200/_nodes/stats)'

-\> Timeout

Search requests won't work either anymore:

curl '[http://testnode00:9200/\_search/?q=name:test](http://testnode00:9200/_search/?q=name:test)'

-\> Timeout

There's nothing visible on elasticsearch log if shutting down the server.  
Iif I manually shut down both Elasticsearch instances on the server, then I  
will get the node disconnect messages on the log and everything fails over  
properly and all the above requests work.

[2013-12-16 15:57:59,945][DEBUG][action.admin.cluster.node.stats]  
[testnode00\_ebs] failed to execute on node [Th4-MYtTTdGh3wZFh3W4vA]

org.elasticsearch.transport.NodeDisconnectedException:  
[testnode01\_ebs][inet[/10.43.129.161:9300]][cluster/nodes/stats/n]  
disconnected

Any ideas why the unicast discovery won't detect missing servers?  
discovery.zen.ping.timeout does not seem to help. And why \_nodes/stats  
request doesn't work if one of the nodes is unresponsive?  
Is there a way to tune TTL values for requests between Elasticsearch nodes?

Additional question:  
Is there a way to tell cloud-aws ec2 discovery plugin to find two instances  
on a single server or does it detect only the first one (on port 9300 and  
not the one on 9301)?

Regards,  
// Janne

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/72c88b17-8c26-4d0c-b8e6-3ef034614c96%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/72c88b17-8c26-4d0c-b8e6-3ef034614c96%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jjheinon](https://avatars.discourse-cdn.com/v4/letter/j/f04885/32.png) [@jjheinon](https://discuss.elastic.co/u/jjheinon)\
**Post date:** [December 16, 2013, 9:45pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-timeout-when-node-dies/14892/2 "2013-12-16T21:45:33Z")

</div>

Actually figured out a solution.

For some reason, the node was running out of threads. There was plenty of  
resources available, but Elasticsearch default threadpool settings were too  
low.

Setting threadpool sizes manually fixed the issue, now the server failure  
detection and timeouts work again:

> <https://gist.github.com/jjheinon/7994672>

Not quite sure yet what exact setting caused the problem.

// Janne

On Monday, December 16, 2013 6:39:38 PM UTC+2, jjheinon wrote:

> I have an Elasticsearch cluster with three servers (testnode00,  
> testnode01, testnode02), with two Elasticsearch instances running on each  
> server (ports 9300 and 9301). Total 6 instances.  
> The instances have been configured with  
> cluster.routing.allocation.awareness.attributes=zone,tag setting so that  
> instances running on the same server can both die and the cluster still  
> works properly.
> 
> Config file in [Elasticsearch config · GitHub](https://gist.github.com/jjheinon/7989423)
> 
> This works in real life too, I can shut down both instances on the same  
> server and everything still works.
> 
> Everything works fine, until I actually shut down one of the servers (i.e.  
> testnode01)
> 
> Then the whole cluster will become unresponsive.
> 
> The basic status requests do work:
> 
> curl '[http://testnode00:9200/](http://testnode00:9200/)'  
> -\>  
> {  
> "ok" : true,  
> "status" : 200,  
> "name" : "testnode00\_ebs",  
> "version" : {  
> "number" : "0.90.5",  
> "build\_hash" : "c8714e8e0620b62638f660f6144831792b9dedee",  
> "build\_timestamp" : "2013-09-17T13:09:46Z",  
> "build\_snapshot" : false,  
> "lucene\_version" : "4.4"  
> },  
> "tagline" : "You Know, for Search"  
> }
> 
> Cluster health request also works:
> 
> curl '[http://testnode00:9200/\_cluster/health](http://testnode00:9200/_cluster/health)'  
> -\>  
> {  
> "active\_primary\_shards":120,"active\_shards":240,"cluster\_name":  
> "test\_cluster",  
> "initializing\_shards":0,"number\_of\_data\_nodes":6,"number\_of\_nodes":6,"  
> relocating\_shards":2,"status":  
> "green",  
> "timed\_out":false,"unassigned\_shards":0}
> 
> but node status request times out:  
> curl '[http://testnode00:9200/\_nodes/stats](http://testnode00:9200/_nodes/stats)'
> 
> -\> Timeout
> 
> Search requests won't work either anymore:
> 
> curl '[http://testnode00:9200/\_search/?q=name:test](http://testnode00:9200/_search/?q=name:test)'
> 
> -\> Timeout
> 
> There's nothing visible on elasticsearch log if shutting down the server.  
> Iif I manually shut down both Elasticsearch instances on the server, then  
> I will get the node disconnect messages on the log and everything fails  
> over properly and all the above requests work.
> 
> [2013-12-16 15:57:59,945][DEBUG][action.admin.cluster.node.stats]  
> [testnode00\_ebs] failed to execute on node [Th4-MYtTTdGh3wZFh3W4vA]
> 
> org.elasticsearch.transport.NodeDisconnectedException:  
> [testnode01\_ebs][inet[/10.43.129.161:9300]][cluster/nodes/stats/n]  
> disconnected
> 
> Any ideas why the unicast discovery won't detect missing servers?  
> discovery.zen.ping.timeout does not seem to help. And why \_nodes/stats  
> request doesn't work if one of the nodes is unresponsive?  
> Is there a way to tune TTL values for requests between Elasticsearch nodes?
> 
> Additional question:  
> Is there a way to tell cloud-aws ec2 discovery plugin to find two  
> instances on a single server or does it detect only the first one (on port  
> 9300 and not the one on 9301)?
> 
> Regards,  
> // Janne

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/fc82ed5a-7c72-44de-ac68-99c225fa315e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/fc82ed5a-7c72-44de-ac68-99c225fa315e%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:00am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-timeout-when-node-dies/14892/3 "2017-07-06T02:00:50Z")

</div>


