# Cluster hanging on node failure

**URL:** <https://discuss.elastic.co/t/cluster-hanging-on-node-failure/22246>\
**Category:** Elasticsearch\
**Created:** [February 18, 2015, 7:30pm UTC](https://discuss.elastic.co/t/cluster-hanging-on-node-failure/22246 "2015-02-18T19:30:46Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Max\_Charas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/max_charas/32/889_2.png) [@Max\_Charas](https://discuss.elastic.co/u/Max_Charas)\
**Post date:** [February 18, 2015, 7:30pm UTC](https://discuss.elastic.co/t/cluster-hanging-on-node-failure/22246/1 "2015-02-18T19:30:46Z")

</div>

Hello all of you bright people,

We’re currently running a smallish 300 GB cluster in production on 5 nodes  
with around 30 mil docs. Everything works flawlessly except when a node  
really goes down (I mean like network/ HW failure/ kill -9).

When we lose a node the cluster becomes more or less completely  
unresponsive for a few minutes. Both regarding indexing and querying. This  
is of course, less than ideal as we have load 24/7.

I would really appreciate some help with understanding best practice  
settings to have a robust cluster.

First goal for us is for the cluster to not become unresponsive in the  
event of a node crash. After reading everything I could find on the web I  
can't really understand if ES is designed to be unresponsive for  
ping\_retries\*ping\_timeout seconds or if the cluster will continue to server  
query requests even during this time. Could anyone help me shed light on  
this?

Secondly in the event of a even worse failure where the cluster goes into  
red state, would it be possible to allow the cluster to still serve  
read/query requests?

I would be ever so grateful for anyone willing to help me understand how  
this works or what we would need to change to make our ES installation more  
robust.

I’ve included our config here:

cluster.name: clustername

node.name: nodename

path.data: /index

node.master: true

node.data: true

discovery.zen.minimum\_master\_nodes: 3

discovery.zen.ping.multicast.enabled: false

discovery.zen.ping.multicast.ping.enabled: false

discovery.zen.ping.unicast.enabled: true

discovery.zen.ping.unicast.hosts: ["host1","host2","host3"]

bootstrap.mlockall: true

index.number\_of\_shards: 10

action.disable\_delete\_all\_indices: true

marvel.agent.exporter.es.hosts: ["marvel:9200"]

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/bb1d307b-8c00-469d-81fb-8067942d02ad%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/bb1d307b-8c00-469d-81fb-8067942d02ad%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Max\_Charas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/max_charas/32/889_2.png) [@Max\_Charas](https://discuss.elastic.co/u/Max_Charas)\
**Post date:** [February 19, 2015, 6:21pm UTC](https://discuss.elastic.co/t/cluster-hanging-on-node-failure/22246/2 "2015-02-19T18:21:25Z")

</div>

I posted here too:

> <https://stackoverflow.com/questions/28601885/cluster-hanging-on-node-failure>

Would love to get some help with this.

Best,  
Max

Den onsdag 18 februari 2015 kl. 20:30:46 UTC+1 skrev Max Charas:

> Hello all of you bright people,
> 
> We’re currently running a smallish 300 GB cluster in production on 5 nodes  
> with around 30 mil docs. Everything works flawlessly except when a node  
> really goes down (I mean like network/ HW failure/ kill -9).
> 
> When we lose a node the cluster becomes more or less completely  
> unresponsive for a few minutes. Both regarding indexing and querying. This  
> is of course, less than ideal as we have load 24/7.
> 
> I would really appreciate some help with understanding best practice  
> settings to have a robust cluster.
> 
> First goal for us is for the cluster to not become unresponsive in the  
> event of a node crash. After reading everything I could find on the web I  
> can't really understand if ES is designed to be unresponsive for  
> ping\_retries\*ping\_timeout seconds or if the cluster will continue to server  
> query requests even during this time. Could anyone help me shed light on  
> this?
> 
> Secondly in the event of a even worse failure where the cluster goes into  
> red state, would it be possible to allow the cluster to still serve  
> read/query requests?
> 
> I would be ever so grateful for anyone willing to help me understand how  
> this works or what we would need to change to make our ES installation more  
> robust.
> 
> I’ve included our config here:
> 
> cluster.name: clustername
> 
> node.name: nodename
> 
> path.data: /index
> 
> node.master: true
> 
> node.data: true
> 
> discovery.zen.minimum\_master\_nodes: 3
> 
> discovery.zen.ping.multicast.enabled: false
> 
> discovery.zen.ping.multicast.ping.enabled: false
> 
> discovery.zen.ping.unicast.enabled: true
> 
> discovery.zen.ping.unicast.hosts: ["host1","host2","host3"]
> 
> bootstrap.mlockall: true
> 
> index.number\_of\_shards: 10
> 
> action.disable\_delete\_all\_indices: true
> 
> marvel.agent.exporter.es.hosts: ["marvel:9200"]

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/fb7171cd-a55e-4ccb-b15f-a6159931b3ff%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/fb7171cd-a55e-4ccb-b15f-a6159931b3ff%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:31am UTC](https://discuss.elastic.co/t/cluster-hanging-on-node-failure/22246/3 "2017-07-06T00:31:29Z")

</div>


