# Cluster stalls when nodes are removed (or the true meaning of expected\_nodes)

**URL:** <https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307>\
**Category:** Elasticsearch\
**Created:** [January 11, 2013, 1:35am UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307 "2013-01-11T01:35:17Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 11, 2013, 1:35am UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/1 "2013-01-11T01:35:17Z")

</div>

One of my clusters has 8 nodes running 0.20.0.RC1. Most settings are the  
default except for:

bootstrap.mlockall: true  
transport.tcp.connect\_timeout: 5s  
gateway.expected\_nodes: 8 \<-- we'll get to this in a second  
discovery.zen.minimum\_master\_nodes: 5  
discovery.zen.ping.multicast.enabled: true

As you can see, the number of expected nodes is equals to the total number  
of nodes in the cluster. If one of the nodes disappears from the cluster,  
the clusters stalls completely for about 1-2 minutes.

As a test, gateway.expected\_nodes was reduced to 6. After changing the  
setting, the cluster no longer stalls if a node disappears.

The general consensus is to set gateway.expected\_nodes to the number of  
nodes in the cluster. Is gateway recovery what is affecting the cluster? If  
the cluster does not respond to request during recovery, shouldn't  
the gateway.expected\_nodes value be set to something lower in case a node  
goes down?

Ivan

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 15, 2013, 12:23am UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/2 "2013-01-15T00:23:21Z")

</div>

This setting should have no effect after initial recovery. Are you sure  
that the effect that you observed wasn't a coincidence?

On Thursday, January 10, 2013 8:35:17 PM UTC-5, Ivan Brusic wrote:

> One of my clusters has 8 nodes running 0.20.0.RC1. Most settings are the  
> default except for:
> 
> bootstrap.mlockall: true  
> transport.tcp.connect\_timeout: 5s  
> gateway.expected\_nodes: 8 \<-- we'll get to this in a second  
> discovery.zen.minimum\_master\_nodes: 5  
> discovery.zen.ping.multicast.enabled: true
> 
> As you can see, the number of expected nodes is equals to the total number  
> of nodes in the cluster. If one of the nodes disappears from the cluster,  
> the clusters stalls completely for about 1-2 minutes.
> 
> As a test, gateway.expected\_nodes was reduced to 6. After changing the  
> setting, the cluster no longer stalls if a node disappears.
> 
> The general consensus is to set gateway.expected\_nodes to the number of  
> nodes in the cluster. Is gateway recovery what is affecting the cluster? If  
> the cluster does not respond to request during recovery, shouldn't  
> the gateway.expected\_nodes value be set to something lower in case a node  
> goes down?
> 
> Ivan

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 15, 2013, 1:03am UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/3 "2013-01-15T01:03:45Z")

</div>

Expected nodes was a red herring. The true issue might be ping timeout for  
zen discovery. If a node is no longer ping-able, the cluster stalls. Doing  
some tests, will write more later with more facts.

If a node disappears completely, should the cluster stall?

--  
Ivan

On Mon, Jan 14, 2013 at 4:23 PM, Igor Motov [imotov@gmail.com](mailto:imotov@gmail.com) wrote:

> ld have no effect after initial recovery. Are you sure tha

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 15, 2013, 2:20am UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/4 "2013-01-15T02:20:11Z")

</div>

Yeah, more information would be useful. It might also help to set logging  
level for "discovery" to TRACE to see what's actually going on with pings  
and connections between nodes. I would suspect that when a node disappears,  
elasticsearch might not detect it quickly enough and during this time some  
of the requests are getting directed to the disappeared node. How do you  
simulate node disappearance by the way?

On Monday, January 14, 2013 8:03:45 PM UTC-5, Ivan Brusic wrote:

> Expected nodes was a red herring. The true issue might be ping timeout for  
> zen discovery. If a node is no longer ping-able, the cluster stalls. Doing  
> some tests, will write more later with more facts.
> 
> If a node disappears completely, should the cluster stall?
> 
> --  
> Ivan
> 
> On Mon, Jan 14, 2013 at 4:23 PM, Igor Motov \<[imo...@gmail.com](mailto:imo...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > ld have no effect after initial recovery. Are you sure tha

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 15, 2013, 3:25am UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/5 "2013-01-15T03:25:23Z")

</div>

Back at home, so I don't have much info. Already set the levels to TRACE  
(uncommented the line in logging.yml). What is the difference between  
discovery.zen.ping.timeout and the ping\_timeout setting referenced here:  
[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86)

Once again, I don't have my notes right now, but after a cluster stalls,  
the master logs (after the cluster comes back) something about retrying [3]  
times for [30s]. Logging was set to INFO for that test, so I don't have  
finer details (right now). The cluster was unresponsive during that time.

The nodes are on VMs, so the nodes disappeared when we took down the entire  
VM host (2 ES nodes per host).

Discovery is done via multicast. I am assuming that the pings are multicast  
pings? Correct? Not a networking guru, but are these pings different from  
"normal" pings? If so, is there a command line utility that does multicast  
ping?

Cheers,

Ivan

On Mon, Jan 14, 2013 at 6:20 PM, Igor Motov [imotov@gmail.com](mailto:imotov@gmail.com) wrote:

> Yeah, more information would be useful. It might also help to set logging  
> level for "discovery" to TRACE to see what's actually going on with pings  
> and connections between nodes. I would suspect that when a node disappears,  
> elasticsearch might not detect it quickly enough and during this time some  
> of the requests are getting directed to the disappeared node. How do you  
> simulate node disappearance by the way?
> 
> On Monday, January 14, 2013 8:03:45 PM UTC-5, Ivan Brusic wrote:
> 
> > Expected nodes was a red herring. The true issue might be ping timeout  
> > for zen discovery. If a node is no longer ping-able, the cluster stalls.  
> > Doing some tests, will write more later with more facts.
> > 
> > If a node disappears completely, should the cluster stall?
> > 
> > --  
> > Ivan
> > 
> > On Mon, Jan 14, 2013 at 4:23 PM, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> > 
> > > ld have no effect after initial recovery. Are you sure tha
> > 
> > --

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 15, 2013, 2:35pm UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/6 "2013-01-15T14:35:04Z")

</div>

Discovery is done via multicast, but when nodes join the cluster they  
establish connections that are used for all other communication including  
pings.

On Monday, January 14, 2013 10:25:23 PM UTC-5, Ivan Brusic wrote:

> Back at home, so I don't have much info. Already set the levels to TRACE  
> (uncommented the line in logging.yml). What is the difference between  
> discovery.zen.ping.timeout and the ping\_timeout setting referenced here:  
> [https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86)
> 
> Once again, I don't have my notes right now, but after a cluster stalls,  
> the master logs (after the cluster comes back) something about retrying [3]  
> times for [30s]. Logging was set to INFO for that test, so I don't have  
> finer details (right now). The cluster was unresponsive during that time.
> 
> The nodes are on VMs, so the nodes disappeared when we took down the  
> entire VM host (2 ES nodes per host).
> 
> Discovery is done via multicast. I am assuming that the pings are  
> multicast pings? Correct? Not a networking guru, but are these pings  
> different from "normal" pings? If so, is there a command line utility that  
> does multicast ping?
> 
> Cheers,
> 
> Ivan
> 
> On Mon, Jan 14, 2013 at 6:20 PM, Igor Motov \<[imo...@gmail.com](mailto:imo...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > Yeah, more information would be useful. It might also help to set logging  
> > level for "discovery" to TRACE to see what's actually going on with pings  
> > and connections between nodes. I would suspect that when a node disappears,  
> > elasticsearch might not detect it quickly enough and during this time some  
> > of the requests are getting directed to the disappeared node. How do you  
> > simulate node disappearance by the way?
> > 
> > On Monday, January 14, 2013 8:03:45 PM UTC-5, Ivan Brusic wrote:
> > 
> > > Expected nodes was a red herring. The true issue might be ping timeout  
> > > for zen discovery. If a node is no longer ping-able, the cluster stalls.  
> > > Doing some tests, will write more later with more facts.
> > > 
> > > If a node disappears completely, should the cluster stall?
> > > 
> > > --  
> > > Ivan
> > > 
> > > On Mon, Jan 14, 2013 at 4:23 PM, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> > > 
> > > > ld have no effect after initial recovery. Are you sure tha
> > > 
> > > --

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 15, 2013, 7:38pm UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/7 "2013-01-15T19:38:42Z")

</div>

Updated the config with the following timeouts

transport.tcp.connect\_timeout: 5s  
discovery.zen.ping.timeout: 1s \<- ignored due to precedence  
discovery.zen.ping\_timeout: 2s  
discovery.zen.fd.ping\_timeout: 2s

A VM host, containing nodes search8 and search11 was taken offline. Node  
search6 was the master. All nodes took over a minute between node removed  
messages.

[2013-01-15 11:09:34,696][INFO][cluster.service] [search12]  
removed {[search8][P7UNCh9oTE623RI8w\_zsPw][inet[/:9300]],}, reason:  
zen-disco-receive(from master  
[[search6][7d\_0aK3XTiiWh\_OGLSffig][inet[/:9300]]])  
[2013-01-15 11:10:52,616][INFO][cluster.service] [search12]  
removed {[[srch-lv111.corp.shop.com](http://srch-lv111.corp.shop.com)][OMMr4k2DRgSsvMRU8vE-eQ][inet[/:9300]],},  
reason: zen-disco-receive(from master  
[[search6][7d\_0aK3XTiiWh\_OGLSffig][inet[/:9300]]])

The log for the master is here: [Cluster stalls upon node removal · GitHub](https://gist.github.com/60bdfafd6273c05a3417)

The cluster was in a red state and unresponsive during this time. These  
outages are us testing the failover capabilities of both the VMs and the  
cluster. Having the cluster go offline completely is not a good situation  
to be in, but elevated search times would be acceptable.

Let me know what else I can provide to help fine-tune the issue.

Ivan

On Tue, Jan 15, 2013 at 6:35 AM, Igor Motov [imotov@gmail.com](mailto:imotov@gmail.com) wrote:

> Discovery is done via multicast, but when nodes join the cluster they  
> establish connections that are used for all other communication including  
> pings.
> 
> On Monday, January 14, 2013 10:25:23 PM UTC-5, Ivan Brusic wrote:
> 
> > Back at home, so I don't have much info. Already set the levels to TRACE  
> > (uncommented the line in logging.yml). What is the difference between  
> > discovery.zen.ping.timeout and the ping\_timeout setting referenced here:  
> > [https://github.com/\*\*elasticsearch/elasticsearch/](https://github.com/**elasticsearch/elasticsearch/)\*\*  
> > blob/master/src/main/java/org/ **elasticsearch/discovery/zen/**  
> > fd/NodesFaultDetection.java#\*\*L86[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86)
> > 
> > Once again, I don't have my notes right now, but after a cluster stalls,  
> > the master logs (after the cluster comes back) something about retrying [3]  
> > times for [30s]. Logging was set to INFO for that test, so I don't have  
> > finer details (right now). The cluster was unresponsive during that time.
> > 
> > The nodes are on VMs, so the nodes disappeared when we took down the  
> > entire VM host (2 ES nodes per host).
> > 
> > Discovery is done via multicast. I am assuming that the pings are  
> > multicast pings? Correct? Not a networking guru, but are these pings  
> > different from "normal" pings? If so, is there a command line utility that  
> > does multicast ping?
> > 
> > Cheers,
> > 
> > Ivan
> > 
> > On Mon, Jan 14, 2013 at 6:20 PM, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> > 
> > > Yeah, more information would be useful. It might also help to set  
> > > logging level for "discovery" to TRACE to see what's actually going on with  
> > > pings and connections between nodes. I would suspect that when a node  
> > > disappears, elasticsearch might not detect it quickly enough and during  
> > > this time some of the requests are getting directed to the disappeared  
> > > node. How do you simulate node disappearance by the way?
> > > 
> > > On Monday, January 14, 2013 8:03:45 PM UTC-5, Ivan Brusic wrote:
> > > 
> > > > Expected nodes was a red herring. The true issue might be ping timeout  
> > > > for zen discovery. If a node is no longer ping-able, the cluster stalls.  
> > > > Doing some tests, will write more later with more facts.
> > > > 
> > > > If a node disappears completely, should the cluster stall?
> > > > 
> > > > --  
> > > > Ivan
> > > > 
> > > > On Mon, Jan 14, 2013 at 4:23 PM, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> > > > 
> > > > > ld have no effect after initial recovery. Are you sure tha
> > > > 
> > > > --
> > 
> > --

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 15, 2013, 11:03pm UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/8 "2013-01-15T23:03:05Z")

</div>

Are you running any plugins that listen to cluster state changes on your  
nodes?

On Tuesday, January 15, 2013 2:38:42 PM UTC-5, Ivan Brusic wrote:

> Updated the config with the following timeouts
> 
> transport.tcp.connect\_timeout: 5s  
> discovery.zen.ping.timeout: 1s \<- ignored due to precedence  
> discovery.zen.ping\_timeout: 2s  
> discovery.zen.fd.ping\_timeout: 2s
> 
> A VM host, containing nodes search8 and search11 was taken offline. Node  
> search6 was the master. All nodes took over a minute between node removed  
> messages.
> 
> [2013-01-15 11:09:34,696][INFO][cluster.service] [search12]  
> removed {[search8][P7UNCh9oTE623RI8w\_zsPw][inet[/:9300]],}, reason:  
> zen-disco-receive(from master  
> [[search6][7d\_0aK3XTiiWh\_OGLSffig][inet[/:9300]]])  
> [2013-01-15 11:10:52,616][INFO][cluster.service] [search12]  
> removed {[[srch-lv111.corp.shop.com](http://srch-lv111.corp.shop.com)][OMMr4k2DRgSsvMRU8vE-eQ][inet[/:9300]],},  
> reason: zen-disco-receive(from master  
> [[search6][7d\_0aK3XTiiWh\_OGLSffig][inet[/:9300]]])
> 
> The log for the master is here:  
> [Cluster stalls upon node removal · GitHub](https://gist.github.com/60bdfafd6273c05a3417)
> 
> The cluster was in a red state and unresponsive during this time. These  
> outages are us testing the failover capabilities of both the VMs and the  
> cluster. Having the cluster go offline completely is not a good situation  
> to be in, but elevated search times would be acceptable.
> 
> Let me know what else I can provide to help fine-tune the issue.
> 
> Ivan
> 
> On Tue, Jan 15, 2013 at 6:35 AM, Igor Motov \<[imo...@gmail.com](mailto:imo...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > Discovery is done via multicast, but when nodes join the cluster they  
> > establish connections that are used for all other communication including  
> > pings.
> > 
> > On Monday, January 14, 2013 10:25:23 PM UTC-5, Ivan Brusic wrote:
> > 
> > > Back at home, so I don't have much info. Already set the levels to  
> > > TRACE (uncommented the line in logging.yml). What is the difference between  
> > > discovery.zen.ping.timeout and the ping\_timeout setting referenced here:  
> > > [https://github.com/\*\*elasticsearch/elasticsearch/](https://github.com/**elasticsearch/elasticsearch/)\*\*  
> > > blob/master/src/main/java/org/ **elasticsearch/discovery/zen/**  
> > > fd/NodesFaultDetection.java#\*\*L86[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/discovery/zen/fd/NodesFaultDetection.java#L86)
> > > 
> > > Once again, I don't have my notes right now, but after a cluster stalls,  
> > > the master logs (after the cluster comes back) something about retrying [3]  
> > > times for [30s]. Logging was set to INFO for that test, so I don't have  
> > > finer details (right now). The cluster was unresponsive during that time.
> > > 
> > > The nodes are on VMs, so the nodes disappeared when we took down the  
> > > entire VM host (2 ES nodes per host).
> > > 
> > > Discovery is done via multicast. I am assuming that the pings are  
> > > multicast pings? Correct? Not a networking guru, but are these pings  
> > > different from "normal" pings? If so, is there a command line utility that  
> > > does multicast ping?
> > > 
> > > Cheers,
> > > 
> > > Ivan
> > > 
> > > On Mon, Jan 14, 2013 at 6:20 PM, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> > > 
> > > > Yeah, more information would be useful. It might also help to set  
> > > > logging level for "discovery" to TRACE to see what's actually going on with  
> > > > pings and connections between nodes. I would suspect that when a node  
> > > > disappears, elasticsearch might not detect it quickly enough and during  
> > > > this time some of the requests are getting directed to the disappeared  
> > > > node. How do you simulate node disappearance by the way?
> > > > 
> > > > On Monday, January 14, 2013 8:03:45 PM UTC-5, Ivan Brusic wrote:
> > > > 
> > > > > Expected nodes was a red herring. The true issue might be ping timeout  
> > > > > for zen discovery. If a node is no longer ping-able, the cluster stalls.  
> > > > > Doing some tests, will write more later with more facts.
> > > > > 
> > > > > If a node disappears completely, should the cluster stall?
> > > > > 
> > > > > --  
> > > > > Ivan
> > > > > 
> > > > > On Mon, Jan 14, 2013 at 4:23 PM, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> > > > > 
> > > > > > ld have no effect after initial recovery. Are you sure tha
> > > > > 
> > > > > --
> > > 
> > > --

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 15, 2013, 11:13pm UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/9 "2013-01-15T23:13:59Z")

</div>

None.

[INFO][plugins] [search8] loaded , sites [bigdesk,  
head]

The timeouts proved to be too low and the cluster has been removing nodes  
too quickly (duh!)). Uping the timeouts for now.

--  
Ivan

On Tue, Jan 15, 2013 at 3:03 PM, Igor Motov [imotov@gmail.com](mailto:imotov@gmail.com) wrote:

> Are you running any plugins that listen to cluster state changes on your  
> nodes?

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 16, 2013, 6:20pm UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/10 "2013-01-16T18:20:34Z")

</div>

Looking into the various timeouts. The first 30 second pause occurs here:

[2013-01-15 11:09:35,108][DEBUG][discovery.zen.fd] [search6] [node  
] failed to ping [[search11][OMMr4k2DRgSsvMRU8vE-eQ][inet[/:9300]]],  
tried  
[3] times, each with maximum [2s] timeout

[2013-01-15 11:10:04,693][DEBUG][indices.store] [search6]  
failed to execute on node [OMMr4k2DRgSsvMRU8vE-eQ]  
org.elasticsearch.transport.ReceiveTimeoutTransportException:  
[search11][inet[/:9300]][/cluster/nodes/indices/shard/store/n]  
request\_id [289991] timed out after [30000ms]  
at  
org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:342)  
at  
java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)  
at  
java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)  
at java.lang.Thread.run(Thread.java:722)

So TransportService blocked for 30 seconds. Can't find where how this  
timeout is set. The closet I can find is the ping\_timeout set  
in TransportClientNodesService. I am assuming it is  
client.transport.ping\_timeout. Transport is threaded, so I am unsure why  
the cluster stalls during fault detection.

--  
Ivan

On Tue, Jan 15, 2013 at 3:13 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> None.
> 
> [INFO][plugins] [search8] loaded , sites [bigdesk,  
> head]
> 
> The timeouts proved to be too low and the cluster has been removing nodes  
> too quickly (duh!)). Uping the timeouts for now.
> 
> --  
> Ivan
> 
> On Tue, Jan 15, 2013 at 3:03 PM, Igor Motov [imotov@gmail.com](mailto:imotov@gmail.com) wrote:
> 
> > Are you running any plugins that listen to cluster state changes on your  
> > nodes?

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:56am UTC](https://discuss.elastic.co/t/cluster-stalls-when-nodes-are-removed-or-the-true-meaning-of-expected-nodes/10307/11 "2017-07-06T02:56:06Z")

</div>


