# Tons of IMMEDIATE Tasks piling up in cluster state after node failures

**URL:** <https://discuss.elastic.co/t/tons-of-immediate-tasks-piling-up-in-cluster-state-after-node-failures/60645>\
**Category:** Elasticsearch\
**Created:** [September 15, 2016, 8:49pm UTC](https://discuss.elastic.co/t/tons-of-immediate-tasks-piling-up-in-cluster-state-after-node-failures/60645 "2016-09-15T20:49:56Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![mkelkar](https://avatars.discourse-cdn.com/v4/letter/m/ea666f/32.png) [@mkelkar](https://discuss.elastic.co/u/mkelkar)\
**Post date:** [September 15, 2016, 8:49pm UTC](https://discuss.elastic.co/t/tons-of-immediate-tasks-piling-up-in-cluster-state-after-node-failures/60645/1 "2016-09-15T20:49:56Z")

</div>

Hi All,  
I am using ES 1.7.3. I was trying to reallocate shards using shard allocation filtering, and then started seeing tons of tasks piling up on master nodes after a couple of nodes failed to start -

```
2072096 7.7h IMMEDIATE zen-disco-node_failed([polloi-node-96d0f116][GeU0XNjcQXudnSg5jr3m9w][anon-polloi-famin-seventhreetwosrcc-polloi-44-9][inet[/10.100.44.9:29300]]{data=false, client=true}), reason transport disconnected

```

2072194 7.7h IMMEDIATE zen-disco-node\_failed([polloi-node-dc090774][RTolJcurQ1quNoC7mHjLSw][anon-polloi-famin-seventhreetwosrcc-polloi-39-139][inet[/10.100.39.139:29300]]{data=false, client=true}), reason transport disconnected  
2072842 7.7h IMMEDIATE zen-disco-node\_failed([polloi-node-c9cf794d][T4lAoXYOSgSF8awZhcX33Q][anon-polloi-famin-seventhreetwosrcc-polloi-39-139][inet[/10.100.39.139:29300]]{data=false, client=true}), reason transport disconnected  
2073587 7.7h IMMEDIATE zen-disco-node\_failed([polloi-node-c9cf794d][T4lAoXYOSgSF8awZhcX33Q][anon-polloi-famin-seventhreetwosrcc-polloi-39-139][inet[/10.100.39.139:29300]]{data=false, client=true}), reason transport disconnected  
2073738 7.7h IMMEDIATE zen-disco-node\_failed([polloi-node-ac6c5562][Srl7SEs\_QkOYyG0USzyfUg][anon-polloi-famin-seventhreetwosrcc-polloi-39-139][inet[/10.100.39.139:29300]]{data=false, client=true}), reason transport disconnected  
2074334 7.6h IMMEDIATE zen-disco-node\_failed([polloi-node-c9cf794d][T4lAoXYOSgSF8awZhcX33Q][anon-polloi-famin-seventhreetwosrcc-polloi-39-139][inet[/10.100.39.139:29300]]{data=false, client=true}), reason transport disconnected  
2075085 7.6h IMMEDIATE zen-disco-node\_failed([polloi-node-8f766892][L-bdY30xRWy6T64MnIV3Rw][anon-polloi-famin-seventhreetwosrcc-polloi-44-9][inet[/10.100.44.9:29300]]{data=false, client=true}), reason transport disconnected  
2075270 7.6h IMMEDIATE zen-disco-node\_failed([polloi-node-91112d2e][H8\_CQKTBTjaBAsX\_LyUhyw][anon-polloi-famin-seventhreetwosrcc-polloi-39-139][inet[/10.100.39.139:29300]]{data=false, client=true}), reason transport disconnected  
2075831 7.5h IMMEDIATE zen-disco-node\_failed([polloi-node-96d0f116][GeU0XNjcQXudnSg5jr3m9w][anon-polloi-famin-seventhreetwosrcc-polloi-44-9][inet[/10.100.44.9:29300]]{data=false, client=true}), reason transport disconnected

These nodes are terminated and not part of cluster anymore. But even then ES somehow thinks they are...

Also, after relocations were kicked off, I started seeing this -

```
marked shard as initializing, but shard state is [POST_RECOVERY], mark shard as started

```

## And then relocations do not happen, they just get stuck. Trying to change cluster settings also fails because the task gets queued up at the end of pending tasks. Here is what cluster health output looks like

curl localhost:9200/\_cluster/health?pretty  
{  
"cluster\_name" : "xxxx\_elasticsearch",  
"status" : "green",  
"timed\_out" : false,  
"number\_of\_nodes" : 120,  
"number\_of\_data\_nodes" : 30,  
"active\_primary\_shards" : 8603,  
"active\_shards" : 25809,  
"relocating\_shards" : 360,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 0,  
"delayed\_unassigned\_shards" : 0,  
"number\_of\_pending\_tasks" : 288770,  
"number\_of\_in\_flight\_fetch" : 0  
}

Any clues on whats going on?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 17, 2016, 5:21am UTC](https://discuss.elastic.co/t/tons-of-immediate-tasks-piling-up-in-cluster-state-after-node-failures/60645/2 "2016-09-17T05:21:45Z")

</div>

> [@mkelkar](#):
>
> "number\_of\_nodes" : 120,  
> "number\_of\_data\_nodes" : 30,

Do you have a lot of clients, or....?

> [@mkelkar](#):
>
> I am using ES 1.7.3

> [@mkelkar](#):
>
> "number\_of\_pending\_tasks" : 288770,

You should upgrade to 2.X, with the high node count, any cluster state update needs to be sent in full to all nodes.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:19pm UTC](https://discuss.elastic.co/t/tons-of-immediate-tasks-piling-up-in-cluster-state-after-node-failures/60645/3 "2017-07-05T22:19:29Z")

</div>


