# Elasticsearch cluster have millions of pending tasks

**URL:** <https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349>\
**Category:** Elasticsearch\
**Created:** [May 7, 2021, 2:05am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349 "2021-05-07T02:05:56Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 7, 2021, 2:05am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/1 "2021-05-07T02:05:56Z")

</div>

i have a es cluster , 7.4 version .

There are a lot of pending tasks around 8 a.m. every day。 pending task list info：

639479 "source": "ilm-execute-cluster-state-steps",  
82666 "source": "ilm-move-to-step",  
186 "source": "cluster\_reroute(reroute after starting shards)",  
12 "source": "ilm-set-step-info",  
1 "source": "update-settings",

my cluster info：

my health info:  
epoch timestamp cluster status node.total node.data shards pri relo init unassign pending\_tasks max\_task\_wait\_time active\_shards\_percent  
1620352457 01:54:17 sre-elasticsearch green 34 30 18001 11216 0 0 0 610953 46.8m 100.0%

my node info:  
master node 3  
hot node 10  
warm node 14  
cold node 6  
client node 1

my ilm info:  
{  
"hotwarm-norollover-15days-for-hot-index" : {  
"version" : 18,  
"modified\_date" : "2021-04-25T11:13:59.657Z",  
"policy" : {  
"phases" : {  
"warm" : {  
"min\_age" : "5d",  
"actions" : {  
"allocate" : {  
"number\_of\_replicas" : 1,  
"include" : { },  
"exclude" : { },  
"require" : {  
"box\_type" : "warm"  
}  
},  
"set\_priority" : {  
"priority" : 100  
}  
}  
},  
"cold" : {  
"min\_age" : "9d",  
"actions" : {  
"allocate" : {  
"number\_of\_replicas" : 0,  
"include" : { },  
"exclude" : { },  
"require" : {  
"box\_type" : "cold"  
}  
},  
"freeze" : { }  
}  
},  
"hot" : {  
"min\_age" : "0ms",  
"actions" : {  
"set\_priority" : {  
"priority" : 100  
}  
}  
},  
"delete" : {  
"min\_age" : "15d",  
"actions" : {  
"delete" : { }  
}  
}  
}  
}  
}  
}

I have already created the index ahead of time at 1:30 a.m。but there are still many pending tasks。I don't know what caused it。  
Can anyone tell me why？ thanks。

---

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 8, 2021, 2:01am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/2 "2021-05-08T02:01:04Z")

</div>

anyone konw?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 10, 2021, 12:29am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/3 "2021-05-10T00:29:53Z")

</div>

> [@fengxiaobai](#):
>
> i have a es cluster , 7.4 version .

FYI 7.4 just reached [EOL](https://www.elastic.co/support/eol), you will want to upgrade ASAP.

> [@fengxiaobai](#):
>
> I have already created the index ahead of time at 1:30 a.m

That is not a good idea, it's a waste of resources. Let Elasticsearch create them as needed.

What is the output from the `_cluster/stats?pretty&human` API?

---

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 10, 2021, 1:27am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/4 "2021-05-10T01:27:51Z")

</div>

@warkolm Thanks for your response. Are you suggesting an upgrade to fix the problem？

GET \_cluster/stats?pretty&human ,info list(i get this info at 9:20 a.m; but 8:00 a.m, maybe it's a little different):

{  
"\_nodes" : {  
"total" : 34,  
"successful" : 34,  
"failed" : 0  
},  
"cluster\_name" : "sre-elasticsearch",  
"cluster\_uuid" : "1WLLkjixT4WGg1TJ8li2zQ",  
"timestamp" : 1620609728494,  
"status" : "green",  
"indices" : {  
"count" : 10676,  
"shards" : {  
"total" : 18279,  
"primaries" : 11312,  
"replication" : 0.6158946251768034,  
"index" : {  
"shards" : {  
"min" : 1,  
"max" : 16,  
"avg" : 1.7121581116523041  
},  
"primaries" : {  
"min" : 1,  
"max" : 8,  
"avg" : 1.0595728737354815  
},  
"replication" : {  
"min" : 0.0,  
"max" : 1.0,  
"avg" : 0.6214874484825778  
}  
}  
},  
"docs" : {  
"count" : 46945133375,  
"deleted" : 1689084  
},  
"store" : {  
"size" : "32.6tb",  
"size\_in\_bytes" : 35905473914951  
},  
"fielddata" : {  
"memory\_size" : "4.2mb",  
"memory\_size\_in\_bytes" : 4406328,  
"evictions" : 0  
},  
"query\_cache" : {  
"memory\_size" : "290.7mb",  
"memory\_size\_in\_bytes" : 304876793,  
"total\_count" : 23108719,  
"hit\_count" : 4070526,  
"miss\_count" : 19038193,  
"cache\_size" : 14305,  
"cache\_count" : 30540,  
"evictions" : 16235  
},  
"completion" : {  
"size" : "0b",  
"size\_in\_bytes" : 0  
},  
"segments" : {  
"count" : 138490,  
"memory" : "28.6gb",  
"memory\_in\_bytes" : 30721990535,  
"terms\_memory" : "15gb",  
"terms\_memory\_in\_bytes" : 16195784675,  
"stored\_fields\_memory" : "12.1gb",  
"stored\_fields\_memory\_in\_bytes" : 13012623640,  
"term\_vectors\_memory" : "0b",  
"term\_vectors\_memory\_in\_bytes" : 0,  
"norms\_memory" : "164.4mb",  
"norms\_memory\_in\_bytes" : 172430208,  
"points\_memory" : "1.1gb",  
"points\_memory\_in\_bytes" : 1268757264,  
"doc\_values\_memory" : "69mb",  
"doc\_values\_memory\_in\_bytes" : 72394748,  
"index\_writer\_memory" : "2.4gb",  
"index\_writer\_memory\_in\_bytes" : 2610843890,  
"version\_map\_memory" : "5.1mb",  
"version\_map\_memory\_in\_bytes" : 5376215,  
"fixed\_bit\_set" : "99.1mb",  
"fixed\_bit\_set\_memory\_in\_bytes" : 103998912,  
"max\_unsafe\_auto\_id\_timestamp" : 1620604808932,  
"file\_sizes" : { }  
}  
},  
"nodes" : {  
"count" : {  
"total" : 34,  
"coordinating\_only" : 0,  
"data" : 30,  
"ingest" : 30,  
"master" : 3,  
"ml" : 1,  
"voting\_only" : 0  
},  
"versions" : [  
"7.4.0"  
],  
"os" : {  
"available\_processors" : 744,  
"allocated\_processors" : 744,  
"names" : [  
{  
"name" : "Linux",  
"count" : 34  
}  
],  
"pretty\_names" : [  
{  
"pretty\_name" : "CentOS Linux 7 (Core)",  
"count" : 34  
}  
],  
"mem" : {  
"total" : "1.4tb",  
"total\_in\_bytes" : 1555088515072,  
"free" : "133.2gb",  
"free\_in\_bytes" : 143066071040,  
"used" : "1.2tb",  
"used\_in\_bytes" : 1412022444032,  
"free\_percent" : 9,  
"used\_percent" : 91  
}  
},  
"process" : {  
"cpu" : {  
"percent" : 300  
},  
"open\_file\_descriptors" : {  
"min" : 1472,  
"max" : 11777,  
"avg" : 6738  
}  
},  
"jvm" : {  
"max\_uptime" : "102.9d",  
"max\_uptime\_in\_millis" : 8898528373,  
"versions" : [  
{  
"version" : "13",  
"vm\_name" : "OpenJDK 64-Bit Server VM",  
"vm\_version" : "13+33",  
"vm\_vendor" : "AdoptOpenJDK",  
"bundled\_jdk" : true,  
"using\_bundled\_jdk" : true,  
"count" : 34  
}  
],  
"mem" : {  
"heap\_used" : "266.6gb",  
"heap\_used\_in\_bytes" : 286297579840,  
"heap\_max" : "725.9gb",  
"heap\_max\_in\_bytes" : 779466833920  
},  
"threads" : 8241  
},  
"fs" : {  
"total" : "196.9tb",  
"total\_in\_bytes" : 216603026358272,  
"free" : "164.1tb",  
"free\_in\_bytes" : 180494084050944,  
"available" : "155.7tb",  
"available\_in\_bytes" : 171248387346432  
},  
"plugins" : ,  
"network\_types" : {  
"transport\_types" : {  
"security4" : 34  
},  
"http\_types" : {  
"security4" : 34  
}  
},  
"discovery\_types" : {  
"zen" : 34  
},  
"packaging\_types" : [  
{  
"flavor" : "default",  
"type" : "rpm",  
"count" : 34  
}  
]  
}  
}

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 10, 2021, 1:32am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/5 "2021-05-10T01:32:45Z")

</div>

Please format your code/logs/config using the `</>` button, or markdown style back ticks. It helps to make things easy to read which helps us help you 🙂

> [@fengxiaobai](#):
>
> @warkolm Thanks for your response. Are you suggesting an upgrade to fix the problem？

It's possible, there is always bug fixes and performance improvements and it'd be a recommended step as part of troubleshooting.

> [@fengxiaobai](#):
>
> "version" : "13",  
> "vm\_name" : "OpenJDK 64-Bit Server VM",  
> "vm\_version" : "13+33",  
> "vm\_vendor" : "AdoptOpenJDK",  
> "bundled\_jdk" : true,  
> "using\_bundled\_jdk" : true,  
> "count" : 34  
> }

You should definitely upgrade your JVM, that's pretty old these days.

It also looks like your average shard size is about 5GB, which is inefficient given you have that many shards. You should look to shrink some of your indices and adjust your index creation strategy. Look at using [ILM](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-lifecycle-management.html) as well.

---

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 10, 2021, 1:58am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/6 "2021-05-10T01:58:38Z")

</div>

@warkolm  
Please format your code/logs/config using the `</>` button, or markdown style back ticks. It helps to make things easy to read which helps us help you  
about this i'm so sorry . now i konw.  
Bye the way , this version of the JDK comes with ES. So i need to upgrade to change the JDK version

Is there any other solution？

---

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 10, 2021, 2:06am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/7 "2021-05-10T02:06:17Z")

</div>

> [@warkolm](#):
>
> It also looks like your average shard size is about 5GB

Please tell me，how is this value calculated？

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 10, 2021, 2:11am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/8 "2021-05-10T02:11:55Z")

</div>

> [@fengxiaobai](#):
>
> Please tell me，how is this value calculated？

This;

> [@fengxiaobai](#):
>
> "size" : "32.6tb",

Divided by;

> [@fengxiaobai](#):
>
> "total" : 18279,

---

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 10, 2021, 2:32am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/9 "2021-05-10T02:32:41Z")

</div>

> [@warkolm](#):
>
> > [@fengxiaobai](#):
> >
> > "size" : "32.6tb",
> 
> Divided by;
> 
> > [@fengxiaobai](#):
> >
> > "total" : 18279,

Forgive me . I get this value:

`echo "scale=2;(32.6*1024)/18279"|bc=1.82`

not 5GB. I can't understand.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 10, 2021, 3:26am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/10 "2021-05-10T03:26:43Z")

</div>

Yeah sorry, my math was bad. That's still not good though.

---

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 10, 2021, 5:06am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/11 "2021-05-10T05:06:32Z")

</div>

> [@warkolm](#):
>
> Yeah sorry, my math was bad. That's still not good though.

ok. I'm going to optimize the sharding problem。  
thansks。

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 10, 2021, 6:05am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/12 "2021-05-10T06:05:13Z")

</div>

Based on the stats it looks like you are generating over 300 daily indices that as @warkolm pointed out are very small. Given that you seem to be creating around 1 TB of indices (primary and replica) per day that is quite inefficient and means that there are a lot of indices to move at specific times as they likely all created within a small time frame.

I would recommend consolidating the data into a much smaller number of daily indices and potentially increase the number of primary shards if required. Aim for a shard size of 30GB to 50GB and I think the situation should improve.

The retention period on the cold nodes also seem quite short so 1/15 of the data held there (1TB or so?) will be replaced there every day. If you have very slow storage that might also slow things down and contribute to the problem.

---

<div class="post-metadata">

**Author:** ![fengxiaobai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fengxiaobai/32/88361_2.png) [@fengxiaobai](https://discuss.elastic.co/u/fengxiaobai)\
**Post date:** [May 11, 2021, 3:08am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/13 "2021-05-11T03:08:46Z")

</div>

@Christian_Dahlqvist Thanks for your response。

Yes, generate 600+ day-level indexes per day about 8:00 am.

Now, in order to solve this problem,I have already created the index ahead of time at 1:30 a.m

The current scenario is that the business logs are stored in a cluster. But the business log has many categories and creates an index for each category.

At the same time, most services have a very small amount of logging.  
Now the minimum shard has been set to 1 in the template.

There is no other solution now. 😂

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 11, 2021, 3:14am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/14 "2021-05-11T03:14:48Z")

</div>

> [@fengxiaobai](#):
>
> There is no other solution now. 😂

Look at using [ILM](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-lifecycle-management.html) as I mentioned.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 11, 2021, 5:33am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/15 "2021-05-11T05:33:57Z")

</div>

Why does each category have each own index? Why can you not store all categories in a single index, or at least a considerably smaller number?

The current approach sounds very inefficient and is likely contributing to your problems.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 8, 2021, 5:34am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-have-millions-of-pending-tasks/272349/16 "2021-06-08T05:34:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
