# The transport thread cannot be executed by CPU for a long time due to node-level scheduling/resource contention in Elasticsearch cluster

**URL:** https://discuss.elastic.co/t/the-transport-thread-cannot-be-executed-by-cpu-for-a-long-time-due-to-node-level-scheduling-resource-contention-in-elasticsearch-cluster/384060
**Category:** Elasticsearch
**Created:** [December 15, 2025, 9:02am UTC](https://discuss.elastic.co/t/the-transport-thread-cannot-be-executed-by-cpu-for-a-long-time-due-to-node-level-scheduling-resource-contention-in-elasticsearch-cluster/384060 "2025-12-15T09:02:40Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![arT1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/art1/32/137703_2.png) [@arT1](https://discuss.elastic.co/u/arT1)
#### Post date: [December 15, 2025, 9:02am UTC](https://discuss.elastic.co/t/the-transport-thread-cannot-be-executed-by-cpu-for-a-long-time-due-to-node-level-scheduling-resource-contention-in-elasticsearch-cluster/384060/1 "2025-12-15T09:02:40Z")

</div>

The transport thread cannot be executed by CPU for a long time due to node-level scheduling/resource contention in Elasticsearch cluster  
Environment:  
Elasticsearch version: v8.13.3  
Deploy: tar.gz  
Service management: systemd  
The physical host is dual-way NUMA with 80CPU logical cores and 380GB memory. This physical host deploys two hot data nodes, one warm data node, and three logstash instances. This is the only way to deploy due to resource constraints. By observing the node process CPU monitoring, the following scheme is given:

1、CPU Allocation Strategy  
The Systemd isolation configuration is implemented to maximize performance by taking advantage of NUMA architecture features to limit the memory and CPU of the same JVM to the same NUMA Node as much as possible to reduce cross-node memory access latency.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/e/d/ed7730e29560f97c5516dc32940afba5218196a2.png)

> [Service]  
> #...  
> CPUAffinity=1-12 41-52

2、node.processors ( [Thread pools | Elasticsearch Guide [8.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/8.13/modules-threadpool.html#node.processors) )  
Modify elasticsearch.yml

> node.processors: 24

Modify jvm.options

> -XX:ActiveProcessorCount=24

Can this solution solve the problem of transport timeout caused by resource contention?

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [December 15, 2025, 1:08pm UTC](https://discuss.elastic.co/t/the-transport-thread-cannot-be-executed-by-cpu-for-a-long-time-due-to-node-level-scheduling-resource-contention-in-elasticsearch-cluster/384060/2 "2025-12-15T13:08:44Z")

</div>

> [@arT1](#):
>
> The transport thread cannot be executed by CPU for a long time due to node-level scheduling/resource contention in Elasticsearch cluster

Where are you getting this information from?

> [@arT1](#):
>
> Can this solution solve the problem of transport timeout caused by resource contention?

I wouldn’t expect so, no.

---

<div class="post-metadata">

### Author: ![arT1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/art1/32/137703_2.png) [@arT1](https://discuss.elastic.co/u/arT1)
#### Post date: [December 15, 2025, 2:09pm UTC](https://discuss.elastic.co/t/the-transport-thread-cannot-be-executed-by-cpu-for-a-long-time-due-to-node-level-scheduling-resource-contention-in-elasticsearch-cluster/384060/3 "2025-12-15T14:09:07Z")

</div>

master node logs

> [WARN][o.e.c.InternalClusterInfoService] [xxxx] failed to retrieve shard stats from node [3BVEF\_zmSceXcMw7Uv3Qyg]org.elasticsearch.transport.ReceiveTimeoutTransportException: [xxxx][xxxx][indices:monitor/stats[n]] request\_id [1820674096] timed out after [15006ms]
> 
> [WARN][o.e.c.InternalClusterInfoService] [xxxx] failed to retrieve shard stats from node [MVfdVFkoQqeimCtvkvgsjQ]org.elasticsearch.transport.ReceiveTimeoutTransportException: [xxxx][xxxx][indices:monitor/stats[n]] request\_id [1820674098]timed out after [15006ms]
> 
> [ERROR][o.e.x.m.c.c.ClusterStatsCollector] [xxxx] collector [cluster\_stats] timed out when collecting data: nodes [3BVEF\_zmSceXcMw7Uv3Qyg, MVfdV  
> FkoQqeimCtvkvgsjQ, aonbQRHEQxW4deheX-pWww] did not respond within [10s]
> 
> [INFO][o.e.c.r.a.AllocationService] [xxxx] current.health="RED" message="Cluster health status changed from [GREEN] to [RED] (reason: [{xxxxxx}{MVfdVFkoQqeimCtvkvgsjQ}{cJGvljffQ-mQqFxBeOy7Zw}{xxxxxx}{xxxxxx}{xxxxxx}{hs}{8.13.3}{7000099-8503000} reason: followers check retry count exceeded [timeouts=3, failures=0]])." previous.health="GREEN" reason="{xxxxxx}{MVfdVFkoQqeimCtvkvgsjQ}{cJGvljffQ-mQqFxBeOy7Zw}{xxxxxx}{10.109.97.51}{xxxxxx}{hs}{8.13.3}{7000099-8503000} reason: followers check retry count exceeded [timeouts=3, failures=0]"

At the same time, the overall CPU utilization of the physical machine is about 28.8%, and the average 5-minute load is about 63.58. CPU single core usage is between 10% and 50%. thread\_pool.write.rejected and thread\_pool.write.queue increased by about 10,000.The stack monitoring shows that the JVM of this node is normal and there is no Full GC. The network between the cluster nodes is all right.NVMe high performance disk was used for the data disk, and the basic monitoring showed that the disk IO was normal.

Two minutes later the anomalous node rejoins the cluster. After two minutes, the cluster state changes to YELLOW. The cluster state changes to GREEN after 16 minutes.

It's peak logging time, I think it's because the physical machine multiple Elasticsearch processes are not isolated --\> a large number of logs are written to hot data nodes --\> CPU scheduling delay --\> ES transport thread is starved --\> trigger follower check retry countexceeded --\> The cluster state changes to red

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [December 15, 2025, 2:51pm UTC](https://discuss.elastic.co/t/the-transport-thread-cannot-be-executed-by-cpu-for-a-long-time-due-to-node-level-scheduling-resource-contention-in-elasticsearch-cluster/384060/4 "2025-12-15T14:51:55Z")

</div>

> [@arT1](#):
>
> multiple Elasticsearch processes are not isolated --\> a large number of logs are written to hot data nodes --\> CPU scheduling delay --\> ES transport thread is starved --\> trigger follower check retry countexceeded --\> The cluster state changes to red

Possibly, but there’s a bunch of other explanations that seem more likely IMO. Quite possibly this is a bug that’s been fixed in the 18+ months since 8.13.3 was released. The [manual contains the proper troubleshooting process](https://www.elastic.co/docs/troubleshoot/elasticsearch/troubleshooting-unstable-cluster?version=9.2#troubleshooting-unstable-cluster-follower-check) that you need to follow, although it’d be simpler to upgrade to a newer version to pick up all the relevant bugfixes first.
