# High CPU Usage on a few data nodes / Hotspotting of data

**URL:** <https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954>\
**Category:** Elasticsearch\
**Created:** [December 10, 2025, 11:45am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954 "2025-12-10T11:45:07Z")\
**Posts on this page:** 20\
**Page:** 6

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 1:10pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/110 "2025-12-15T13:10:41Z")

</div>

> [@leandrojmp](#):
>
> No, I would mention to check the mount options because per default ext4 are mounted with the parameter relatime then it should be changed to noatime, this adds a boost in speed, but you already shared that the mount options include `noatime`.

Do you recommend switching to XFS instead?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 15, 2025, 1:52pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/111 "2025-12-15T13:52:17Z")

</div>

> [@Lakshya\_Gupta](#):
>
> Do you recommend switching to XFS instead?

I don't think this will make any difference, besides that, to make this change you would need to reinstall and reconfigure your nodes.

Being honest, I'm not sure you will be able to solve this issue as it seems to be a problem with the design choice, you mentioned that you are having this problem for the past 4 years and that it is getting worst.

Not sure if it was asked already, but what is the refresh\_interval in your indices? The default would be `1s`, but increasing this value can help with the index performance and reduce the CPU usage, but it can also have other effects on search, so you would need to test.

Also, I could not find, but how many indices and what is the total size of your cluster? You mentioned something around 60 data nodes with 320 GB of disk, which seems a weird size, considering the watermark you would have something close to 16 TB of data, which seems a small size to be split on 60 nodes, I had an on-premises cluster with a hot tier of 16 TB split into 4 nodes only.

I think you will need to review your cluster design to be able to fix this problem.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 6:57pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/112 "2025-12-15T18:57:12Z")

</div>

> [@leandrojmp](#):
>
> the refresh\_interval in your indices

30 seconds.

> [@leandrojmp](#):
>
> I think you will need to review your cluster design to be able to fix this problem.

We currently have 63 data nodes cluster (remaining are 3 master and 5 co-ordinator)…these 63 data nodes have 340/440 gb storage attached to each of them…can you suggest a better design? Also, we only use 3TB out of the 16TB storage

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 7:14pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/113 "2025-12-15T19:14:25Z")

</div>

> [@RainTown](#):
>
> `iostat -ztcxd 60 60` or `iostat -ztcxd 10 360`

Wasn’t able to generate enough load today 😕 Will get back on this tomorrow.

---

<div class="post-metadata">

**Author:** ![linkerc](https://avatars.discourse-cdn.com/v4/letter/l/13edae/32.png) [@linkerc](https://discuss.elastic.co/u/linkerc)\
**Post date:** [December 15, 2025, 7:15pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/114 "2025-12-15T19:15:59Z")

</div>

dropping replica down from 2 to 1 should help.

From my testing in our production cluster, more replicas than 1 doesn’t seem to speed up read noticeably , but it incurs heavy CPU during write. Especially when all of your writes are update.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [December 15, 2025, 7:16pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/115 "2025-12-15T19:16:00Z")

</div>

> [@Lakshya\_Gupta](#):
>
> We currently have 63 data nodes cluster (remaining are 3 master and 5 co-ordinator)…these 63 data nodes have 340/440 gb storage attached to each of them…can you suggest a better design? Also, we only use 3TB out of the 16TB storage

Originally (5 days ago):

> [@Lakshya\_Gupta](#):
>
> - ES Version 7.17.0
> - 60 data nodes cluster
> - 50 primary shards and
> - 2 replica for each primary shards

So 3 new data nodes added?

Anyways, you keep posting these things like these are real hardware. You have _virtual_ machines. We (at least I) don't know yet

- How the storage is "sliced and diced" before presented to the VMs, and what are you sharing it with?
- How much storage there is - you quoted a colleague saying you had 2x 3.8TB PM1733 per server which "are shared among the VMs scheduled on that server". Well, that's almost 8TB per server.
- We "only use 3TB out of the 16TB storage". That's say 2 servers? Or there are 20 servers, or 50, or 250, or .. ?
- Is there any over-provisioning, of anything, going on here?
- What _exactly_ is the Virtual layer? Some VMware product? Something else?

I alos note:

> [@Lakshya\_Gupta](#):
>
> _The read and write IOPs limits for p4i instances (this is the type of instance we use) is 100 per GiB_.”

p4i isn't an instance type I recognize, so I'm guessing this is some kind of corporate cloud solution copying (a bit) from AWS type terminology. Maths tells me

100 IOPs / GiB

would mean

32,000 IOPS / 320 GiB (which we have no evidence you are seeing anything remotely close to that)

on assumption of linear mapping. but it would also translate to

380,000 IOPS / 3.8TB disk

and I'm a little skeptical on that.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 7:24pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/116 "2025-12-15T19:24:28Z")

</div>

> [@RainTown](#):
>
> So 3 new data nodes added?

It was 63 since the start 😅

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 7:30pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/117 "2025-12-15T19:30:53Z")

</div>

> [@RainTown](#):
>
> - How the storage is "sliced and diced" before presented to the VMs, and what are you sharing it with?
> - How much storage there is - you quoted a colleague saying you had 2x 3.8TB PM1733 per server which "are shared among the VMs scheduled on that server". Well, that's almost 8TB per server.
> - We "only use 3TB out of the 16TB storage". That's say 2 servers? Or there are 20 servers, or 50, or 250, or .. ?
> - Is there any over-provisioning, of anything, going on here?
> - What _exactly_ is the Virtual layer? Some VMware product? Something else?

Will come back on these tomorrow. I have very less idea about this, will need to discuss internally.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 7:31pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/118 "2025-12-15T19:31:28Z")

</div>

> [@linkerc](#):
>
> dropping replica down from 2 to 1 should help.

Thanks for that suggestion, but we can’t go with that 😕 to make our system fault tolerant, we need at least 2 replica shards.

---

<div class="post-metadata">

**Author:** ![linkerc](https://avatars.discourse-cdn.com/v4/letter/l/13edae/32.png) [@linkerc](https://discuss.elastic.co/u/linkerc)\
**Post date:** [December 15, 2025, 7:49pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/119 "2025-12-15T19:49:11Z")

</div>

That’s what we thought too. But we eventually reduce all but few indices down to 1 replica. Unless you are expecting multiple nodes to be down simultaneously. ES is pretty good with shard assignments.

Granted you have custom routing. I assume you let the replica assigns automatically?

Another potential solution is to increase the shard of your heavy indices. Spreading the CPU across more nodes will reduce hotspot.

You will need to reindex for your use case. I assume you guys don’t create new index periodically since it’s update.

What we have done is creating monthly index say “user\_202512”. The index contains user info and constantly being updated, etc.

When the user’s data is being updated, we write the data to 2 indices. “user\_202512” & “user\_202601”.

We simply write current user data to both this month and next month.

The reader simply reads “user\_currentMonth” to get the latest data.

That way, our “user” index will rotate monthly to allow changes to settings and/or mappings. Periodically creating new index in ES is beneficial. It spots stale data (bad application) very easily.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 16, 2025, 4:47am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/120 "2025-12-16T04:47:20Z")

</div>

> [@linkerc](#):
>
> I assume you let the replica assigns automatically?

Correct.

> [@linkerc](#):
>
> I assume you guys don’t create new index periodically since it’s update.

We cannot create a new index because our use case involves updates to data which can be 3 years old as well.

Also, we tried re-indexing by increasing the routing\_partition\_size from 5 to 8, but after re-indexing, all the data for a single routing key went to 1 shard (instead of going to 8 shards).

We used the below command, are you aware why this could have happened?

```auto
POST _reindex?slices=50&requests_per_second=-1&wait_for_completion=false
{
  "conflicts": "proceed",
  "source": {
    "index": "listings"
  },
  "dest": {
    "index": "listings_v2",
    "routing": "keep"
  }
}

```

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 16, 2025, 5:09am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/121 "2025-12-16T05:09:07Z")

</div>

In the next post I’ll make sure to consolidate and post all the below things

- open questions around the internals of our VM etc
- iostat metrics along with ES graphs during period of high CPU spikes
- cluster configuration

Since collecting this data can take 1-2 days, I’ll post all the details in a single shot, maybe then we should be able to find a solve 😃

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 16, 2025, 8:56am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/122 "2025-12-16T08:56:11Z")

</div>

Hi @RainTown @Christian_Dahlqvist , here are some of the new points:

1. storage is sliced using LVMs and presented to VMs using virtio-blk

2. virtualization layer - qemu/kvm

3. we have 100 iops configured per gb at a blocksize of 4kb from the infra team. (The 100 iops per gb is committed only for a blocksize of 4kb)

4. I checked internally within my team and we don’t seem to have a noise neighbour problem

5. cluster configuration

6. this is how our disk is configured

I’m generating load meanwhile to run the iostat commands, will share the results soon.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [December 16, 2025, 9:12am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/123 "2025-12-16T09:12:13Z")

</div>

So this is not elasticsearch question. now, its just IO tuning in a KVM environment.

> [@Lakshya\_Gupta](#):
>
> ```auto
> <iotune>
> <read_iops_sec>32000</read_iops_sec>
> <write_iops_sec>32000</write_iops_sec>
> </iotune>
> 
> ```

take that on one of your hot nodes, for a test, and see what you get.

But I am guessing you need to enable multi-queue, add `queues='4'` (or some bigger number) in the config, after the `discard=..`. I believe, if unset, the default is 1.

On your VM, see if there are dirs under:

`ls /sys/block/vdb/mq/`

If you have just a directory called `0` then you have 1 queue.

e.g. for me

```auto
[root@rhel10x1 ~]# ls /sys/block/vda/mq/
0 1 2 3

```

The queue count should not exceed processor count.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 16, 2025, 9:32am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/124 "2025-12-16T09:32:54Z")

</div>

ls /sys/block/vdb/mq/

Got below output

```auto
0 1 10 11 12 13 14 15 16 17 18 19 2 3 4 5 6 7 8 9

```

Number of cores = 20, should be good then?

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 16, 2025, 10:42am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/125 "2025-12-16T10:42:49Z")

</div>

There is 1 hypothesis we have right now, just thinking out loud.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/9/79de35383739761cd961ac5b618936704651a05b.png)

Let’s debug the highlighted node - From the above graph, we can see that the disk throughput is reaching a max of 385mbps = 3,94,240 KBs.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/6/3647bd91e146e7ba9e01050bb218a45032ad02c0.png)

From this graph, we can see that the disk iops is 3.21K.

Hence, block size = total disk throughput / disk iops = 3,94,240 / 3210 = 122.8161993769 KB.

Our infra team has committed a write iops of 100 per gb only for a block size of 4Kb, however, our block size seems to be reaching ~ 122KB, does that mean we are not actually getting 100iops per gb and are hitting the wall?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 16, 2025, 11:12am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/126 "2025-12-16T11:12:06Z")

</div>

> [@Lakshya\_Gupta](#):
>
> Sure, this metric and the 1 suggested by Christian “Did you have a chance to verify that your deduplication logic is working as expected so we can remove that as a potential issue/contributing factor?”…I’ll try to get these both and post it here as soon as I’m able to see the load.

Did we get anywhere on this?

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 16, 2025, 11:25am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/127 "2025-12-16T11:25:35Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Did we get anywhere on this?

Not yet 😕

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [December 16, 2025, 11:47am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/128 "2025-12-16T11:47:37Z")

</div>

> [@Lakshya\_Gupta](#):
>
> There is 1 hypothesis we have right now, just thinking out loud.

Yes, its a decent theory.

And I think I took us down a wrong alley. The

> [@Lakshya\_Gupta](#):
>
> ls /sys/block/vdb/mq/

> [@Lakshya\_Gupta](#):
>
> `0 1 10 11 12 13 14 15 16 17 18 19 2 3 4 5 6 7 8 9`

set of 20 queues wont help much here, as we have only one (hot) shard per node, right? So the available parallelism doesn't help much, as each shard has one writer thread.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 16, 2025, 12:25pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/129 "2025-12-16T12:25:08Z")

</div>

> [@RainTown](#):
>
> as we have only one (hot) shard per node, right?

Not necessary, in total we have 5 hot primary shards, meaning in total 15 hot shards (5 primary and 10 of their replicas)…considering our cluster configuration, every node usually has 3 shards, so in the worst case we can get all 15 shards (primary + replica) boiled down to 5 data nodes.

So, in the worst case, a data node can have 3 hot shards.

[Previous page](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954.md?page=5)

[Next page](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954.md?page=7)
