# High CPU Usage on a few data nodes / Hotspotting of data

**URL:** <https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954>\
**Category:** Elasticsearch\
**Created:** [December 10, 2025, 11:45am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954 "2025-12-10T11:45:07Z")\
**Posts on this page:** 20\
**Page:** 5

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 14, 2025, 11:25am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/90 "2025-12-14T11:25:31Z")

</div>

> [@Lakshya-Gupta](#):
>
> True, but we have 1 billion documents 😕

Updates to large number of different documents does not necessarily result in a lot of small segments.

> [@Lakshya-Gupta](#):
>
> We are using a key-value based redis data structure in order to implement deduping. The id here will be same as the document id in Elasticsearch, hence, the requests coming up are queued for 30 seconds, so that requests on the same key are overridden in the map, and after 30 seconds, only 1 update per ID is done.

Are you doing this using a single Redis instance for 1 billion documents?

I assume you are writing to Elasticsearch concurrently using a number of clients. How do you manage the partition of the documentID space across these to avoid duplicate updates?

Are you only sending update requests every 30 seconds and idling in between or is the processing stacked in some way?

If this works the way you describe I do not think you should see a lot of small segments created even for 1 billion total documents. Do you know the document ID of a document that is heavily updated? If you have this I guess you could retrieve this from Elasticsearch using the [get document API](https://www.elastic.co/guide/en/elasticsearch/reference/8.19/docs-get.html#docs-get-api-response-body) once per minute and verify that the `_version` field increments at a nice slow pace in line with the expected update frequency.

It is quite possible that frequent updates is not at all causing the issues you are describing, so it would be useful to rule that out if possible. You may also use the [cat segments API](https://www.elastic.co/guide/en/elasticsearch/reference/8.19/cat-segments.html) to check the size of the new segments generated.

---

<div class="post-metadata">

**Author:** ![Lakshya-Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya-gupta/32/146362_2.png) [@Lakshya-Gupta](https://discuss.elastic.co/u/Lakshya-Gupta)\
**Post date:** [December 14, 2025, 11:51am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/91 "2025-12-14T11:51:59Z")

</div>

Okay, I think I didn’t explain it quite well, let me walk you through the entire indexing process we use.

1. We have 1 billion documents, and 99% of the requests we get are update requests; not many new documents are inserted.
2. We have an initial topology which writes all these update requests at a document ID level to a redis system wherein the key is the document ID.
3. While ingesting data to redis, we set the polling time as +30 seconds from the first time a document for a particular key was inserted, this means that if we get 20 update requests for the same document id, there will only be 1 key-value pair in redis with the polling time of 30 secoonds after the first update request came in.
4. We have an apache storm topology where we can configure the parallel number of workers, and these pick the requests from the redis database. This topology then batches the result and sends a bulk update request to Elasticsearch.

This is the complete process, hope that clarifies things…let me know in case of any more additional details required 😃 .

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 14, 2025, 11:55am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/92 "2025-12-14T11:55:33Z")

</div>

I would run the checks I mentioned in order to try to verify it is indeed working as intended.

Does this mean that you are only indexing in bursts every 30 seconds or so?

---

<div class="post-metadata">

**Author:** ![Lakshya-Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya-gupta/32/146362_2.png) [@Lakshya-Gupta](https://discuss.elastic.co/u/Lakshya-Gupta)\
**Post date:** [December 14, 2025, 11:55am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/93 "2025-12-14T11:55:45Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> It is quite possible that frequent updates is not at all causing the issues you are describing, so it would be useful to rule that out if possible

True, and from the graphs I’ve visualised, im 99% sure that the cpu spike graph is matching that if the number of segments graph. As soon as the number of segments icreases, I always see a cpu increase.

---

<div class="post-metadata">

**Author:** ![Lakshya-Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya-gupta/32/146362_2.png) [@Lakshya-Gupta](https://discuss.elastic.co/u/Lakshya-Gupta)\
**Post date:** [December 14, 2025, 12:06pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/94 "2025-12-14T12:06:34Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> I would run the checks I mentioned in order to try to verify it is indeed working as intended.

Sure, I will run them during high indexing load and let you know.

> [@Christian\_Dahlqvist](#):
>
> Does this mean that you are only indexing in bursts every 30 seconds or so?

No no, since it’s a continuous streaming topology, you can say that the burst is happening every second (for the previous 30 seconds request), somewhat like a sliding window.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [December 14, 2025, 12:33pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/95 "2025-12-14T12:33:05Z")

</div>

/aside - a while ago I suggested to one of the admins to create an kibana/elasticsearch cluster populated with this forums comments. Cos, you know, for search. This (longish) thread demonstrates the use case. Did someone other than me mention VMware in this thread? No so easy to see without a lot of scrolling?

Picking a few small things:

- you can find out the filesystem by just typing the `mount` command on a data node, look for your data partition, likely vdb or vdb1, and you will see ext4/xfs/... Paste here if unsure. `lsblk -f` would also tell you
- you have mentioned few times "320GB/340gb ssd" but the disk quoted is 440+ GB, I hope this disk is _exclusively_ allocated to elasticsearch?

> [@Lakshya\_Gupta](#):
>
> ```auto
> logical name: /dev/vdb
> size: 446GiB (478GB)
> 
> ```

- If you are using your own `_id`, indexing _new_ docs has a higher cost when the index size gets really large. Not likely critical, but you should know that.
- I might have confused this with another thread, are your "VMs" from VMware or some other virtualization platform?
- You seemed to very easily increase (v)CPU from 20 to 32, and it's a little mystery why it seemed to make no difference at all. But in end a VM runs on a host, maybe with a bunch of other VMs on same host. VMware, and other platforms, will let you over-provision resources massively. It would likely even allow you to put 2 or more data nodes on same host, unless you specify otherwise. Point is elasticsearch works best on dedicated resources, especially if you want best performance. The assumption here is you are in a corp IT environment, where _my experience_ is that the 'VM team" don't know/care a lot about your use cases, they just satisfy your requests as asked. So onus is on you to be asking right questions.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [December 14, 2025, 12:55pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/96 "2025-12-14T12:55:53Z")

</div>

> [@Lakshya-Gupta](#):
>
> No no, since it’s a continuous streaming topology, you can say that the burst is happening every second (for the previous 30 seconds request), somewhat like a sliding window.

Between second N and second N+1, their preceding 30-second window has 29 seconds overlapping. Might be me, but I have not understood your explanation here, sorry.

Can you share the index monitoring screens from kibana, equivalent of mine below but for your much more interesting index, with say a typical 1 hour time interval.

 ![Screenshot 2025-12-14 at 13.53.04](https://us1.discourse-cdn.com/elastic/original/3X/d/8/d81ec1ecf899c370e33705277709decf11330e2b.jpeg)  
 ![Screenshot 2025-12-14 at 13.53.13](https://us1.discourse-cdn.com/elastic/original/3X/2/2/227c1d8f6268cfd606c16fd51d4d47cafaf4dfc4.jpeg)  
 ![Screenshot 2025-12-14 at 13.53.08](https://us1.discourse-cdn.com/elastic/original/3X/f/c/fcb8baa5907f0b292e84e0a5501d84e083a40b52.jpeg)

---

<div class="post-metadata">

**Author:** ![Lakshya-Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya-gupta/32/146362_2.png) [@Lakshya-Gupta](https://discuss.elastic.co/u/Lakshya-Gupta)\
**Post date:** [December 14, 2025, 4:15pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/97 "2025-12-14T16:15:17Z")

</div>

Here are the snapshots during high cpu spikes

 ![IMG_0362](https://us1.discourse-cdn.com/elastic/original/3X/3/f/3f2d448b2863ff460a2e059e9974377961a1b0d8.jpeg)

 ![IMG_0363](https://us1.discourse-cdn.com/elastic/original/3X/1/1/11ef98b2c76f934cb7d7f0f58dbcbe3829f27f95.jpeg)

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 14, 2025, 4:23pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/98 "2025-12-14T16:23:47Z")

</div>

> [@RainTown](#):
>
> a while ago I suggested to one of the admins to create an kibana/elasticsearch cluster populated with this forums comments. Cos, you know, for search. This (longish) thread demonstrates the use case. Did someone other than me mention VMware in this thread? No so easy to see without a lot of scrolling?

Wow, that’s crazy 😮.

> [@RainTown](#):
>
> `mount` command

Got this response @leandrojmp

`/dev/vdb1 on /var/lib/elasticsearch type ext4 (rw,noatime,nodelalloc)`

> [@RainTown](#):
>
> `lsblk -f`

Got the below response

```auto
NAME FSTYPE FSVER LABEL UUID FSAVAIL FSUSE% MOUNTPOINT
vda                                                                            
├─vda1 ext4 1.0 363680f3-0237-4987-a303-e236995e1017 6G 33% /
├─vda14                                                                        
└─vda15 vfat FAT16 E76C-DD8F 113M 9% /boot/efi
vdb                                                                            
└─vdb1 ext4 1.0 9263cd42-3be5-4efb-95ed-7b0f218f3589 299G 27% /var/lib/elasticsearch

```

> [@RainTown](#):
>
> - you have mentioned few times "320GB/340gb ssd" but the disk quoted is 440+ GB, I hope this disk is _exclusively_ allocated to elasticsearch?
> 
> > [@Lakshya\_Gupta](#):
> >
> > ```auto
> > 
> > ```

We have a heterogenous mix of data nodes, some of them are 340gb and some of them are 440gb, the node which I’m running the commands on is that of 440gb.

> [@RainTown](#):
>
> I might have confused this with another thread, are your "VMs" from VMware or some other virtualization platform?

Will get back on this.

> [@RainTown](#):
>
> So onus is on you to be asking right questions.

Understood, can you let me know what all I shall be asking the IT team?

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [December 14, 2025, 6:31pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/99 "2025-12-14T18:31:24Z")

</div>

> [@RainTown](#):
>
> Can you share the index monitoring screens from kibana ... with say a typical 1 hour time interval.

> [@Lakshya-Gupta](#):
>
> Here are the snapshots during high cpu spikes

From July? For a period of a few days, rather than the 1 hour I asked (and I implicitly meant current data). Anyways, they show what you wrote to me, there were periods of high CPU for hours. You have changed a lot of things since July. You ran the `lsblk` command today, why not get the screenshots today? The IOPS are not high at all, not ins same ballpark as the Samsung product spec sheet you shared.

You have settled that you are using ext4 as filesystem. Other might suggests tweaks there.

But for me outstanding is really the larger scale setup - be it VMware or something else. How is IO provided to the VMs. In detail.

> [@Lakshya\_Gupta](#):
>
> ... can you let me know what all I shall be asking the IT team?

It's all written in this thread. You have an application (elasticsearch) that needs significant IO capability, specifically IOPS, **on many instances simultaneously**. Your elasticsearch instances need as close to exclusive access to resources (CPU and memory too btw) as is possible.

---

<div class="post-metadata">

**Author:** ![Lakshya-Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya-gupta/32/146362_2.png) [@Lakshya-Gupta](https://discuss.elastic.co/u/Lakshya-Gupta)\
**Post date:** [December 14, 2025, 6:45pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/100 "2025-12-14T18:45:13Z")

</div>

> [@RainTown](#):
>
> You have changed a lot of things since July. You ran the `lsblk` command today, why not get the screenshots today? The IOPS are not high at all, not ins same ballpark as the Samsung product spec sheet you shared.

I had those screenshots handy, hence shared those…but those were the final screenshots after all the optimizations (adding dedup layer, reduce writes, add co-ordinate nodes etc) we had done.

I can share some recent graphs in the morning as well, but the data is pretty much the same.

> [@RainTown](#):
>
> You have an application (elasticsearch) that needs significant IO capability, specifically IOPS, **on many instances simultaneously**. Your elasticsearch instances need as close to exclusive access to resources (CPU and memory too btw) as is possible.

Sure, I’ll check this with the team which manages these virtual machine instances and then back to you with the full details 😃

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 3:20am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/101 "2025-12-15T03:20:01Z")

</div>

> [@RainTown](#):
>
> You have settled that you are using ext4 as filesystem. Other might suggests tweaks there.

@leandrojmp any tweaks you might want to suggest here?

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 4:29am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/102 "2025-12-15T04:29:19Z")

</div>

> [@RainTown](#):
>
> But for me outstanding is really the larger scale setup - be it VMware or something else. How is IO provided to the VMs. In detail.

Some information I was able to get today states “_We don't use SRIOV on the PM1733 drives. We use virtio-block to expose the disk to the VM after creating a logical volume on the 3.84T PM1733 drive. The read and write IOPs limits for p4i instances (this is the type of instance we use) is 100 per GiB_.”

I’m not having much knowledge about how these operate internally and what questions to ask to the team, if you can please list down some questions pin-pointing the exact details required, I’ll be really grateful 😁.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 8:05am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/103 "2025-12-15T08:05:16Z")

</div>

@RainTown / @Christian_Dahlqvist got these inputs when asked my IT team regarding the hardware -

1. No your 60 VMs are not sharing a single PM1733 disk. 60 VMs are scheduled on different physical servers. Every physical server has 2 PM1733 disks which are shared among the VMs scheduled on that server.  
2. We don't use SRIOV for PM1733 disks. Rather we just create LVMs and expose them using virtio-blk to VMs  
3. Disk formatting is controlled by the VM owners. We just give raw block devices and the VM owners decide what filesystem to use.

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 8:31am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/104 "2025-12-15T08:31:18Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Your issue seems related to merging and indexingb, which coordinating-only nodes are not involved in, so I do not understand the question. Coordinating-only nodes im my experience primarily help with queries and aggregations where the can take some load off the data nodes.

Right, so my question is that we have a read aggregation query which runs at a certain QPS that it utilises 25% cpu of the data nodes (this was before we added co-ordinator nodes).

Now after adding the co-ordinator nodes, the co-ordinator nodes only experience a spike of 5%, whereas the load on the data nodes is the same…why are the co-ordinator nodes not helping in the aggregation over here?

> [@Christian\_Dahlqvist](#):
>
> If you change the shard partition size to an odd number, e.g. 9, do you get the same result?

Is it necessary for the number of primary shards to be a multiple of the routing\_partition number?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 15, 2025, 8:43am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/105 "2025-12-15T08:43:37Z")

</div>

> [@Lakshya\_Gupta](#):
>
> Now after adding the co-ordinator nodes, the co-ordinator nodes only experience a spike of 5%, whereas the load on the data nodes is the same…why are the co-ordinator nodes not helping in the aggregation over here?

It is helping with part of the aggregation work, but as far as I know most of the work is done at the shard level, which has to be executed on the data nodes.

> [@Lakshya\_Gupta](#):
>
> Is it necessary for the number of primary shards to be a multiple of the routing\_partition number?

I do not know. Even if you increase the number of partitions I do not see it resolving the issue. At best it probably just gives you a bit of added time.

> [@Lakshya\_Gupta](#):
>
> We don't use SRIOV for PM1733 disks. Rather we just create LVMs and expose them using virtio-blk to VMs

I have no experience setting this up but would suspectvthere is a level of configurability involved. I ran a quick serach onbline and found some guides around optimising e.g. the number of IO threads. If this is all correctly configured it sounds like you should be able to achieve near native performance corresponding to your IOPS quotas, but there is no way for us to know whether that is the case or not.

> [@Christian\_Dahlqvist](#):
>
> Do you know the document ID of a document that is heavily updated? If you have this I guess you could retrieve this from Elasticsearch using the [get document API](https://www.elastic.co/guide/en/elasticsearch/reference/8.19/docs-get.html#docs-get-api-response-body) once per minute and verify that the `_version` field increments at a nice slow pace in line with the expected update frequency.

Did you have a chance to verify that your deduplication logic is working as expected so we can remove that as a potential issue/contributing factor?

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 8:51am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/106 "2025-12-15T08:51:20Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Did you have a chance to verify that your deduplication logic is working as expected so we can remove that as a potential issue/contributing factor?

Not yet, have noted that down as one of the action item to complete by today (hopefully :P) 😃

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [December 15, 2025, 10:06am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/107 "2025-12-15T10:06:35Z")

</div>

Approaching 100 posts on this thread ...

@Lakshya_Gupta please try to share the `iostat -ztcxd 60 60` or `iostat -ztcxd 10 360` output and kibana screenshots, ideally covering same period that includes the 90%+ CPU load. What I want to see is do you ever get a sustained period with IOPS higher than a couple of K on the at-the-time loaded data nodes?

I probably looked at similar posts to @Christian_Dahlqvist about possible issues (gotchas even) with different virtualization setups in terms of IO. But via this forum that sort of low level stuff is going to be really difficult to pin down, just too many variables.

> [@Lakshya\_Gupta](#):
>
> Every physical server has 2 PM1733 disks which are shared among the VMs scheduled on that server.

If I've interpreted it right, thats kind of good news, the disks are in the physical servers running the VMs. But a critical word there is "shared". I'd like to know how "shared" these resources are, what you are sharing with, what else are you sharing, are you even having multiple elasticsearch data node VMs on the same physical server? e.g. maybe your increase of CPU count from 20 to 32 made little difference because the CPUs were already overcommitted.

Also:

> [@Lakshya\_Gupta](#):
>
> We don't use SRIOV for PM1733 disks. Rather we just create LVMs and expose them using virtio-blk to VMs

LVMs here might mean Linux Logical Volume Manager (lvm) devices, but can have other interpretations. But again it's just another level of abstraction, which makes seeing things clearly hard. And, without a doubt, SR-IOV would work better if setup correctly.

As a senior Elastic contributor wrote recently, effectively the TLDR here is:

> **Elastic node with Dedicated Host or VMs (with dedicated resources) = Best / Stable Outcome**

---

<div class="post-metadata">

**Author:** ![Lakshya\_Gupta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lakshya_gupta/32/146116_2.png) [@Lakshya\_Gupta](https://discuss.elastic.co/u/Lakshya_Gupta)\
**Post date:** [December 15, 2025, 11:13am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/108 "2025-12-15T11:13:21Z")

</div>

> [@RainTown](#):
>
> `iostat -ztcxd 60 60` or `iostat -ztcxd 10 360` output and kibana screenshots, ideally covering same period that includes the 90%+ CPU load.

Sure, this metric and the 1 suggested by Christian “Did you have a chance to verify that your deduplication logic is working as expected so we can remove that as a potential issue/contributing factor?”…I’ll try to get these both and post it here as soon as I’m able to see the load.

> [@RainTown](#):
>
> I'd like to know how "shared" these resources are, what you are sharing with, what else are you sharing, are you even having multiple elasticsearch data node VMs on the same physical server?

Sure, let me check this as well.

> [@RainTown](#):
>
> SR-IOV would work better if setup correctly

Sure, thanks for that input, checking on this as well with the infra team.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 15, 2025, 12:15pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954/109 "2025-12-15T12:15:21Z")

</div>

> [@Lakshya\_Gupta](#):
>
> @leandrojmp any tweaks you might want to suggest here?

No, I would mention to check the mount options because per default ext4 are mounted with the parameter relatime then it should be changed to noatime, this adds a boost in speed, but you already shared that the mount options include `noatime`.

[Previous page](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954.md?page=4)

[Next page](https://discuss.elastic.co/t/high-cpu-usage-on-a-few-data-nodes-hotspotting-of-data/383954.md?page=6)
