# Is there a way to throttle or stagger ILM?

**URL:** https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324
**Category:** Elasticsearch
**Tags:** ilm-index-lifecycle-management
**Created:** [November 2, 2023, 10:32pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324 "2023-11-02T22:32:55Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![jerrac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerrac/32/52980_2.png) [@jerrac](https://discuss.elastic.co/u/jerrac)
#### Post date: [November 2, 2023, 10:32pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/1 "2023-11-02T22:32:55Z")

</div>

So, a system I'm building that uses ES heavily has been periodically not getting the data that it should out of ES. After digging in a bit, I finally noticed that there was a pattern.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/e/cea96e596add10fbaa32457a899ff46f8146f009.png)  
The elastic agent queue depth keeps spiking periodically. A bit more digging and my logs showed me that those spikes are when ILM is rolling over indices and downsampling my data.

I'm guessing that a large part of the problem is that I'm running a single ES node. Long story short, we need to trim down as much as possible if we're going to keep using ES. So, increasing my ES nodes is not an ideal solution.

My thought would be to stagger ILM jobs somehow so it's doing a few at a time all day long instead of all of them all at once. Is there a way to do that?

My other (not ideal) thought would be to add extra processing nodes, while keeping only one master/data node, but would ILM even be able to run on a non-data node?

Any other ideas?

Thanks!

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [November 2, 2023, 11:11pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/2 "2023-11-02T23:11:50Z")

</div>

> [@jerrac](#):
>
> My thought would be to stagger ILM jobs somehow so it's doing a few at a time all day long instead of all of them all at once. Is there a way to do that?

Hmm ILM-triggered activities should be trying to stay out of the way of your production workload, it sounds like we might need a bit more throttling on the downsampling action. Yet it's only supposed to use a tiny threadpool, 1/8th of your CPUs, so I wonder why it's having such a big impact.

Could you grab `GET _nodes/hot_threads?threads=9999` from a time when it's struggling, and share it here (or likely on [https://gist.github.com/](https://gist.github.com/) since it'll be too big)?

---

<div class="post-metadata">

### Author: ![jerrac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerrac/32/52980_2.png) [@jerrac](https://discuss.elastic.co/u/jerrac)
#### Post date: [November 6, 2023, 10:06pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/3 "2023-11-06T22:06:00Z")

</div>

@DavidTurner Here are a few different runs of that command.

> <https://gist.github.com/jerrac/20f71080b0ce75c38338db384b1655db>

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [November 7, 2023, 1:58pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/4 "2023-11-07T13:58:11Z")

</div>

Thanks, that's helpful. Are you running on spinning disks or SSDs?

---

<div class="post-metadata">

### Author: ![jerrac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerrac/32/52980_2.png) [@jerrac](https://discuss.elastic.co/u/jerrac)
#### Post date: [November 7, 2023, 2:22pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/5 "2023-11-07T14:22:08Z")

</div>

Data is on an iscsi lun backed by SSD's. I believe they are pretty fast SSD's as well. You ever hear of an Kaminario? That's what the storage is on.

Edit:

Also, possibly relevant, ES is running in a single node Docker stack service. We have 3 Docker Swarm nodes it can run on, so each of those nodes mounts the lun, and we have OCFS configured for the filesystem. The idea being if the 1 instance of ES has to be restarted on another node, it will be using the same data as the old instance.

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [November 7, 2023, 2:51pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/6 "2023-11-07T14:51:10Z")

</div>

> [@jerrac](#):
>
> Data is on an iscsi lun backed by SSD's. I believe they are pretty fast SSD's as well.

Hmm. These stack dumps show that your system is _heavily_ bottlenecked on IO, with many threads stuck for several hundreds of milliseconds waiting for a `write()` or similar to complete. I don't think your storage is performing as well as you think it should.

---

<div class="post-metadata">

### Author: ![jerrac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerrac/32/52980_2.png) [@jerrac](https://discuss.elastic.co/u/jerrac)
#### Post date: [November 18, 2023, 2:05am UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/7 "2023-11-18T02:05:38Z")

</div>

Well, I am 90% sure the issue is OCFS2. I moved the ES instance to a different server where I could use a normal xfs iscsi lun, and that seems to have resolved the issues I was having.

Thanks for the help @DavidTurner !

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [November 18, 2023, 7:11am UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/8 "2023-11-18T07:11:54Z")

</div>

Ah yes that'd explain it indeed, thanks for closing the loop. Clustered filesystems seem to be a rich source of performance (and sometimes correctness) issues, and the complexity they add is largely unnecessary when Elasticsearch is also doing its own clustering and replication work. XFS is a better choice IMO.

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [November 18, 2023, 7:13am UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/9 "2023-11-18T07:13:26Z")

</div>

Closing another loop on the ES side, we still think it might be a good idea to limit the resources needed by downsampling anyway:

> <https://github.com/elastic/elasticsearch/issues/101970>
>
> \### Description
> 
> The downsampling task is a single thread task when it comes t…o metric aggregations. Anyway, indexing documents into the target index happens using the \`BulkProcessor2\`. The downsampling thread submits indexing requests without waiting for a response so to achieve maximum throughput. As a result of that, normally, there are multiple outstanding indexing requests consuming threads from the search/indexing thread pool. That can result in using all available threads for downsampling (indexing) without leaving room for other tasks, like regular indexing, to be executed. Ideally we would like to implement a mechanism by which we limit the number of outstanding indexing requests so to limit the number of threads used for indexing by the downsampling thread. Also we would like to expose this limit as a setting that users can control. A possibility would be to expose the \`maxBytesInFlight\` of \`BulkProcessor2\` as a setting (instead of setting it as a constant as it is right now).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [December 16, 2023, 7:14am UTC](https://discuss.elastic.co/t/is-there-a-way-to-throttle-or-stagger-ilm/346324/10 "2023-12-16T07:14:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
