# Using ILM for huge size of indexes

**URL:** <https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496>\
**Category:** Elasticsearch\
**Tags:** ilm-index-lifecycle-management\
**Created:** [February 25, 2023, 9:38am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496 "2023-02-25T09:38:32Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 25, 2023, 9:38am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/1 "2023-02-25T09:38:32Z")

</div>

Use case: We have a few indexes which have huge amounts of data and they are growing. We need to figure out a way to optimize the search time. We are thinking of using ILM to manage the indexes. But there are a few roadblocks we are facing. Hoping that we can get help here

1. Is there a way we can be updated when an index is created through ILM? Maybe through a webhook. We need this to keep a list of indexes and when they are created so whenever we need to update an existing document, we can directly update that document in that particular index rather than searching all the indexes. If not, let us know what the best way to achieve this.
2. What's the best way to optimize the search for a time filter? We don't want to query all the indexes unnecessarily. We are thinking of keeping track of indexes along with created time and filtering on our end to search on the particular indexes which fall under the time filter.  
We are using Elasticsearch 6.8 currently.  
Also, if things are possible in the latest version, do let us know.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 25, 2023, 9:38am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/2 "2023-02-25T09:38:32Z")

</div>

Elasticsearch 6.8 is [EOL](https://www.elastic.co/support/eol) and no longer supported. Please upgrade ASAP.

(This is an automated response from your friendly Elastic bot. Please report this post if you have any suggestions or concerns :elasticheart: )

---

<div class="post-metadata">

**Author:** ![Wave](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wave/32/117242_2.png) [@Wave](https://discuss.elastic.co/u/Wave)\
**Post date:** [February 25, 2023, 8:27pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/3 "2023-02-25T20:27:28Z")

</div>

Can you provide us with some specific numbers? How many indices and how big is each index and how many shards per index? There are recommended sizes of shards on average and I want to make sure we are in the same ballpark. As far as number 2 goes, elastic does a great job to know what indices contain what data and you usually don't have to worry about tailoring your search query at the index level. Having shards that are too large (in the hundreds of gigabytes is where you usually run into trouble). With all that said it really depends on the data and cluster, so with out specifics it is hard to be give recommendations. Also, like the bot said you should upgrade the cluster at your earlier convenance. There have been lots of upgrades from 6.8 to 8.6 that you will probably make use of.

---

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 26, 2023, 5:29am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/4 "2023-02-26T05:29:47Z")

</div>

We are having 2 indexes above 100 GB and 6 indexes in the range of 50GB to 100 GB. Each of them has 2 shards. How does Elastic figure out which indices have what data?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 26, 2023, 10:27am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/5 "2023-02-26T10:27:56Z")

</div>

Elasticsearch uses a routing value on which we compute a modulo to define in which shard a document should go.  
By default the routing value is the id of the document.

---

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 26, 2023, 12:22pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/6 "2023-02-26T12:22:16Z")

</div>

@dadoonet This should be for creating documents, correct?  
I am looking to optimize my searching as I have a time-based query and was thinking to refine the search by specifying indexes on my end. I wanted to know does Elastic uses any method to optimise which indexes/shard to look for while searching. If yes, then how does it do that.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 26, 2023, 3:04pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/7 "2023-02-26T15:04:03Z")

</div>

> [@Chirag\_Poddar](#):
>
> I wanted to know does Elastic uses any method to optimise which indexes/shard to look for while searching. If yes, then how does it do that.

Elasticsearch has had many improvements in search, performance, resource usage and many other aspects in last couple of years, it would be very hard to explain what changed, but you may read the release notes for every version since 6.8 if you want to know what was changed and added.

But in your example, you do not need to specify the exact index that may have your data while search, elasticsearch is able to already do that and from version 7.16 this got even better.

Take a look at the this [blog post](https://www.elastic.co/blog/three-ways-improved-elasticsearch-scalability) which has an example of some improvements made.

You should look into upgrade as soon as possible, first you will need to upgrade to the last 7.17 and then you would be able to upgrade to the last 8 version.

---

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 26, 2023, 3:35pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/8 "2023-02-26T15:35:07Z")

</div>

@leandrojmp Okay. Thank you  
Is there a way to know when ILM creates an index rather than querying the alias details at an interval? We need this to update the old documents.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 26, 2023, 4:03pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/9 "2023-02-26T16:03:34Z")

</div>

I do not see the point with ILM in this scenario. It was designed to handle immutable data, which is not what you have.

If you have a timestamp associated with every document that you can access when updating I would use "traditional" time based indices where the time period covered, e.g. a day or a month, is specified as part of the index name.

---

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 27, 2023, 10:29am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/10 "2023-02-27T10:29:00Z")

</div>

Ok. So should we use Rollover API for creating such indexes? And we should use the created time of that doc to find the index for updation.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 27, 2023, 10:33am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/11 "2023-02-27T10:33:21Z")

</div>

If you use rollover you have the same problem as with ILM (uses rollover behind the scenes) that you do not know which index data resides in. Instead create monthly indices with the year and month in the name, e.g. `index-2023.02`. If you have the timestamp of a document elasewhere you know exactly which index to update based on this.

---

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 27, 2023, 10:39am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/12 "2023-02-27T10:39:07Z")

</div>

For that exact reason, I am trying to figure out a way in which I can keep a mapping on my own end of index name and creation time while using ILM. Is there a way to find this mapping apart from calling ILM/Rollover explain API at a fixed interval?

I am trying to avoid index creation on my own end as this will add an overhead on the system when I am trying to create a document.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 27, 2023, 10:45am UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/13 "2023-02-27T10:45:16Z")

</div>

I do not understand what the issue is. The way I described is how time-based indices were managed for years before rollover came into the picture. You would have an index template that applied to all new indices related to a specific pattern, and you would create the index name based on the timestamp in the document when writing it. If you now what the original timestamp is for documents you want to update (kept track of outside Elasticsearch) that is all you need to create the correct index name.

---

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 27, 2023, 1:05pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/14 "2023-02-27T13:05:42Z")

</div>

Okay.  
Is it a bad idea to track the indexex created through rollover and maintain the mapping of the index name and creation time of the index?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 27, 2023, 5:45pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/15 "2023-02-27T17:45:45Z")

</div>

What would be the benefit of this over the approach I suggested?

I think you are overcomplicating things, but it could be that I have misunderstood something.

---

<div class="post-metadata">

**Author:** ![Chirag\_Poddar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chirag_poddar/32/117722_2.png) [@Chirag\_Poddar](https://discuss.elastic.co/u/Chirag_Poddar)\
**Post date:** [February 27, 2023, 7:41pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/16 "2023-02-27T19:41:37Z")

</div>

I am trying to remove the overhead of creating indexes on the application side. Nothing else.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 27, 2023, 8:11pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/17 "2023-02-27T20:11:22Z")

</div>

If you have an index template that matches the set of time-based indices I am suggesting, you just derive the correct index name for each document based on the stored timestamp and send a bulk request to Elasticsearch. When Elasticsearch sees you are indexing into a new index, it will automatically create it. You should therefore not explicitly need to create any indices from the application.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 27, 2023, 8:11pm UTC](https://discuss.elastic.co/t/using-ilm-for-huge-size-of-indexes/326496/18 "2023-03-27T20:11:36Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
