# Archive old indices with data compression

**URL:** <https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879>\
**Category:** Elasticsearch\
**Created:** [October 16, 2019, 3:49pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879 "2019-10-16T15:49:40Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![kevinray0030](https://avatars.discourse-cdn.com/v4/letter/k/71c47a/32.png) [@kevinray0030](https://discuss.elastic.co/u/kevinray0030)\
**Post date:** [October 16, 2019, 3:49pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879/1 "2019-10-16T15:49:40Z")

</div>

Hey all,

I am trying to find a solution to where I can keep roughly 90 days of live data on my cluster but then archive anything over 90 days up to a year. This is a compliance requirement.

I am currently looking into Curator as an automated solution, but the Snapshot function only does metadata and I am looking for something that will also compress my data, is there a function of Elasticsearch that I am missing?

Thanks

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [October 16, 2019, 4:43pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879/2 "2019-10-16T16:43:43Z")

</div>

> [@kevinray0030](#):
>
> the Snapshot function only does metadata

Hmm? What do you mean?

> [@kevinray0030](#):
>
> I am looking for something that will also compress my data, is there a function of Elasticsearch that I am missing?

Snapshots are full copies of all of the segments within the indices/shards you tell the API you want to have in the snapshot. Compression of data occurs at the segment-level, so technically, your data is already compressed, but in the exact same format it exists in on your nodes.

If you want to compress data further, you would have to use other means, like a compressed filesystem (ZFS, perhaps, though it won't further compress the LZ4 or DEFLATE segments much further), or exporting and compressing the raw JSON data.

---

<div class="post-metadata">

**Author:** ![kevinray0030](https://avatars.discourse-cdn.com/v4/letter/k/71c47a/32.png) [@kevinray0030](https://discuss.elastic.co/u/kevinray0030)\
**Post date:** [October 16, 2019, 5:06pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879/3 "2019-10-16T17:06:51Z")

</div>

> [@theuntergeek](#):
>
> > [@kevinray0030](#):
> >
> > the Snapshot function only does metadata
> 
> Hmm? What do you mean?

Sorry, I realize this is confusing now looking back at it, when looking at the Elasticsearch documentation in regards to Snapshot, the repository is created with compression but the compression only applies to metadata files (index mapping and settings).

> [@theuntergeek](#):
>
> If you want to compress data further, you would have to use other means, like a compressed filesystem (ZFS, perhaps, though it won't further compress the LZ4 or DEFLATE segments much further), or exporting and compressing the raw JSON data.

So, the compressing of data further is what I am looking for. I essentially have an index, for example, that is 10GB of data. After we hit our 90 day period, we will likely not need to look at that data again, ever, but due to compliance reasons, we need to keep it in case someone does want to look at it.

Would it be possibly for me to just use something like gzip on those files?

Another question that just occurred to me is that when creating the snapshots, is there a way to have the index folder that is created be renamed to whatever it's alias is?

For example, if I have a folder with the name of **leCqaJuESyCYg\_VL7Xl3og** but the english version of the index is **index\_123** , is there a way, during the snapshot process to name that folder as the easier to read version?

Thanks.

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [October 16, 2019, 5:31pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879/4 "2019-10-16T17:31:28Z")

</div>

> [@kevinray0030](#):
>
> Would it be possibly for me to just use something like gzip on those files?

No. This does not work.

As stated, a file like **leCqaJuESyCYg\_VL7Xl3og** corresponds to an index name, and the files in that directory are index data/segments. This data is already compressed using at least LZ4, and potentially DEFLATE if you enabled `best_compression` (and further compressing it may not even yield the benefits you are hoping for, while incurring considerable extra cost/effort). You cannot rename or alter these names in any way or it will invalidate your snapshot. The cluster metadata requires that these remain completely consistent. As such, you cannot just gzip the directory.

Some people have taken to putting a full snapshot into a dedicated S3 repository/bucket path, and then putting that full path into Glacier for long-term storage, which is less expensive and can persist for a long time. That may be an option for you, as you can have multiple snapshot repositories, and simply choose which one to snapshot into at snapshot time.

The same is true for a local/NFS snapshot. You could create a clean, new snapshot repository in a dedicated filesystem path, and then tar/gzip that path. You won't get much better compression, but it would be something. You would then not want to re-use that path for any other snapshots, but probably even remove the repository altogether.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 16, 2019, 5:36pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879/5 "2019-10-16T17:36:42Z")

</div>

If you are on a reasonably recent version you may want to look into [source only snapshots](https://www.elastic.co/guide/en/elasticsearch/reference/7.4/modules-snapshots.html#_source_only_repository) which can save space but will take longer to restore into the cluster if needed.

---

<div class="post-metadata">

**Author:** ![kevinray0030](https://avatars.discourse-cdn.com/v4/letter/k/71c47a/32.png) [@kevinray0030](https://discuss.elastic.co/u/kevinray0030)\
**Post date:** [October 16, 2019, 5:51pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879/6 "2019-10-16T17:51:33Z")

</div>

I am running ES 7.1, I will take a look into those source only snapshots.  
Thanks for the recommendation.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 13, 2019, 5:57pm UTC](https://discuss.elastic.co/t/archive-old-indices-with-data-compression/203879/7 "2019-11-13T17:57:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
