# S3 bucket size holding the snapshots is 2-2.5x of the total disk used of the cluster

**URL:** <https://discuss.elastic.co/t/s3-bucket-size-holding-the-snapshots-is-2-2-5x-of-the-total-disk-used-of-the-cluster/278944>\
**Category:** Elasticsearch\
**Tags:** snapshot-and-restore\
**Created:** [July 16, 2021, 8:09pm UTC](https://discuss.elastic.co/t/s3-bucket-size-holding-the-snapshots-is-2-2-5x-of-the-total-disk-used-of-the-cluster/278944 "2021-07-16T20:09:40Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![MitParekh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mitparekh/32/91768_2.png) [@MitParekh](https://discuss.elastic.co/u/MitParekh)\
**Post date:** [July 16, 2021, 8:09pm UTC](https://discuss.elastic.co/t/s3-bucket-size-holding-the-snapshots-is-2-2-5x-of-the-total-disk-used-of-the-cluster/278944/1 "2021-07-16T20:09:41Z")

</div>

Hi,

I have a cluster (ES 7.11.1) with 6 nodes and the **total disk used** is roughly **~4.5TB**. I store logging data in it using data streams. My data stream patterns are like `abc-lmn-xyz`. The data retention period is 14 days.

I have an SLM policy that runs every hour where I configured it to pick up indices matching the pattern `*abc*`. The expiry set in s3 for snapshots is 15 days.

```json
{
  "name": "<abc-hourly-{now{yyyy-MM-dd't'HH:mm:ss.SSS'z'}}>",
  "schedule": "0 0 * * * ?",
  "repository": "s3_repository_ds",
  "config": {
    "indices": [
      "*abc*"
    ],
    "ignore_unavailable": true
  },
  "retention": {
    "expire_after": "15d"
  }
}

```

The **total size of data in s3** was found to be **~11TB** even though my total disk size used is ~4.5TB

Is this s3 storage size expected, or is my policy misconfigured?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 17, 2021, 7:28am UTC](https://discuss.elastic.co/t/s3-bucket-size-holding-the-snapshots-is-2-2-5x-of-the-total-disk-used-of-the-cluster/278944/2 "2021-07-17T07:28:22Z")

</div>

Doesn't seem totally unreasonable to me. Data in snapshots is deduplicated where possible, but if you take a snapshot, do some more indexing and take another snapshot then it's possible that all the files in the shard are different (no deduplication is possible) which would take up double the storage size.

Do you really need to retain every hourly snapshot for the full 15 days? You could retain hourlies for 2 days say and then just 12-hourly ones for the remainder of the time.

---

<div class="post-metadata">

**Author:** ![MitParekh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mitparekh/32/91768_2.png) [@MitParekh](https://discuss.elastic.co/u/MitParekh)\
**Post date:** [July 17, 2021, 7:29pm UTC](https://discuss.elastic.co/t/s3-bucket-size-holding-the-snapshots-is-2-2-5x-of-the-total-disk-used-of-the-cluster/278944/3 "2021-07-17T19:29:52Z")

</div>

Sorry, I am a bit confused here. When you say store hourly for 2 days and 12 hourlies for 15 days. What value does it add? Won't the size of both the snapshots be the same? Am I missing out something here?

I understand that the snapshots are incremental so be it hourly or 12 hourly, the total storage is going to be the same.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 17, 2021, 8:01pm UTC](https://discuss.elastic.co/t/s3-bucket-size-holding-the-snapshots-is-2-2-5x-of-the-total-disk-used-of-the-cluster/278944/4 "2021-07-17T20:01:49Z")

</div>

Snapshots are not incremental, at least not at the document level. Each snapshot contains the full set of data but segments are reused if the have add ready been copied and are unchanged. The API keeps track of which segments are in use by which segments are in use by which snapshots. When merging occurs data is copied over to a new segment which means the repository can hold the same data multiple times which explains the larger size.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 14, 2021, 8:01pm UTC](https://discuss.elastic.co/t/s3-bucket-size-holding-the-snapshots-is-2-2-5x-of-the-total-disk-used-of-the-cluster/278944/5 "2021-08-14T20:01:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
