# Feasible solution of snapshot for 30 GB and increasing every day data

**URL:** <https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-monitoring\
**Created:** [February 3, 2020, 7:05am UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568 "2020-02-03T07:05:23Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![priyanka10](https://avatars.discourse-cdn.com/v4/letter/p/43a26b/32.png) [@priyanka10](https://discuss.elastic.co/u/priyanka10)\
**Post date:** [February 3, 2020, 7:05am UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/1 "2020-02-03T07:05:23Z")

</div>

I want to know that how can I take differential snapshots in ElasticSearch and also how it works?

We received around 30 GB of data monthly in all indices of ElasticSearch. Few indices get update daily and few indices data get purge after certain retention days. So I was thinking to go with incremental snapshot so that it will not take time and only modified data will get into a snapshot. But I don't know how it works and will it be feasible for my case?

Could you please help me to design a snapshot process so that it can work permanently and will not be impacted with time.

---

<div class="post-metadata">

**Author:** ![Armin\_Braun](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/armin_braun/32/20092_2.png) [@Armin\_Braun](https://discuss.elastic.co/u/Armin_Braun)\
**Post date:** [February 3, 2020, 9:53am UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/2 "2020-02-03T09:53:45Z")

</div>

Hi @priyanka10

all snapshots to a single repository are incremental. You will not have to do anything specifically to take differential snapshots. If you just take snapshots to the same repository, each snapshot will try to reuse as much data as possible from prior snapshots automatically.

---

<div class="post-metadata">

**Author:** ![nsuthar93](https://avatars.discourse-cdn.com/v4/letter/n/8dc957/32.png) [@nsuthar93](https://discuss.elastic.co/u/nsuthar93)\
**Post date:** [February 5, 2020, 1:15pm UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/3 "2020-02-05T13:15:47Z")

</div>

Thanks Armin for reply, could you please help us to understand your statement:

"_If you just take snapshots to the same repository, each snapshot will try to reuse as much data as possible from prior snapshots automatically_."

What we are understanding, suppose there is one index which has initially three records as below.

```
    {
            "_index": "test_inx",
            "_type": "doc",
            "_id": "2",
            "_score": 1,
            "_source": {
              "Empid": "2",
              "Name": "BCD"
            }
          },
          {
            "_index": "test_inx",
            "_type": "doc",
            "_id": "1",
            "_score": 1,
            "_source": {
              "Empid": "1",
              "Name": "ABC"
            }
          },
          {
            "_index": "test_inx",
            "_type": "doc",
            "_id": "3",
            "_score": 1,
            "_source": {
              "Empid": "3",
              "Name": "EFG"
            }

```

Now, we take a snapshot of above index into **snapshot-1**.  
On next day, few new records get insert (Records with EmpId with 4 & 5) into index and one record get update (record with EmpId 2) so final stage of index will be as below.

```
{
        "_index": "priority_inx",
        "_type": "doc",
        "_id": "2",
        "_score": 1,
        "_source": {
          "Empid": "2",
          "Name": "NEW_BCD"
        }
      },
      {
        "_index": "priority_inx",
        "_type": "doc",
        "_id": "1",
        "_score": 1,
        "_source": {
          "Empid": "1",
          "Name": "ABC"
        }
      },
      {
        "_index": "priority_inx",
        "_type": "doc",
        "_id": "3",
        "_score": 1,
        "_source": {
          "Empid": "3",
          "Name": "EFG"
        }
     {
        "_index": "priority_inx",
        "_type": "doc",
        "_id": "4",
        "_score": 1,
        "_source": {
          "Empid": "4",
          "Name": "LMN"
        }
     {
        "_index": "priority_inx",
        "_type": "doc",
        "_id": "5",
        "_score": 1,
        "_source": {
          "Empid": "5",
          "Name": "XYZ"
        }

```

Next, we take new snapshot into **snapshot-2** within same repository.

Now, my concern is what happen in background, **since records with EmpId 1 and 3 are as it is in both snapshots so are they again restore into snapshot-2?** or as per your statement "_it use as much data from previous snapshot_" then will **it take these two records from snapshot-1 when we will try to restore snapshot-2?**

**if it stored all records into snapshot-2 then is it possible to delete snapshot-1 since with time repository size will be increase if it keep same records multiple times?**

---

<div class="post-metadata">

**Author:** ![Armin\_Braun](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/armin_braun/32/20092_2.png) [@Armin\_Braun](https://discuss.elastic.co/u/Armin_Braun)\
**Post date:** [February 5, 2020, 1:32pm UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/4 "2020-02-05T13:32:13Z")

</div>

First off , the unit of snapshotting isn't individual documents but rather Lucene segments, which can be roughly interpreted as a group of documents. So the level of incrementally isn't that granular. Still, the points your question raises remain the same logically, just figured I'd point this out:

> [@nsuthar93](#):
>
> Now, my concern is what happen in background, **since records with EmpId 1 and 3 are as it is in both snapshots so are they again restore into snapshot-2?** or as per your statement " _it use as much data from previous snapshot_ " then will **it take these two records from snapshot-1 when we will try to restore snapshot-2?**

Yes roughly (if there are Lucene merges in the meantime or primary failover this may not always work) this is true, unchanged segments/documents will be reused in cases like this one and the segment containing document `1` and `3` will not be re-uploaded to the repository.

> [@nsuthar93](#):
>
> if it stored all records into snapshot-2 then is it possible to delete snapshot-1 since with time repository size will be increase if it keep same records multiple times?

Yes, you can delete snapshots as you see fit. The snapshot functionality will then simply remove the data that was only referenced by the deleted snapshot (in your example snapshot-1) but will leave the data that is still required for other snapshots in the repository. The repository will not needlessly keep files around that aren't used by any snapshots.

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [February 5, 2020, 2:29pm UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/5 "2020-02-05T14:29:05Z")

</div>

if you want to automate this from cron level check this out as well

[https://discuss.elastic.co/t/automatic-elasticsearch-snapshot-backup/173223](https://discuss.elastic.co/t/automatic-elasticsearch-snapshot-backup/173223)

---

<div class="post-metadata">

**Author:** ![nsuthar93](https://avatars.discourse-cdn.com/v4/letter/n/8dc957/32.png) [@nsuthar93](https://discuss.elastic.co/u/nsuthar93)\
**Post date:** [February 6, 2020, 6:25am UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/6 "2020-02-06T06:25:28Z")

</div>

Thank you Armin, Now it is clear 🙂

---

<div class="post-metadata">

**Author:** ![nsuthar93](https://avatars.discourse-cdn.com/v4/letter/n/8dc957/32.png) [@nsuthar93](https://discuss.elastic.co/u/nsuthar93)\
**Post date:** [February 6, 2020, 6:26am UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/7 "2020-02-06T06:26:15Z")

</div>

sure we will check it and let you know in case any concerns.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 5, 2020, 6:26am UTC](https://discuss.elastic.co/t/feasible-solution-of-snapshot-for-30-gb-and-increasing-every-day-data/217568/8 "2020-03-05T06:26:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
