# Query on Incremental backups using snapshots

**URL:** <https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597>\
**Category:** Elasticsearch\
**Created:** [April 24, 2020, 8:16am UTC](https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597 "2020-04-24T08:16:37Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![shivani\_aggarwal](https://avatars.discourse-cdn.com/v4/letter/s/22d042/32.png) [@shivani\_aggarwal](https://discuss.elastic.co/u/shivani_aggarwal)\
**Post date:** [April 24, 2020, 8:16am UTC](https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597/1 "2020-04-24T08:16:37Z")

</div>

Hi,  
We use **elasticsearch snapshot API** to take backup of ES data daily and also regularly copy the entire content from the snapshot repository to another remote system.

Now that we have huge amount of data in ES (in TBs), and the daily increment to actual data is very less (order of few GBs), we do not think it is feasible to transfer TBs of backed-up data to the remote system everyday.

As ES snapshots are incremental: each snapshot of an index only stores data that is not part of an earlier snapshot --

1. Is it possible to clearly identify only the incremental changes to the snapshot repo as part of creation of the snapshot?  
This would help us to only transfer the incremental snapshot content to the remote system and not the entire thing.  
If yes, can you please help us understand how that can be achieved.

2. If point 1 can be achieved, how do you think restore would work? Can the incremental snapshots be combined and placed into the snapshot repo for restore to ES?

Note: Version used - **elasticsearch-oss:7.0.1**  
Any pointers/suggestions would be appreciated.

Thanks!

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 24, 2020, 8:48am UTC](https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597/2 "2020-04-24T08:48:43Z")

</div>

I think you are looking for [https://github.com/elastic/elasticsearch/issues/54944](https://github.com/elastic/elasticsearch/issues/54944) so please indicate your support for this feature on that Github Issue.

Today the best you can do is make sure that no snapshots are currently running and then transfer any new/changed files to the remote system. A tool like `rsync` would do this, but be warned that if it fails part-way through a transfer then it may leave the remote repository in an inconsistent state. It's probably simpler and more reliable to snapshot directly to the remote system using a shared filesystem (or, e.g something like Minio). Or just use a public cloud for your snapshots, they're very reliable and secure and not that expensive given the hassle you're currently facing. It only costs about $25 per month to store 1TB of data on S3.

To restore, you would need to make the entire repository available to Elasticsearch, either by exposing it on a shared filesystem (or e.g. something like Minio) or else by copying all the files to somewhere that Elasticsearch can access.

---

<div class="post-metadata">

**Author:** ![shivani\_aggarwal](https://avatars.discourse-cdn.com/v4/letter/s/22d042/32.png) [@shivani\_aggarwal](https://discuss.elastic.co/u/shivani_aggarwal)\
**Post date:** [April 24, 2020, 12:32pm UTC](https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597/3 "2020-04-24T12:32:38Z")

</div>

Thanks for the reply.  
Sorry if i was not clear with my question.

We already have a process which would take the backup of ES data. The sequence of backup is as below, when backup is triggered (ideally a cron which runs daily)

1. Snapshot API runs and stores the snapshot to a local volume say es\_backup (glusterfs)
2. Another process runs which does tar of the content of es\_backup and sends this to remote repo.

The sequence followed during restore:

1. Clear the es\_backup volume.
2. Copy the content from remote repo and untar to es\_backup.
3. Run snapshot API to restore from es\_backup.

Issue:

1. ES is used to store huge data around 5 TB
2. Amount of data that is pushed everyday is in few GBs.
3. Even though snapshot API does incremental backup, but the backup which is moved to remote repo is the complete content of es\_backup (Not incremental)

Expectation:

1. Just like how ES snapshot api takes incremental backup, is it possible to move **ONLY** this incremental backup to the remote repo? How do we identify only the incremented changes?

2. If this is possible, if in case of disaster and we want to restore last 7 days data, copying back the previous 7 days incrementally backed-up data from remote repo to es\_backup and running the restore api will restore the data in ES ?

I hope I was able to explain my question better now 🙂 Pls let me know if more details are required.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 24, 2020, 1:37pm UTC](https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597/4 "2020-04-24T13:37:31Z")

</div>

I think your original question was quite clear, and my answer to your clarification is the same.

> [@shivani\_aggarwal](#):
>
> 1. How do we identify only the incremented changes?

Compare the files (e.g. by their last-modified date) for instance using a tool like `rsync`, but be warned that this might be unreliable. Better to snapshot directly to the remote system.

> [@shivani\_aggarwal](#):
>
> 1. If this is possible, if in case of disaster and we want to restore last 7 days data, copying back the previous 7 days incrementally backed-up data from remote repo to es\_backup and running the restore api will restore the data in ES ?

No, Elasticsearch needs access to the whole repository, not just the last 7 days of changes.

---

<div class="post-metadata">

**Author:** ![shivani\_aggarwal](https://avatars.discourse-cdn.com/v4/letter/s/22d042/32.png) [@shivani\_aggarwal](https://discuss.elastic.co/u/shivani_aggarwal)\
**Post date:** [April 27, 2020, 5:56pm UTC](https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597/5 "2020-04-27T17:56:49Z")

</div>

Many thanks for your explanation 🙂

About the restore part, just to add on -

- Everyday a new index gets created (in the format log-\<yyyy.mm.dd\> with new data ingested to it.
- We have configured the retention of indices as 20d (using elasticsearch-curator that runs everyday).  
So that means, everyday all indices older than 20d get deleted.  
Then the elasticsearch snapshot API is used to backup the current data in ES incrementally.

Now, can only the last 20 days incrementally backed-up data be restored ?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 25, 2020, 5:57pm UTC](https://discuss.elastic.co/t/query-on-incremental-backups-using-snapshots/229597/6 "2020-05-25T17:57:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
