# Backup data in a robust way

**URL:** <https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214>\
**Category:** Elasticsearch\
**Created:** [January 2, 2013, 9:35am UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214 "2013-01-02T09:35:35Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Robin\_Verlangen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robin_verlangen/32/1542_2.png) [@Robin\_Verlangen](https://discuss.elastic.co/u/Robin_Verlangen)\
**Post date:** [January 2, 2013, 9:35am UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/1 "2013-01-02T09:35:35Z")

</div>

Hi there,

What would be the best way to create backups of the data in ES? In my use  
case the data is append-only (log storage etc).

I was thinking about closing the index, gzip the entire folder on every  
node (not so easy to coordinate), and store it on something like Amazon S3.  
However I would like to know whether there are more solid solutions to this.

Best regards,

Robin Verlangen  
_Software engineer_  
\*  
\*  
W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
E robin@us2.nl

[http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)

Disclaimer: The information contained in this message and attachments is  
intended solely for the attention and use of the named addressee and may be  
confidential. If you are not the intended recipient, you are reminded that  
the information remains the property of the sender. You must not use,  
disclose, distribute, copy, print or rely on this e-mail. If you have  
received this message in error, please contact the sender immediately and  
irrevocably delete this message and any copies.

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 2, 2013, 10:05am UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/2 "2013-01-02T10:05:13Z")

</div>

Hey Robin,

I found this Gist: [Backup and restore an Elastic search index (shamelessly copied from http://tech.superhappykittymeow.com/?p=296) · GitHub](https://gist.github.com/1939828)  
[https://gist.github.com/1939828](https://gist.github.com/1939828)

I never tested it myself.

Have a look also at: [Backup Elasticsearch node · GitHub](https://gist.github.com/4380715)  
[https://gist.github.com/4380715](https://gist.github.com/4380715)  
Never tested it also...

That said, I think that in the next versions, some tools will appears and  
probably will cover this need.

HTH  
David

Le 2 janvier 2013 à 10:35, Robin Verlangen [robin@us2.nl](mailto:robin@us2.nl) a écrit :

> Hi there,
> 
> What would be the best way to create backups of the data in ES? In my use  
> case the data is append-only (log storage etc).
> 
> I was thinking about closing the index, gzip the entire folder on every node  
> (not so easy to coordinate), and store it on something like Amazon S3. However  
> I would like to know whether there are more solid solutions to this.
> 
> Best regards,
> 
> Robin Verlangen  
> Software engineer
> 
> W [http://www.robinverlangen.nl](http://www.robinverlangen.nl) [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> E robin@us2.nl [mailto:robin@us2.nl](mailto:robin@us2.nl)  
> [http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)
> 
> Disclaimer: The information contained in this message and attachments is  
> intended solely for the attention and use of the named addressee and may be  
> confidential. If you are not the intended recipient, you are reminded that the  
> information remains the property of the sender. You must not use, disclose,  
> distribute, copy, print or rely on this e-mail. If you have received this  
> message in error, please contact the sender immediately and irrevocably delete  
> this message and any copies.
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![Robin\_Verlangen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robin_verlangen/32/1542_2.png) [@Robin\_Verlangen](https://discuss.elastic.co/u/Robin_Verlangen)\
**Post date:** [January 2, 2013, 1:52pm UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/3 "2013-01-02T13:52:40Z")

</div>

Hi David,

Thank you for your reply. The second gist is about the same as I was  
thinking about. Seems that I'll have to keep track of my data, or find it  
through the API.

Best regards,

Robin Verlangen  
_Software engineer_  
\*  
\*  
W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
E robin@us2.nl

[http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)

Disclaimer: The information contained in this message and attachments is  
intended solely for the attention and use of the named addressee and may be  
confidential. If you are not the intended recipient, you are reminded that  
the information remains the property of the sender. You must not use,  
disclose, distribute, copy, print or rely on this e-mail. If you have  
received this message in error, please contact the sender immediately and  
irrevocably delete this message and any copies.

On Wed, Jan 2, 2013 at 11:05 AM, David Pilato [david@pilato.fr](mailto:david@pilato.fr) wrote:

> \*\*  
> Hey Robin,
> 
> I found this Gist: [Backup and restore an Elastic search index (shamelessly copied from http://tech.superhappykittymeow.com/?p=296) · GitHub](https://gist.github.com/1939828)
> 
> I never tested it myself.
> 
> Have a look also at: [Backup Elasticsearch node · GitHub](https://gist.github.com/4380715)  
> Never tested it also...
> 
> That said, I think that in the next versions, some tools will appears and  
> probably will cover this need.
> 
> HTH  
> David
> 
> Le 2 janvier 2013 à 10:35, Robin Verlangen [robin@us2.nl](mailto:robin@us2.nl) a écrit :
> 
> Hi there,
> 
> What would be the best way to create backups of the data in ES? In my use  
> case the data is append-only (log storage etc).
> 
> I was thinking about closing the index, gzip the entire folder on every  
> node (not so easy to coordinate), and store it on something like Amazon S3.  
> However I would like to know whether there are more solid solutions to  
> this.
> 
> Best regards,
> 
> Robin Verlangen  
> _Software engineer_
> 
> - 
> - 
> 
> W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> E robin@us2.nl
> 
> [http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)
> 
> Disclaimer: The information contained in this message and attachments is  
> intended solely for the attention and use of the named addressee and may be  
> confidential. If you are not the intended recipient, you are reminded that  
> the information remains the property of the sender. You must not use,  
> disclose, distribute, copy, print or rely on this e-mail. If you have  
> received this message in error, please contact the sender immediately and  
> irrevocably delete this message and any copies.
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> --

--

---

<div class="post-metadata">

**Author:** ![Karussell\_2](https://avatars.discourse-cdn.com/v4/letter/k/54ee81/32.png) [@Karussell\_2](https://discuss.elastic.co/u/Karussell_2)\
**Post date:** [January 5, 2013, 8:15pm UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/4 "2013-01-05T20:15:21Z")

</div>

As David pointed out there will appear more tools for this in the upcoming  
versions.

Until then you can try:

- you need to flush (+ probably disable indexing), then rsync to your  
backup folder. [Backup ElasticSearch with rsync · GitHub](https://gist.github.com/1074906)
- you need to flush, then snapshot via AWS
- use my reindexing plugin to copy a certain index or only a subset of the  
data into a completely different cluster (probably not efficient regarding  
network IO as its using only gzipped+json not binary json etc)  
[GitHub - karussell/elasticsearch-reindex: Simple re-indexing. To backup, apply index settings changes and more ElasticMagic](https://github.com/karussell/elasticsearch-reindex)
- then there is the this plugin  
[GitHub - jprante/elasticsearch-knapsack: Knapsack plugin is an import/export tool for Elasticsearch](https://github.com/jprante/elasticsearch-knapsack) where you can do the same  
probably a bit more efficient as it is tarred + zipped

Regards,  
Peter.

On Wednesday, January 2, 2013 2:52:40 PM UTC+1, Robin Verlangen wrote:

> Hi David,
> 
> Thank you for your reply. The second gist is about the same as I was  
> thinking about. Seems that I'll have to keep track of my data, or find it  
> through the API.
> 
> Best regards,
> 
> Robin Verlangen  
> _Software engineer_  
> \*  
> \*  
> W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> E ro...@us2.nl \<javascript:\>
> 
> [http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)
> 
> Disclaimer: The information contained in this message and attachments is  
> intended solely for the attention and use of the named addressee and may be  
> confidential. If you are not the intended recipient, you are reminded that  
> the information remains the property of the sender. You must not use,  
> disclose, distribute, copy, print or rely on this e-mail. If you have  
> received this message in error, please contact the sender immediately and  
> irrevocably delete this message and any copies.
> 
> On Wed, Jan 2, 2013 at 11:05 AM, David Pilato \<[da...@pilato.fr](mailto:da...@pilato.fr)\<javascript:\>
> 
> > wrote:
> 
> > \*\*  
> > Hey Robin,
> > 
> > I found this Gist: [Backup and restore an Elastic search index (shamelessly copied from http://tech.superhappykittymeow.com/?p=296) · GitHub](https://gist.github.com/1939828)
> > 
> > I never tested it myself.
> > 
> > Have a look also at: [Backup Elasticsearch node · GitHub](https://gist.github.com/4380715)  
> > Never tested it also...
> > 
> > That said, I think that in the next versions, some tools will appears  
> > and probably will cover this need.
> > 
> > HTH  
> > David
> > 
> > Le 2 janvier 2013 à 10:35, Robin Verlangen \<ro...@us2.nl \<javascript:\>\>  
> > a écrit :
> > 
> > Hi there,
> > 
> > What would be the best way to create backups of the data in ES? In my  
> > use case the data is append-only (log storage etc).
> > 
> > I was thinking about closing the index, gzip the entire folder on every  
> > node (not so easy to coordinate), and store it on something like Amazon S3.  
> > However I would like to know whether there are more solid solutions to  
> > this.
> > 
> > Best regards,
> > 
> > Robin Verlangen  
> > _Software engineer_
> > 
> > - 
> > - 
> > 
> > W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> > E ro...@us2.nl \<javascript:\>
> > 
> > [http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)
> > 
> > Disclaimer: The information contained in this message and attachments  
> > is intended solely for the attention and use of the named addressee and may  
> > be confidential. If you are not the intended recipient, you are reminded  
> > that the information remains the property of the sender. You must not use,  
> > disclose, distribute, copy, print or rely on this e-mail. If you have  
> > received this message in error, please contact the sender immediately and  
> > irrevocably delete this message and any copies.
> > 
> > --
> > 
> > --  
> > David Pilato  
> > [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > 
> > --

--

---

<div class="post-metadata">

**Author:** ![Karel\_Minarik\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karel_minarik_2/32/1078_2.png) [@Karel\_Minarik\_2](https://discuss.elastic.co/u/Karel_Minarik_2)\
**Post date:** [January 6, 2013, 7:41am UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/5 "2013-01-06T07:41:16Z")

</div>

I was thinking about closing the index, gzip the entire folder on every  
node (not so easy to coordinate), and store it on something like Amazon S3.  
However I would like to know whether there are more solid solutions to this.

That is essentially correct. You don't have to close the index; as others  
pointed out, flush  
[[http://www.elasticsearch.org/guide/reference/api/admin-indices-flush.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-flush.html)]  
the index, disable flush for the index  
[[http://www.elasticsearch.org/guide/reference/api/admin-indices-update-settings.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-update-settings.html)],  
create the archive or rsync to the backup location, and enable flush when  
that operation is done. You don't have to disable indexing etc. You don't  
have to attempt to "synchronize" the gzipping.

In a near future, there will be a snapshot/restore API in Elasticsearch  
itself.

Karel

--

---

<div class="post-metadata">

**Author:** ![radu\_gheorghe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe/32/556_2.png) [@radu\_gheorghe](https://discuss.elastic.co/u/radu_gheorghe)\
**Post date:** [January 7, 2013, 8:26am UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/6 "2013-01-07T08:26:24Z")

</div>

Hello,

Below is yet another gist. I'm thinking the more, the better 🙂

> <https://gist.github.com/radu-gheorghe/3180985>

This is especially useful if you have time-based data, and you'd want to  
optimize your indices before "archiving" them. This would make better use  
of your disk space and your searches would be faster.

## Best regards, Radu

[http://sematext.com/](http://sematext.com/) -- Elasticsearch -- Solr -- Lucene

On Sun, Jan 6, 2013 at 9:41 AM, Karel Minařík \<  
[karel.minarik@elasticsearch.com](mailto:karel.minarik@elasticsearch.com)\> wrote:

> I was thinking about closing the index, gzip the entire folder on every  
> node (not so easy to coordinate), and store it on something like Amazon S3.  
> However I would like to know whether there are more solid solutions to this.
> 
> That is essentially correct. You don't have to close the index; as others  
> pointed out, flush [  
> [Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-indices-flush.html)]  
> the index, disable flush for the index [  
> [Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-indices-update-settings.html)],  
> create the archive or rsync to the backup location, and enable flush when  
> that operation is done. You don't have to disable indexing etc. You don't  
> have to attempt to "synchronize" the gzipping.
> 
> In a near future, there will be a snapshot/restore API in Elasticsearch  
> itself.
> 
> Karel
> 
> --

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [January 7, 2013, 11:17am UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/7 "2013-01-07T11:17:51Z")

</div>

Just a disclaimer: while it is easy to package your Elasticsearch \_source  
documents with the knapsack plugin into a tar.gz archive, it is not  
intended as a backup/restore solution and you can not expect it to work  
correctly in all situations. Precautions should be taken. For example when  
the index is "hot", the tool can't ensure to capture a consistent set of  
docs.

Jörg

On Saturday, January 5, 2013 9:15:21 PM UTC+1, Karussell wrote:

> - then there is the this plugin  
> [GitHub - jprante/elasticsearch-knapsack: Knapsack plugin is an import/export tool for Elasticsearch](https://github.com/jprante/elasticsearch-knapsack) where you can do the  
> same probably a bit more efficient as it is tarred + zipped
> 
> >

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:57am UTC](https://discuss.elastic.co/t/backup-data-in-a-robust-way/10214/8 "2017-07-06T02:57:35Z")

</div>


