# Persistency

**URL:** <https://discuss.elastic.co/t/persistency/2808>\
**Category:** Elasticsearch\
**Created:** [February 13, 2010, 9:03pm UTC](https://discuss.elastic.co/t/persistency/2808 "2010-02-13T21:03:32Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![aemadrid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aemadrid/32/3185_2.png) [@aemadrid](https://discuss.elastic.co/u/aemadrid)\
**Post date:** [February 13, 2010, 9:03pm UTC](https://discuss.elastic.co/t/persistency/2808/1 "2010-02-13T21:03:32Z")

</div>

I'm very interested in ES and I'm wondering how you can do persistency  
and how it impacts performance. Can you elaborate a little bit more on  
that?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 13, 2010, 9:57pm UTC](https://discuss.elastic.co/t/persistency/2808/2 "2010-02-13T21:57:44Z")

</div>

Sure. Basically, persistency is done in a "write behind manner" and it  
called gateway in elasticsearch. The cluster meta data and including indices  
that you choose to are written to a long term persistent storage  
asynchronously in the background.

The cluster meta data include all the indices (and their settings) that have  
been created in the cluster (not including the indices content). This is  
written to the gateway every time something changes (like creating a new  
index). Each index in turn can also be persistent to a gateway. The index  
content itself is persistent to the gateway, along with its transaction log.

In terms of high availability, if you create an index with 1 replica per  
shard, then you have "real time" high availability, which means that if a  
node fails, there is a replica for it. Of course, you can always create more  
than one replica per shard. The replication to the replicas is done in  
parallel. The nice thing about elasticsearch is that reads go to one of the  
replica shards, which means that you can scale reads/search with more  
replicas.

Each shard replication group has a primary shard, which is responsible for  
persisting the index and the transaction log into the long term gateway  
storage. This is done in a scheduled manner (you can control it), and there  
is even an API to force it (the gateway snapshot API).

The snapshotting of the index into the persistent storage is done in the  
background, so there is no impact on performance of actual operations done  
against the index.

The gateway module itself (both the cluster one, and the index one) are  
completely pluggable. The current implementation is a file system based one.  
There are more to come including cloud based ones to persist to Amazon S3  
for example.

This solution is my preferred solution for handling long term persistency of  
of a cluster since it means that node storage is completely temporal. This  
in turn means that you can store the index in memory for example, get the  
performance benefits that comes with it, without scarifying long term  
persistency.

Some docs links:

[http://www.elasticsearch.com/docs/elasticsearch/modules/gateway/](http://www.elasticsearch.com/docs/elasticsearch/modules/gateway/)  
[http://www.elasticsearch.com/docs/elasticsearch/index\_modules/gateway/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/gateway/)  
[http://www.elasticsearch.com/docs/elasticsearch/index\_modules/store/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/store/)

-shay.banon

On Sat, Feb 13, 2010 at 11:03 PM, aemadrid [aemadrid@gmail.com](mailto:aemadrid@gmail.com) wrote:

> I'm very interested in ES and I'm wondering how you can do persistency  
> and how it impacts performance. Can you elaborate a little bit more on  
> that?

---

<div class="post-metadata">

**Author:** ![aemadrid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aemadrid/32/3185_2.png) [@aemadrid](https://discuss.elastic.co/u/aemadrid)\
**Post date:** [February 14, 2010, 3:55am UTC](https://discuss.elastic.co/t/persistency/2808/3 "2010-02-14T03:55:20Z")

</div>

Thanks so much for the explanation. It makes sense to me although I'm  
getting a little bit lost on the formulas for shards/replicas. Something  
tells me that there would be a couple of scenarios when it comes to HA. Some  
people )like me) are after HA in the smallest sense possible: 2 servers to  
keep everything running in case one fails and everything to disk in case  
both fail. Some other people I bet are after the 3+ servers (10+, 100+).  
Could you write/video how you would go about setting something like those  
scenarios 2 up? One of the selling points for me is how servers find each  
other (only localhost or local network too?) and remaster each other on  
failures. It seems that setting up persistence is a little more involved.

Thanks in advance,

Adrian Madrid  
My eBiz, Senior Developer  
3082 W. Maple Loop Dr  
Lehi, UT 84043  
801-341-3824

On Sat, Feb 13, 2010 at 14:57, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Sure. Basically, persistency is done in a "write behind manner" and it  
> called gateway in elasticsearch. The cluster meta data and including indices  
> that you choose to are written to a long term persistent storage  
> asynchronously in the background.
> 
> The cluster meta data include all the indices (and their settings) that  
> have been created in the cluster (not including the indices content). This  
> is written to the gateway every time something changes (like creating a new  
> index). Each index in turn can also be persistent to a gateway. The index  
> content itself is persistent to the gateway, along with its transaction log.
> 
> In terms of high availability, if you create an index with 1 replica per  
> shard, then you have "real time" high availability, which means that if a  
> node fails, there is a replica for it. Of course, you can always create more  
> than one replica per shard. The replication to the replicas is done in  
> parallel. The nice thing about elasticsearch is that reads go to one of the  
> replica shards, which means that you can scale reads/search with more  
> replicas.
> 
> Each shard replication group has a primary shard, which is responsible for  
> persisting the index and the transaction log into the long term gateway  
> storage. This is done in a scheduled manner (you can control it), and there  
> is even an API to force it (the gateway snapshot API).
> 
> The snapshotting of the index into the persistent storage is done in the  
> background, so there is no impact on performance of actual operations done  
> against the index.
> 
> The gateway module itself (both the cluster one, and the index one) are  
> completely pluggable. The current implementation is a file system based one.  
> There are more to come including cloud based ones to persist to Amazon S3  
> for example.
> 
> This solution is my preferred solution for handling long term persistency  
> of of a cluster since it means that node storage is completely temporal.  
> This in turn means that you can store the index in memory for example, get  
> the performance benefits that comes with it, without scarifying long term  
> persistency.
> 
> Some docs links:
> 
> [http://www.elasticsearch.com/docs/elasticsearch/modules/gateway/](http://www.elasticsearch.com/docs/elasticsearch/modules/gateway/)  
> [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/gateway/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/gateway/)  
> [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/store/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/store/)
> 
> -shay.banon
> 
> On Sat, Feb 13, 2010 at 11:03 PM, aemadrid [aemadrid@gmail.com](mailto:aemadrid@gmail.com) wrote:
> 
> > I'm very interested in ES and I'm wondering how you can do persistency  
> > and how it impacts performance. Can you elaborate a little bit more on  
> > that?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 14, 2010, 5:56am UTC](https://discuss.elastic.co/t/persistency/2808/4 "2010-02-14T05:56:50Z")

</div>

The discovery module uses jgroups, which in turn can use multicast/unicast.  
I will document it later today/tomorrow more properly. It works across  
machines.

In case of simple, two servers, I suggest you deploy either 2 shards, each  
with 1 replica, or 4 shards, each with one replica. In this case, you will  
have either 2 shard on each machine , or 4 shards on each machine. The  
reason for the extra shards is the ability to grow in the future and for  
added concurrency.

If things are a bit more difficult, just leave everything as is (it defaults  
to 5 shards, each with one replica), and change the gateway configuration in  
the elasticsearch.yml fie to:

gateway:  
type : fs  
fs :  
location : /shared/fs/location

-shay.banon

On Sun, Feb 14, 2010 at 5:55 AM, Adrian Madrid [aemadrid@gmail.com](mailto:aemadrid@gmail.com) wrote:

> Thanks so much for the explanation. It makes sense to me although I'm  
> getting a little bit lost on the formulas for shards/replicas. Something  
> tells me that there would be a couple of scenarios when it comes to HA. Some  
> people )like me) are after HA in the smallest sense possible: 2 servers to  
> keep everything running in case one fails and everything to disk in case  
> both fail. Some other people I bet are after the 3+ servers (10+, 100+).  
> Could you write/video how you would go about setting something like those  
> scenarios 2 up? One of the selling points for me is how servers find each  
> other (only localhost or local network too?) and remaster each other on  
> failures. It seems that setting up persistence is a little more involved.
> 
> Thanks in advance,
> 
> Adrian Madrid  
> My eBiz, Senior Developer  
> 3082 W. Maple Loop Dr  
> Lehi, UT 84043  
> 801-341-3824
> 
> On Sat, Feb 13, 2010 at 14:57, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> 
> > Sure. Basically, persistency is done in a "write behind manner" and it  
> > called gateway in elasticsearch. The cluster meta data and including indices  
> > that you choose to are written to a long term persistent storage  
> > asynchronously in the background.
> > 
> > The cluster meta data include all the indices (and their settings) that  
> > have been created in the cluster (not including the indices content). This  
> > is written to the gateway every time something changes (like creating a new  
> > index). Each index in turn can also be persistent to a gateway. The index  
> > content itself is persistent to the gateway, along with its transaction log.
> > 
> > In terms of high availability, if you create an index with 1 replica per  
> > shard, then you have "real time" high availability, which means that if a  
> > node fails, there is a replica for it. Of course, you can always create more  
> > than one replica per shard. The replication to the replicas is done in  
> > parallel. The nice thing about elasticsearch is that reads go to one of the  
> > replica shards, which means that you can scale reads/search with more  
> > replicas.
> > 
> > Each shard replication group has a primary shard, which is responsible for  
> > persisting the index and the transaction log into the long term gateway  
> > storage. This is done in a scheduled manner (you can control it), and there  
> > is even an API to force it (the gateway snapshot API).
> > 
> > The snapshotting of the index into the persistent storage is done in the  
> > background, so there is no impact on performance of actual operations done  
> > against the index.
> > 
> > The gateway module itself (both the cluster one, and the index one) are  
> > completely pluggable. The current implementation is a file system based one.  
> > There are more to come including cloud based ones to persist to Amazon S3  
> > for example.
> > 
> > This solution is my preferred solution for handling long term persistency  
> > of of a cluster since it means that node storage is completely temporal.  
> > This in turn means that you can store the index in memory for example, get  
> > the performance benefits that comes with it, without scarifying long term  
> > persistency.
> > 
> > Some docs links:
> > 
> > [http://www.elasticsearch.com/docs/elasticsearch/modules/gateway/](http://www.elasticsearch.com/docs/elasticsearch/modules/gateway/)  
> > [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/gateway/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/gateway/)  
> > [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/store/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/store/)
> > 
> > -shay.banon
> > 
> > On Sat, Feb 13, 2010 at 11:03 PM, aemadrid [aemadrid@gmail.com](mailto:aemadrid@gmail.com) wrote:
> > 
> > > I'm very interested in ES and I'm wondering how you can do persistency  
> > > and how it impacts performance. Can you elaborate a little bit more on  
> > > that?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:25am UTC](https://discuss.elastic.co/t/persistency/2808/5 "2017-07-06T04:25:46Z")

</div>


