# How persistence works in ElasticSearch

**URL:** <https://discuss.elastic.co/t/how-persistence-works-in-elasticsearch/2872>\
**Category:** Elasticsearch\
**Created:** [March 26, 2010, 8:57pm UTC](https://discuss.elastic.co/t/how-persistence-works-in-elasticsearch/2872 "2010-03-26T20:57:41Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Berkay\_Mollamustafao](https://avatars.discourse-cdn.com/v4/letter/b/22d042/32.png) [@Berkay\_Mollamustafao](https://discuss.elastic.co/u/Berkay_Mollamustafao)\
**Post date:** [March 26, 2010, 8:57pm UTC](https://discuss.elastic.co/t/how-persistence-works-in-elasticsearch/2872/1 "2010-03-26T20:57:41Z")

</div>

I can use some help verifying/understanding how persistence works. Here is  
my understanding of how it works:

Regardless of whether the index is stored in memory or file system, it is  
considered temporary and removed when the node is stopped, hence if all the  
nodes in the cluster stop, indices would be lost.

As such, for persistence write behind gateway needs to be used. Gateway  
keeps a transaction log and (periodically?) creates indices. If all the  
nodes in the cluster were stopped and restarting, the indices and the  
transaction logs created by the gateway are used to recreate node indices.

Is this right?

Regards,  
Berkay Mollamustafaoglu  
mberkay on yahoo, google and skype

On Fri, Mar 26, 2010 at 1:10 PM, Tim Robertson [timrobertson100@gmail.com](mailto:timrobertson100@gmail.com)wrote:

> Thanks Shay,
> 
> So... reading between the lines, does it then use protobufs (or other?) for  
> RPC instead of JSON and serializing and deserializing?
> 
> Cheers  
> Tim
> 
> On Fri, Mar 26, 2010 at 3:58 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> 
> > Hi,
> > 
> > You won't enjoy locally between elasticsearch and hadoop in any case  
> > since both use different distribution model. The locality would only make  
> > sense for the indexing part, and think that you probably won't really need  
> > it (it should be fast enough).
> > 
> > What language are you going to write your jobs at? If Java, then make  
> > use of the native Java client (obtained from a "non data" Server started)  
> > and not HTTP. More here:  
> > [http://www.elasticsearch.com/docs/elasticsearch/java\_api/client/#Server\_Client](http://www.elasticsearch.com/docs/elasticsearch/java_api/client/#Server_Client)
> > 
> > -shay.banon
> > 
> > On Fri, Mar 26, 2010 at 5:35 PM, Tim Robertson \<[timrobertson100@gmail.com](mailto:timrobertson100@gmail.com)
> > 
> > > wrote:
> > 
> > > Hey,
> > > 
> > > Is anyone building their indexes using Hadoop? If so, are they deploying  
> > > ES across the same cluster as Hadoop and trying to reduce network noise by  
> > > making use of data locality, or keeping the clusters separate and just  
> > > calling over HTTP from MapReduce when building the indexes? I am about to  
> > > set up on EC2, and planned to keep the search and processing machines  
> > > separate.
> > > 
> > > Cheers,  
> > > Tim

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 26, 2010, 9:04pm UTC](https://discuss.elastic.co/t/how-persistence-works-in-elasticsearch/2872/2 "2010-03-26T21:04:10Z")

</div>

Yes, thats basically how it works. Regarding the transaction log, basically,  
the gateway is responsible for mirroring the current shard lucene index, and  
the delta transaction log. When a commit occurs (either through an API call,  
or automatically by elasticsearch), the transaction log is flushed.

The benefit of this architecture is that the indexable state of the cluster  
can be written in an async manner, and, the actual storage of the index is  
irrelevant for long term persistency, which means you can still store the  
index in memory (or just parts of it, with the upcoming cacheable FS  
storage), and not loose it on failure.

-shay.banon

On Fri, Mar 26, 2010 at 11:57 PM, Berkay Mollamustafaoglu \<[mberkay@gmail.com](mailto:mberkay@gmail.com)

> wrote:

> I can use some help verifying/understanding how persistence works. Here is  
> my understanding of how it works:
> 
> Regardless of whether the index is stored in memory or file system, it is  
> considered temporary and removed when the node is stopped, hence if all the  
> nodes in the cluster stop, indices would be lost.
> 
> As such, for persistence write behind gateway needs to be used. Gateway  
> keeps a transaction log and (periodically?) creates indices. If all the  
> nodes in the cluster were stopped and restarting, the indices and the  
> transaction logs created by the gateway are used to recreate node indices.
> 
> Is this right?
> 
> Regards,  
> Berkay Mollamustafaoglu  
> mberkay on yahoo, google and skype
> 
> On Fri, Mar 26, 2010 at 1:10 PM, Tim Robertson [timrobertson100@gmail.com](mailto:timrobertson100@gmail.com)wrote:
> 
> > Thanks Shay,
> > 
> > So... reading between the lines, does it then use protobufs (or other?)  
> > for RPC instead of JSON and serializing and deserializing?
> > 
> > Cheers  
> > Tim
> > 
> > On Fri, Mar 26, 2010 at 3:58 PM, Shay Banon \<[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)
> > 
> > > wrote:
> > 
> > > Hi,
> > > 
> > > You won't enjoy locally between elasticsearch and hadoop in any case  
> > > since both use different distribution model. The locality would only make  
> > > sense for the indexing part, and think that you probably won't really need  
> > > it (it should be fast enough).
> > > 
> > > What language are you going to write your jobs at? If Java, then make  
> > > use of the native Java client (obtained from a "non data" Server started)  
> > > and not HTTP. More here:  
> > > [http://www.elasticsearch.com/docs/elasticsearch/java\_api/client/#Server\_Client](http://www.elasticsearch.com/docs/elasticsearch/java_api/client/#Server_Client)
> > > 
> > > -shay.banon
> > > 
> > > On Fri, Mar 26, 2010 at 5:35 PM, Tim Robertson \<  
> > > [timrobertson100@gmail.com](mailto:timrobertson100@gmail.com)\> wrote:
> > > 
> > > > Hey,
> > > > 
> > > > Is anyone building their indexes using Hadoop? If so, are they  
> > > > deploying ES across the same cluster as Hadoop and trying to reduce network  
> > > > noise by making use of data locality, or keeping the clusters separate and  
> > > > just calling over HTTP from MapReduce when building the indexes? I am about  
> > > > to set up on EC2, and planned to keep the search and processing machines  
> > > > separate.
> > > > 
> > > > Cheers,  
> > > > Tim

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:25am UTC](https://discuss.elastic.co/t/how-persistence-works-in-elasticsearch/2872/3 "2017-07-06T04:25:03Z")

</div>


