# Replication Strategies

**URL:** <https://discuss.elastic.co/t/replication-strategies/3081>\
**Category:** Elasticsearch\
**Created:** [July 8, 2010, 12:40pm UTC](https://discuss.elastic.co/t/replication-strategies/3081 "2010-07-08T12:40:08Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Franz\_Allan\_Valencia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/franz_allan_valencia/32/3326_2.png) [@Franz\_Allan\_Valencia](https://discuss.elastic.co/u/Franz_Allan_Valencia)\
**Post date:** [July 8, 2010, 12:40pm UTC](https://discuss.elastic.co/t/replication-strategies/3081/1 "2010-07-08T12:40:08Z")

</div>

Curious, if you guys are using ElasticSearch for replication, how do you  
replicate your data?

I'm assuming your main data source is slow, so I was wondering how you go  
around it to have a faster replication.

Do you do your huge/slow query on your main data store and store that in  
ElasticSearch, or do you guys query expensive tables and store those in  
ElasticSearch, and then later recombine them? ...or? 🙂

Do you guys do lazy replication wherein the first time a query is made, you  
retrieve data from your main data source, then replicate in ElasticSearch?  
Or do you do a scheduled replication? ...or? 🙂

Thanks,

--  
Franz Allan Valencia See | Java Software Engineer  
[franz.see@gmail.com](mailto:franz.see@gmail.com)  
LinkedIn: [http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
Twitter: [http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 8, 2010, 2:34pm UTC](https://discuss.elastic.co/t/replication-strategies/3081/2 "2010-07-08T14:34:21Z")

</div>

Replication means several things when it comes to elasticsearch, you mean  
"replicating" changes done to the actual data source to elasticsearch.  
Usually the best thing is to build a system that can apply changes done to  
the data source back to elasticsearch. How you go about applying changes  
done to the data source into elasticsearch can be done in several manners,  
depending on the data source:

1. If the data source has hooks that allow to be notified when something  
changes in it, then a custom elasticsearch hook can be written to apply  
those changes.

2. If the data source provides a stream of changes done to it, that stream  
can be used to apply the changes to elasticsearch.

3. A custom poller for changes can be written that polls the data source for  
changes (for example, based on last poll timestamp) and apply them to  
elasticsearch.

-shay.banon

On Thu, Jul 8, 2010 at 3:40 PM, Franz Allan Valencia See \<  
[franz.see@gmail.com](mailto:franz.see@gmail.com)\> wrote:

> Curious, if you guys are using Elasticsearch for replication, how do you  
> replicate your data?
> 
> I'm assuming your main data source is slow, so I was wondering how you go  
> around it to have a faster replication.
> 
> Do you do your huge/slow query on your main data store and store that in  
> Elasticsearch, or do you guys query expensive tables and store those in  
> Elasticsearch, and then later recombine them? ...or? 🙂
> 
> Do you guys do lazy replication wherein the first time a query is made, you  
> retrieve data from your main data source, then replicate in Elasticsearch?  
> Or do you do a scheduled replication? ...or? 🙂
> 
> Thanks,
> 
> --  
> Franz Allan Valencia See | Java Software Engineer  
> [franz.see@gmail.com](mailto:franz.see@gmail.com)  
> LinkedIn: [http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
> Twitter: [http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

---

<div class="post-metadata">

**Author:** ![Samuel\_Doyle](https://avatars.discourse-cdn.com/v4/letter/s/e19adc/32.png) [@Samuel\_Doyle](https://discuss.elastic.co/u/Samuel_Doyle)\
**Post date:** [July 8, 2010, 5:03pm UTC](https://discuss.elastic.co/t/replication-strategies/3081/3 "2010-07-08T17:03:41Z")

</div>

At the moment I'm using polling but since the main application is making use  
of Hibernate I may try to leverage a post flush interceptor for incremental  
updating.  
The initial main indexing is done via straight SQL.

S.D.

On Thu, Jul 8, 2010 at 5:40 AM, Franz Allan Valencia See \<  
[franz.see@gmail.com](mailto:franz.see@gmail.com)\> wrote:

> Curious, if you guys are using Elasticsearch for replication, how do you  
> replicate your data?
> 
> I'm assuming your main data source is slow, so I was wondering how you go  
> around it to have a faster replication.
> 
> Do you do your huge/slow query on your main data store and store that in  
> Elasticsearch, or do you guys query expensive tables and store those in  
> Elasticsearch, and then later recombine them? ...or? 🙂
> 
> Do you guys do lazy replication wherein the first time a query is made, you  
> retrieve data from your main data source, then replicate in Elasticsearch?  
> Or do you do a scheduled replication? ...or? 🙂
> 
> Thanks,
> 
> --  
> Franz Allan Valencia See | Java Software Engineer  
> [franz.see@gmail.com](mailto:franz.see@gmail.com)  
> LinkedIn: [http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
> Twitter: [http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

---

<div class="post-metadata">

**Author:** ![Franz\_Allan\_Valencia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/franz_allan_valencia/32/3326_2.png) [@Franz\_Allan\_Valencia](https://discuss.elastic.co/u/Franz_Allan_Valencia)\
**Post date:** [July 9, 2010, 2:25am UTC](https://discuss.elastic.co/t/replication-strategies/3081/4 "2010-07-09T02:25:02Z")

</div>

Thanks everyone.

Pardon Shay, let me elaborate 🙂

Currently, my problems are on  
a.) Getting the initial data from the main data source (MDS) to replicated  
data source (RDS) (which is Elasticsearch), and  
b.) Keeping the data in the RDS updated

For a.) these are my choices so far:  
a.1.) Query from the MDS everything that I will replicate and put them all  
in the RDS  
a.2.) Query from the MDS upon demand the information and put that in the RDS  
a.3.) Query from the MDS the heavy parts and put those in the RDS and then  
query the rest of the data on demand from the MDS and combine that with the  
data in the RDS, and then put the result back in the RDS (i.e. `original query is select * from tbl_small, tbl_big where ...`. So i first replicate  
`tbl_big`. and then when the user does an action, I will start querying from  
tbl\_small then combine that to the replicated tbl\_big and put the result in  
the query\_aggregate RDS so that future queries only hit the RDS)

For b.) these are my choices so far:  
b.1.) Have an event listener that when data in the MDS, the listener will  
update the data in the RDS (we're currently considering hibernate listener  
or db trigger)  
b.2.) Have a scheduled refresh of those that changed since last refresh

But since I don't know much about Replication & Caching  
Patterns/Anti-Patterns, I would like to inquire on how the community does it  
🙂

Thoughts?

--  
Franz Allan Valencia See | Java Software Engineer  
[franz.see@gmail.com](mailto:franz.see@gmail.com)  
LinkedIn: [http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
Twitter: [http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

On Fri, Jul 9, 2010 at 1:03 AM, Samuel Doyle [samueldoyle@gmail.com](mailto:samueldoyle@gmail.com) wrote:

> At the moment I'm using polling but since the main application is making  
> use of Hibernate I may try to leverage a post flush interceptor for  
> incremental updating.  
> The initial main indexing is done via straight SQL.
> 
> S.D.
> 
> On Thu, Jul 8, 2010 at 5:40 AM, Franz Allan Valencia See \<  
> [franz.see@gmail.com](mailto:franz.see@gmail.com)\> wrote:
> 
> > Curious, if you guys are using Elasticsearch for replication, how do you  
> > replicate your data?
> > 
> > I'm assuming your main data source is slow, so I was wondering how you go  
> > around it to have a faster replication.
> > 
> > Do you do your huge/slow query on your main data store and store that in  
> > Elasticsearch, or do you guys query expensive tables and store those in  
> > Elasticsearch, and then later recombine them? ...or? 🙂
> > 
> > Do you guys do lazy replication wherein the first time a query is made,  
> > you retrieve data from your main data source, then replicate in  
> > Elasticsearch? Or do you do a scheduled replication? ...or? 🙂
> > 
> > Thanks,
> > 
> > --  
> > Franz Allan Valencia See | Java Software Engineer  
> > [franz.see@gmail.com](mailto:franz.see@gmail.com)  
> > LinkedIn: [http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
> > Twitter: [http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [July 9, 2010, 8:22pm UTC](https://discuss.elastic.co/t/replication-strategies/3081/5 "2010-07-09T20:22:42Z")

</div>

Franz, I think you really just need to do something similar to what  
DIH (DataImportHandler) in Solr does, which is roughly this:

- Have a tool that can do a bulk import/indexing
- Have a tool (perhaps it's the same tool) that can do incremental  
import/indexing. It can do that by keeping track of the last imported  
ID or by tracking the timestamp of the last import or some such.

The bulk import is something you'd do once, at the beginning, when  
your ES index is empty. Then you would periodically run the tool to  
do the incremental indexing. You would also use bulk importer/indexer  
if you have to reindex from scratch for whatever reason.

The above will not get your data into ES in real-time, so if you need  
real-time search, you need a real-time approach instead of the  
incremental one. In that case you need to have a mechanism that  
inserts individual records/documents into ES as soon as they are added  
to your data store. Some data stores have hooks to make this kind of  
automatic, like MongoDB, Terrastore, etc. Hibernate Search does this  
for relational databases and Lucene (which is the low level Java  
library ES uses for indexing/searching).

## Otis

Sematext :: [http://sematext.com/](http://sematext.com/) :: Solr - Lucene - Nutch  
Lucene ecosystem search :: [http://search-lucene.com/](http://search-lucene.com/)

On Jul 8, 10:25 pm, Franz Allan Valencia See [franz....@gmail.com](mailto:franz....@gmail.com)  
wrote:

> Thanks everyone.
> 
> Pardon Shay, let me elaborate 🙂
> 
> Currently, my problems are on  
> a.) Getting the initial data from the main data source (MDS) to replicated  
> data source (RDS) (which is Elasticsearch), and  
> b.) Keeping the data in the RDS updated
> 
> For a.) these are my choices so far:  
> a.1.) Query from the MDS everything that I will replicate and put them all  
> in the RDS  
> a.2.) Query from the MDS upon demand the information and put that in the RDS  
> a.3.) Query from the MDS the heavy parts and put those in the RDS and then  
> query the rest of the data on demand from the MDS and combine that with the  
> data in the RDS, and then put the result back in the RDS (i.e. `original query is select * from tbl_small, tbl_big where ...`. So i first replicate  
> `tbl_big`. and then when the user does an action, I will start querying from  
> tbl\_small then combine that to the replicated tbl\_big and put the result in  
> the query\_aggregate RDS so that future queries only hit the RDS)
> 
> For b.) these are my choices so far:  
> b.1.) Have an event listener that when data in the MDS, the listener will  
> update the data in the RDS (we're currently considering hibernate listener  
> or db trigger)  
> b.2.) Have a scheduled refresh of those that changed since last refresh
> 
> But since I don't know much about Replication & Caching  
> Patterns/Anti-Patterns, I would like to inquire on how the community does it  
> 🙂
> 
> Thoughts?
> 
> --  
> Franz Allan Valencia See | Java Software Engineer  
> [franz....@gmail.com](mailto:franz....@gmail.com)  
> LinkedIn:[Franz Allan See - PDAX | LinkedIn](http://www.linkedin.com/in/franzsee)  
> Twitter:[http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)
> 
> On Fri, Jul 9, 2010 at 1:03 AM, Samuel Doyle [samueldo...@gmail.com](mailto:samueldo...@gmail.com) wrote:
> 
> > At the moment I'm using polling but since the main application is making  
> > use of Hibernate I may try to leverage a post flush interceptor for  
> > incremental updating.  
> > The initial main indexing is done via straight SQL.
> 
> > S.D.
> 
> > On Thu, Jul 8, 2010 at 5:40 AM, Franz Allan Valencia See \<  
> > [franz....@gmail.com](mailto:franz....@gmail.com)\> wrote:
> 
> > > Curious, if you guys are using Elasticsearch for replication, how do you  
> > > replicate your data?
> 
> > > I'm assuming your main data source is slow, so I was wondering how you go  
> > > around it to have a faster replication.
> 
> > > Do you do your huge/slow query on your main data store and store that in  
> > > Elasticsearch, or do you guys query expensive tables and store those in  
> > > Elasticsearch, and then later recombine them? ...or? 🙂
> 
> > > Do you guys do lazy replication wherein the first time a query is made,  
> > > you retrieve data from your main data source, then replicate in  
> > > Elasticsearch? Or do you do a scheduled replication? ...or? 🙂
> 
> > > Thanks,
> 
> > > --  
> > > Franz Allan Valencia See | Java Software Engineer  
> > > [franz....@gmail.com](mailto:franz....@gmail.com)  
> > > LinkedIn:[Franz Allan See - PDAX | LinkedIn](http://www.linkedin.com/in/franzsee)  
> > > Twitter:[http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

---

<div class="post-metadata">

**Author:** ![Franz\_Allan\_Valencia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/franz_allan_valencia/32/3326_2.png) [@Franz\_Allan\_Valencia](https://discuss.elastic.co/u/Franz_Allan_Valencia)\
**Post date:** [July 12, 2010, 3:05am UTC](https://discuss.elastic.co/t/replication-strategies/3081/6 "2010-07-12T03:05:16Z")

</div>

I guess I'm in the right path then.

## Thanks,

Franz Allan Valencia See | Java Software Engineer  
[franz.see@gmail.com](mailto:franz.see@gmail.com)  
LinkedIn: [http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
Twitter: [http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

On Sat, Jul 10, 2010 at 4:22 AM, Otis [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com) wrote:

> Franz, I think you really just need to do something similar to what  
> DIH (DataImportHandler) in Solr does, which is roughly this:
> 
> - Have a tool that can do a bulk import/indexing
> - Have a tool (perhaps it's the same tool) that can do incremental  
> import/indexing. It can do that by keeping track of the last imported  
> ID or by tracking the timestamp of the last import or some such.
> 
> The bulk import is something you'd do once, at the beginning, when  
> your ES index is empty. Then you would periodically run the tool to  
> do the incremental indexing. You would also use bulk importer/indexer  
> if you have to reindex from scratch for whatever reason.
> 
> The above will not get your data into ES in real-time, so if you need  
> real-time search, you need a real-time approach instead of the  
> incremental one. In that case you need to have a mechanism that  
> inserts individual records/documents into ES as soon as they are added  
> to your data store. Some data stores have hooks to make this kind of  
> automatic, like MongoDB, Terrastore, etc. Hibernate Search does this  
> for relational databases and Lucene (which is the low level Java  
> library ES uses for indexing/searching).
> 
> ## Otis
> 
> Sematext :: [http://sematext.com/](http://sematext.com/) :: Solr - Lucene - Nutch  
> Lucene ecosystem search :: [http://search-lucene.com/](http://search-lucene.com/)
> 
> On Jul 8, 10:25 pm, Franz Allan Valencia See [franz....@gmail.com](mailto:franz....@gmail.com)  
> wrote:
> 
> > Thanks everyone.
> > 
> > Pardon Shay, let me elaborate 🙂
> > 
> > Currently, my problems are on  
> > a.) Getting the initial data from the main data source (MDS) to  
> > replicated  
> > data source (RDS) (which is Elasticsearch), and  
> > b.) Keeping the data in the RDS updated
> > 
> > For a.) these are my choices so far:  
> > a.1.) Query from the MDS everything that I will replicate and put them  
> > all  
> > in the RDS  
> > a.2.) Query from the MDS upon demand the information and put that in the  
> > RDS  
> > a.3.) Query from the MDS the heavy parts and put those in the RDS and  
> > then  
> > query the rest of the data on demand from the MDS and combine that with  
> > the  
> > data in the RDS, and then put the result back in the RDS (i.e. `original query is select * from tbl_small, tbl_big where ...`. So i first  
> > replicate  
> > `tbl_big`. and then when the user does an action, I will start querying  
> > from  
> > tbl\_small then combine that to the replicated tbl\_big and put the result  
> > in  
> > the query\_aggregate RDS so that future queries only hit the RDS)
> > 
> > For b.) these are my choices so far:  
> > b.1.) Have an event listener that when data in the MDS, the listener will  
> > update the data in the RDS (we're currently considering hibernate  
> > listener  
> > or db trigger)  
> > b.2.) Have a scheduled refresh of those that changed since last refresh
> > 
> > But since I don't know much about Replication & Caching  
> > Patterns/Anti-Patterns, I would like to inquire on how the community does  
> > it  
> > 🙂
> > 
> > Thoughts?
> > 
> > --  
> > Franz Allan Valencia See | Java Software Engineer  
> > [franz....@gmail.com](mailto:franz....@gmail.com)  
> > LinkedIn:[http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
> > Twitter:[http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)
> > 
> > On Fri, Jul 9, 2010 at 1:03 AM, Samuel Doyle [samueldo...@gmail.com](mailto:samueldo...@gmail.com)  
> > wrote:
> > 
> > > At the moment I'm using polling but since the main application is  
> > > making  
> > > use of Hibernate I may try to leverage a post flush interceptor for  
> > > incremental updating.  
> > > The initial main indexing is done via straight SQL.
> > 
> > > S.D.
> > 
> > > On Thu, Jul 8, 2010 at 5:40 AM, Franz Allan Valencia See \<  
> > > [franz....@gmail.com](mailto:franz....@gmail.com)\> wrote:
> > 
> > > > Curious, if you guys are using Elasticsearch for replication, how do  
> > > > you  
> > > > replicate your data?
> > 
> > > > I'm assuming your main data source is slow, so I was wondering how you  
> > > > go  
> > > > around it to have a faster replication.
> > 
> > > > Do you do your huge/slow query on your main data store and store that  
> > > > in  
> > > > Elasticsearch, or do you guys query expensive tables and store those  
> > > > in  
> > > > Elasticsearch, and then later recombine them? ...or? 🙂
> > 
> > > > Do you guys do lazy replication wherein the first time a query is  
> > > > made,  
> > > > you retrieve data from your main data source, then replicate in  
> > > > Elasticsearch? Or do you do a scheduled replication? ...or? 🙂
> > 
> > > > Thanks,
> > 
> > > > --  
> > > > Franz Allan Valencia See | Java Software Engineer  
> > > > [franz....@gmail.com](mailto:franz....@gmail.com)  
> > > > LinkedIn:[http://www.linkedin.com/in/franzsee](http://www.linkedin.com/in/franzsee)  
> > > > Twitter:[http://www.twitter.com/franz\_see](http://www.twitter.com/franz_see)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:22am UTC](https://discuss.elastic.co/t/replication-strategies/3081/7 "2017-07-06T04:22:27Z")

</div>


