# Dealing with latency when indexing

**URL:** <https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146>\
**Category:** Elasticsearch\
**Created:** [March 14, 2013, 2:43am UTC](https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146 "2013-03-14T02:43:49Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dustin\_Lashmar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dustin_lashmar/32/2594_2.png) [@Dustin\_Lashmar](https://discuss.elastic.co/u/Dustin_Lashmar)\
**Post date:** [March 14, 2013, 2:43am UTC](https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146/1 "2013-03-14T02:43:49Z")

</div>

Hi all,

I'm investigating setting up an Elasticsearch cluster that spans multiple  
regions (possibly ec2 regions, but possibly not), and I'm anticipating a  
fair bit of latency between them.  
I think it makes sense to use  
cluster.routing.allocation.awareness.attributes and  
cluster.routing.allocation.awareness.force.region.values and setting the  
number of replicas = number of regions - 1. That way there will be a full  
copy of the data in each region (right?)  
I'm also configuring the client nodes with  
cluster.routing.allocation.awareness.attributes so they should always hit  
nodes in the same region if they are available.  
This is awesome because it means full fail-over if a region goes down, and  
also means that searches won't have to leave the region they started from,  
avoiding the latency (unless nodes in that region go down, but that's ok).

The only issue is when it comes to indexing documents, my understanding  
(and correct me if I'm wrong) is that docs being indexed will need to first  
be indexed on the primary shard, then go to the replicas. So if the primary  
is not in your region it will take at least 2\*latency before you see the  
document in your region.

So is there a way to make sure that each region always has at least one  
primary shard? And to route new documents to that particular shard? I  
thought  
[http://www.elasticsearch.org/guide/reference/api/admin-cluster-reroute.html](http://www.elasticsearch.org/guide/reference/api/admin-cluster-reroute.html)  
might help, but as far as I can tell I can't use it to switch a shard to  
primary status.

Or perhaps my whole approach is wrong, does anyone have strategies for  
dealing with high latency between sections of a cluster?

Thanks in advance,

Dustin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![andrassy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrassy/32/32783_2.png) [@andrassy](https://discuss.elastic.co/u/andrassy)\
**Post date:** [August 30, 2013, 9:19am UTC](https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146/2 "2013-08-30T09:19:28Z")

</div>

Hi Dustin,

Did you get very far on this or find any useful info elsewhere? We're  
building out a similar multi-site cluster.

Thanks,

Neil

On Thursday, 14 March 2013 02:43:49 UTC, Dustin Lashmar wrote:

> Hi all,
> 
> I'm investigating setting up an Elasticsearch cluster that spans multiple  
> regions (possibly ec2 regions, but possibly not), and I'm anticipating a  
> fair bit of latency between them.  
> I think it makes sense to use  
> cluster.routing.allocation.awareness.attributes and  
> cluster.routing.allocation.awareness.force.region.values and setting the  
> number of replicas = number of regions - 1. That way there will be a full  
> copy of the data in each region (right?)  
> I'm also configuring the client nodes with  
> cluster.routing.allocation.awareness.attributes so they should always hit  
> nodes in the same region if they are available.  
> This is awesome because it means full fail-over if a region goes down, and  
> also means that searches won't have to leave the region they started from,  
> avoiding the latency (unless nodes in that region go down, but that's ok).
> 
> The only issue is when it comes to indexing documents, my understanding  
> (and correct me if I'm wrong) is that docs being indexed will need to first  
> be indexed on the primary shard, then go to the replicas. So if the primary  
> is not in your region it will take at least 2\*latency before you see the  
> document in your region.
> 
> So is there a way to make sure that each region always has at least one  
> primary shard? And to route new documents to that particular shard? I  
> thought  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-cluster-reroute.htmlmight) help, but as far as I can tell I can't use it to switch a shard to  
> primary status.
> 
> Or perhaps my whole approach is wrong, does anyone have strategies for  
> dealing with high latency between sections of a cluster?
> 
> Thanks in advance,
> 
> Dustin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Dustin\_Lashmar\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dustin_lashmar_2/32/2123_2.png) [@Dustin\_Lashmar\_2](https://discuss.elastic.co/u/Dustin_Lashmar_2)\
**Post date:** [September 1, 2013, 11:24pm UTC](https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146/3 "2013-09-01T23:24:40Z")

</div>

Hi Neil,

I didn't get much further with this as we found a single region was enough  
for now. From what I've read though most people recommend against using a  
single cluster that spans multiple regions and instead setting up a cluster  
per region and using other techniques to keep them in synch. How exactly  
you'd go about that I don't really know, I never really got that far with  
it.

Sorry I couldn't help more, let me know if you find a nice solution,

Dustin

On Friday, August 30, 2013 7:19:28 PM UTC+10, Neil Andrassy wrote:

> Hi Dustin,
> 
> Did you get very far on this or find any useful info elsewhere? We're  
> building out a similar multi-site cluster.
> 
> Thanks,
> 
> Neil
> 
> On Thursday, 14 March 2013 02:43:49 UTC, Dustin Lashmar wrote:
> 
> > Hi all,
> > 
> > I'm investigating setting up an Elasticsearch cluster that spans multiple  
> > regions (possibly ec2 regions, but possibly not), and I'm anticipating a  
> > fair bit of latency between them.  
> > I think it makes sense to use  
> > cluster.routing.allocation.awareness.attributes and  
> > cluster.routing.allocation.awareness.force.region.values and setting the  
> > number of replicas = number of regions - 1. That way there will be a full  
> > copy of the data in each region (right?)  
> > I'm also configuring the client nodes with  
> > cluster.routing.allocation.awareness.attributes so they should always hit  
> > nodes in the same region if they are available.  
> > This is awesome because it means full fail-over if a region goes down,  
> > and also means that searches won't have to leave the region they started  
> > from, avoiding the latency (unless nodes in that region go down, but that's  
> > ok).
> > 
> > The only issue is when it comes to indexing documents, my understanding  
> > (and correct me if I'm wrong) is that docs being indexed will need to first  
> > be indexed on the primary shard, then go to the replicas. So if the primary  
> > is not in your region it will take at least 2\*latency before you see the  
> > document in your region.
> > 
> > So is there a way to make sure that each region always has at least one  
> > primary shard? And to route new documents to that particular shard? I  
> > thought  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-cluster-reroute.htmlmight) help, but as far as I can tell I can't use it to switch a shard to  
> > primary status.
> > 
> > Or perhaps my whole approach is wrong, does anyone have strategies for  
> > dealing with high latency between sections of a cluster?
> > 
> > Thanks in advance,
> > 
> > Dustin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [September 3, 2013, 8:23pm UTC](https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146/4 "2013-09-03T20:23:58Z")

</div>

I suspect that the immediate answer is to implement two independent  
clusters, one in each region.

Then create a durable queue in each cluster. Make sure the backing file  
store for that queue is properly replicated (for example, iSCSI or NetApp  
or something similar).

Then each consumer is a Java application that consumes a request from the  
queue and sends it to the other cluster via the TransportClient class.  
Defining the TransportClient with all of the addresses of the nodes in the  
remote cluster will handle the needed failover.

Then ensure that the application can tolerate the intra-cluster latency.  
For example, if the link between regions is broken, the clusters can  
operate independently and queue updates to the other. The application must  
be aware of this and treat it as a normal part of its operation. This is  
easier said than done, but for a database with high update rates and that  
is split across two faraway regions, this is the simplest option which  
means it's the easiest to create and test.

On the other hand, putting the database into one highly available cloud  
(for example, Amazon) is a single-region solution that follows Mark Twain's  
advice to "put all of your eggs in one basket, and then watch that  
basket!". Works very well most of the time, but network or other service  
outages must be factored in, such as the recent Amazon AWS outage.

On Sunday, September 1, 2013 7:24:40 PM UTC-4, Dustin Lashmar wrote:

> Hi Neil,
> 
> I didn't get much further with this as we found a single region was enough  
> for now. From what I've read though most people recommend against using a  
> single cluster that spans multiple regions and instead setting up a cluster  
> per region and using other techniques to keep them in synch. How exactly  
> you'd go about that I don't really know, I never really got that far with  
> it.
> 
> Sorry I couldn't help more, let me know if you find a nice solution,
> 
> Dustin
> 
> On Friday, August 30, 2013 7:19:28 PM UTC+10, Neil Andrassy wrote:
> 
> > Hi Dustin,
> > 
> > Did you get very far on this or find any useful info elsewhere? We're  
> > building out a similar multi-site cluster.
> > 
> > Thanks,
> > 
> > Neil

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 4, 2013, 10:31am UTC](https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146/5 "2013-09-04T10:31:33Z")

</div>

A simple mirror cluster indexing technique is firing up two (or more)  
TransportClients and indexing to the clusters in sync.

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:18am UTC](https://discuss.elastic.co/t/dealing-with-latency-when-indexing/11146/6 "2017-07-06T02:18:18Z")

</div>


