# Geo locating shards question

**URL:** <https://discuss.elastic.co/t/geo-locating-shards-question/10518>\
**Category:** Elasticsearch\
**Created:** [January 27, 2013, 12:42pm UTC](https://discuss.elastic.co/t/geo-locating-shards-question/10518 "2013-01-27T12:42:43Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![james\_lewis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james_lewis/32/2296_2.png) [@james\_lewis](https://discuss.elastic.co/u/james_lewis)\
**Post date:** [January 27, 2013, 12:42pm UTC](https://discuss.elastic.co/t/geo-locating-shards-question/10518/1 "2013-01-27T12:42:43Z")

</div>

Just a quick questions:

In the scenario discussed in this RavenDB guide[http://ravendb.net/docs/server/scaling-out/sharding](http://ravendb.net/docs/server/scaling-out/sharding), data  
from different companies from across the globe is to be stored in  
differently located shards based on the region of the company. So if  
company A is based in Asia and company B is based in the UK, all of company  
A's data would be indexed into a shard located in Asia and all of company  
B's data would be indexed into a shard located in the UK.

My questions is: is this geo locating of shards possible without creating a  
different index for each company in elasticsearch? And if it is, is it  
something that would need to be thought about before going live with a  
solution? What I mean by that is: if you weren't bothered about  
geolocating data when you initially went live, would it be possible to  
introduce a solution later when required?

Regards,  
James

--

---

<div class="post-metadata">

**Author:** ![Itamar\_Syn\_Hershko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/itamar_syn_hershko/32/725_2.png) [@Itamar\_Syn\_Hershko](https://discuss.elastic.co/u/Itamar_Syn_Hershko)\
**Post date:** [January 27, 2013, 1:42pm UTC](https://discuss.elastic.co/t/geo-locating-shards-question/10518/2 "2013-01-27T13:42:22Z")

</div>

The design decisions behind sharding is quite different between RavenDB and  
Elasticsearch. To name one such difference, ES does this on the cluster,  
while sharding in RavenDB is powered by the client itself, and the server  
is not at all aware it is just a shard.

In Elasticsearch the best approach would probably be to have different  
indexes per _region_ and have it sharded based on actual load (probably  
reserving some virtual shards). You could then control which nodes it will  
be deployed on using the include/exclude  
tags[http://www.elasticsearch.org/guide/reference/index-modules/allocation.html](http://www.elasticsearch.org/guide/reference/index-modules/allocation.html).  
Since the cluster is split geographically, I wonder if it makes sense to  
still have them in one ES cluster. Only useful if you are going to query  
and aggregate results across regions.

One important benefit for creating an index per region will be the ability  
to worry about scaling each region independently, instead of having one  
huge index, geographically distributed, for everything without the ability  
to re-shard.

You always have to properly plan sharding with Elasticsearch, and that goes  
for RavenDB as well (where you'd have to create a good sharding function).  
In your scenario, you can start with one huge index and reindex to  
different geographically distributed indexes at a later time. Reindexing  
with ES is easy enough, and may the \_source be with you 🙂

On Sun, Jan 27, 2013 at 2:42 PM, [james.lewis@7digital.com](mailto:james.lewis@7digital.com) wrote:

> Just a quick questions:
> 
> In the scenario discussed in this RavenDB guide[http://ravendb.net/docs/server/scaling-out/sharding](http://ravendb.net/docs/server/scaling-out/sharding), data  
> from different companies from across the globe is to be stored in  
> differently located shards based on the region of the company. So if  
> company A is based in Asia and company B is based in the UK, all of company  
> A's data would be indexed into a shard located in Asia and all of company  
> B's data would be indexed into a shard located in the UK.
> 
> My questions is: is this geo locating of shards possible without creating  
> a different index for each company in elasticsearch? And if it is, is it  
> something that would need to be thought about before going live with a  
> solution? What I mean by that is: if you weren't bothered about  
> geolocating data when you initially went live, would it be possible to  
> introduce a solution later when required?
> 
> Regards,  
> James
> 
> --

---

<div class="post-metadata">

**Author:** ![james\_lewis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james_lewis/32/2296_2.png) [@james\_lewis](https://discuss.elastic.co/u/james_lewis)\
**Post date:** [January 27, 2013, 1:48pm UTC](https://discuss.elastic.co/t/geo-locating-shards-question/10518/3 "2013-01-27T13:48:56Z")

</div>

That was my initial plan, to start with my one index and then reindex later  
on if I need to (I'm using aliasing anyway).

I did just have a thought though - if I were more concerned with first time  
to byte on search performance then I wouldn't care which continent I was  
indexing to, I would just need to make sure that there was a replica of the  
data in the continent the search request was made. So if someone searches  
from the US their request would get routed to a US node containing a  
replica of that data (but it would have been indexed on a node hosted in  
the UK for example).

Thanks a lot for confirming the differences between RavenDB and ES - I was  
envisioning a similar sharding function to be required in elasticsearch but  
obviously not!

Thanks,  
James

On Sun, Jan 27, 2013 at 1:42 PM, Itamar Syn-Hershko [itamar@code972.com](mailto:itamar@code972.com)wrote:

> The design decisions behind sharding is quite different between RavenDB  
> and Elasticsearch. To name one such difference, ES does this on the  
> cluster, while sharding in RavenDB is powered by the client itself, and the  
> server is not at all aware it is just a shard.
> 
> In Elasticsearch the best approach would probably be to have different  
> indexes per _region_ and have it sharded based on actual load (probably  
> reserving some virtual shards). You could then control which nodes it will  
> be deployed on using the include/exclude tags[http://www.elasticsearch.org/guide/reference/index-modules/allocation.html](http://www.elasticsearch.org/guide/reference/index-modules/allocation.html).  
> Since the cluster is split geographically, I wonder if it makes sense to  
> still have them in one ES cluster. Only useful if you are going to query  
> and aggregate results across regions.
> 
> One important benefit for creating an index per region will be the ability  
> to worry about scaling each region independently, instead of having one  
> huge index, geographically distributed, for everything without the ability  
> to re-shard.
> 
> You always have to properly plan sharding with Elasticsearch, and that  
> goes for RavenDB as well (where you'd have to create a good sharding  
> function). In your scenario, you can start with one huge index and reindex  
> to different geographically distributed indexes at a later time. Reindexing  
> with ES is easy enough, and may the \_source be with you 🙂
> 
> On Sun, Jan 27, 2013 at 2:42 PM, [james.lewis@7digital.com](mailto:james.lewis@7digital.com) wrote:
> 
> > Just a quick questions:
> > 
> > In the scenario discussed in this RavenDB guide[http://ravendb.net/docs/server/scaling-out/sharding](http://ravendb.net/docs/server/scaling-out/sharding), data  
> > from different companies from across the globe is to be stored in  
> > differently located shards based on the region of the company. So if  
> > company A is based in Asia and company B is based in the UK, all of company  
> > A's data would be indexed into a shard located in Asia and all of company  
> > B's data would be indexed into a shard located in the UK.
> > 
> > My questions is: is this geo locating of shards possible without creating  
> > a different index for each company in elasticsearch? And if it is, is it  
> > something that would need to be thought about before going live with a  
> > solution? What I mean by that is: if you weren't bothered about  
> > geolocating data when you initially went live, would it be possible to  
> > introduce a solution later when required?
> > 
> > Regards,  
> > James
> > 
> > --
> 
> --

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:54am UTC](https://discuss.elastic.co/t/geo-locating-shards-question/10518/4 "2017-07-06T02:54:26Z")

</div>


