# Elasticsearch on a wide scale around Globe

**URL:** <https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513>\
**Category:** Elasticsearch\
**Created:** [August 17, 2015, 2:56pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513 "2015-08-17T14:56:01Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Michel\_Laporte](https://avatars.discourse-cdn.com/v4/letter/m/3bc359/32.png) [@Michel\_Laporte](https://discuss.elastic.co/u/Michel_Laporte)\
**Post date:** [August 17, 2015, 2:56pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/1 "2015-08-17T14:56:02Z")

</div>

Hi,

We are due to deploy Elasticsearch (Along with Graylog) to our Office in London , New York, Seattle & San Francisco

We have a a server in each remote office (Graylog & Elasticsearch). However i have a small question / Problem.

We will be logging quite a lot of information. What i would like to do is:

Main Office : London  
Store all information in the cluster in Elasticsearch DB here

New York (Main US Office):  
Store all Northern American office Logs in Elasticsearch Database here (NY, Seattle & San Fran logs will be saved in the Elasticsearch DB

(Seattle, Sanfran will store logs ONLY from their office. So Network devices , Syslog messages in their respective office will be stored on their own Elasticsearch database only)

I dont want Seattle & San Fran to hold the whole Elasticsearch DB as the offices are only small and only about 10 devices will be logging to ES / Graylog . Whereas NYC and UK will have \> 100 devices/Servers logging to itself..

I want to be able to have UK and NYC holding ALL the information in the ES cluster and Sea and San fran to only hold their own logs but also send it to NYC so we have a backup. Is that possible?

I've seen Shard Allocation filtering but unsure on how to go ahead with it. Sorry if this is confusing it's been a project for \> 6 months and ideally want to roll it out in the next 30-40 days.

Thank you for all your help,  
Michel  
Junior Sys Admin

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 17, 2015, 10:25pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/2 "2015-08-17T22:25:20Z")

</div>

You _do not_ want to create a single cluster than spans all these sites. ES is latency sensitive and any networking issues would cause you dramallamas.

Best option if you want to do this is to use snapshot + restore to copy data around.

---

<div class="post-metadata">

**Author:** ![Michel\_Laporte](https://avatars.discourse-cdn.com/v4/letter/m/3bc359/32.png) [@Michel\_Laporte](https://discuss.elastic.co/u/Michel_Laporte)\
**Post date:** [August 18, 2015, 9:10am UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/3 "2015-08-18T09:10:49Z")

</div>

Okay.  
So would you recommend a cluster for Northern America and a Cluster for UK?

Or a different cluster for each office? (Even if SEA and San Fran will be fairly small)

Thanks

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 18, 2015, 9:57pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/4 "2015-08-18T21:57:26Z")

</div>

If you are happy shipping things over the wire then I'd have a cluster per continent, with indices split into per site and per source (ie network, system etc).

---

<div class="post-metadata">

**Author:** ![Michel\_Laporte](https://avatars.discourse-cdn.com/v4/letter/m/3bc359/32.png) [@Michel\_Laporte](https://discuss.elastic.co/u/Michel_Laporte)\
**Post date:** [August 19, 2015, 8:30am UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/5 "2015-08-19T08:30:33Z")

</div>

Okay thank you so much for clarifying this.

I will set up a multi cluster.

How do you split indices ?

---

<div class="post-metadata">

**Author:** ![ronchalant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ronchalant/32/3611_2.png) [@ronchalant](https://discuss.elastic.co/u/ronchalant)\
**Post date:** [August 19, 2015, 6:26pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/6 "2015-08-19T18:26:30Z")

</div>

I'm looking for similar functionality but for disaster recovery. We have two datacenters, one that houses our DR environment and another that houses our production. Our production environment runs robust servers (SSDs, etc.) while our DR environment would be more spartan with VMs mounting NFS shared (definitely not ideal, but it would allow us to operate our business).

Based on what I've been able to piece together, we'd want a separate "disaster recovery" cluster setup at our DR center that operates completely independently of production, and we'd want to sync data between them somehow.

Is there a way this can be automated on some kind of schedule? Is that something we'd have to do on our own with cron jobs or is there support for this sort of thing natively?

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [August 20, 2015, 10:11pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/7 "2015-08-20T22:11:03Z")

</div>

Hi,

You could feed the same data to multiple clusters (see Kafka, Kafka MirrorMaker) or maybe you could use ES snapshots if not being 100% up to date in secondary DCs is acceptable.

## Otis

Monitoring \* Alerting \* Anomaly Detection \* Centralized Log Management  
Solr & Elasticsearch Support \* [http://sematext.com/](http://sematext.com/)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 21, 2015, 8:06am UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/8 "2015-08-21T08:06:41Z")

</div>

Just name them differently for each data type/source 🙂

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 21, 2015, 8:07am UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/9 "2015-08-21T08:07:22Z")

</div>

What Otis mentioned! There's also a bunch of other threads with similar questions that may provide other info.

---

<div class="post-metadata">

**Author:** ![Michel\_Laporte](https://avatars.discourse-cdn.com/v4/letter/m/3bc359/32.png) [@Michel\_Laporte](https://discuss.elastic.co/u/Michel_Laporte)\
**Post date:** [August 21, 2015, 9:17am UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/10 "2015-08-21T09:17:12Z")

</div>

Thanks for your help Mark 🙂

---

<div class="post-metadata">

**Author:** ![ronchalant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ronchalant/32/3611_2.png) [@ronchalant](https://discuss.elastic.co/u/ronchalant)\
**Post date:** [August 21, 2015, 2:48pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/11 "2015-08-21T14:48:06Z")

</div>

> [@otisg](#):
>
> Kafka MirrorMaker

Thanks @otisg, I'll look into that!

---

<div class="post-metadata">

**Author:** ![ronchalant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ronchalant/32/3611_2.png) [@ronchalant](https://discuss.elastic.co/u/ronchalant)\
**Post date:** [August 21, 2015, 3:04pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/12 "2015-08-21T15:04:54Z")

</div>

> [@warkolm](#):
>
> What Otis mentioned! There's also a bunch of other threads with similar questions that may provide other info.

yeah I've been looking around but there doesn't seem to be an "official" way to do it .. most seem to point to doing a snapshot from prod and restore in your "mirrored" environment, or if you need it live to do a "distributed indexing" or something where whenever you index in one cluster you distribute that indexing request to multiple clusters.

We'll probably end up doing something where we regularly snapshot the prod index to a network file share, then periodically restore a full snapshot to DR just to keep it "close" to production. In our case we don't _need_ the DR index to be always 100% caught up with production, just relatively close.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 21, 2015, 10:33pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/13 "2015-08-21T22:33:30Z")

</div>

As with most things ES, it depends on your use case and requirements 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:54pm UTC](https://discuss.elastic.co/t/elasticsearch-on-a-wide-scale-around-globe/27513/14 "2017-07-05T23:54:27Z")

</div>


