# Multi-datacenter deployments

**URL:** <https://discuss.elastic.co/t/multi-datacenter-deployments/5039>\
**Category:** Elasticsearch\
**Created:** [August 2, 2011, 10:33pm UTC](https://discuss.elastic.co/t/multi-datacenter-deployments/5039 "2011-08-02T22:33:51Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ask\_Bjorn\_Hansen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ask_bjorn_hansen/32/2995_2.png) [@Ask\_Bjorn\_Hansen](https://discuss.elastic.co/u/Ask_Bjorn_Hansen)\
**Post date:** [August 2, 2011, 10:33pm UTC](https://discuss.elastic.co/t/multi-datacenter-deployments/5039/1 "2011-08-02T22:33:51Z")

</div>

Hi everyone,

We're working on making one of our ES using applications run in multiple  
data centers (active/passive for now).

The other data stores we have can replicate seamlessly and have some idea of  
"local" vs "remote"; so we'd like to not run the ES instances as completely  
separate.

A minimal-ish TODO list I could think of is:

- Being able to discover nodes across networks (I think that works already  
with a bit of configuration, no?)
- Have ES know to have at least one copy of each shard in each datacenter.
- When copying shards on startup then pull it from a local node if possible.
- When doing queries, prefer shards that are local.
- Have the client prefer local servers.

Is any of this on the roadmap? I think we'd be interested in helping  
sponsor this work if possible.

- ask

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 6, 2011, 5:30pm UTC](https://discuss.elastic.co/t/multi-datacenter-deployments/5039/2 "2011-08-06T17:30:44Z")

</div>

Hi,

Smarter / more controllable shard allocation is on the roadmap, where you  
could control shard allocation in a similar manner that you described, but  
it will only work properly for DC that have a fast connection between them.  
Otherwise, a different approach is needed, where a write ahead log based  
replication will be needed between two different es clusters.

On Wed, Aug 3, 2011 at 1:33 AM, Ask Bjørn Hansen [ask@develooper.com](mailto:ask@develooper.com) wrote:

> Hi everyone,
> 
> We're working on making one of our ES using applications run in multiple  
> data centers (active/passive for now).
> 
> The other data stores we have can replicate seamlessly and have some idea  
> of "local" vs "remote"; so we'd like to not run the ES instances as  
> completely separate.
> 
> A minimal-ish TODO list I could think of is:
> 
> - Being able to discover nodes across networks (I think that works already  
> with a bit of configuration, no?)
> - Have ES know to have at least one copy of each shard in each datacenter.
> - When copying shards on startup then pull it from a local node if  
> possible.
> - When doing queries, prefer shards that are local.
> - Have the client prefer local servers.
> 
> Is any of this on the roadmap? I think we'd be interested in helping  
> sponsor this work if possible.
> 
> - ask

---

<div class="post-metadata">

**Author:** ![loren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loren/32/44942_2.png) [@loren](https://discuss.elastic.co/u/loren)\
**Post date:** [July 9, 2012, 9:42pm UTC](https://discuss.elastic.co/t/multi-datacenter-deployments/5039/3 "2012-07-09T21:42:44Z")

</div>

I am looking to do something very similar here with two datacenters that  
are about 20ms apart. All of the indexing happens in one DC, so I would  
prefer all the primary shards to be in one DC and all the replication to  
happen across the VPN to the second DC. Otherwise I could be sending data  
across the interconnect twice.

It sounds like right now the WAL replication functionality is not yet  
implemented, so this sort of master-slave cluster replication is not (yet)  
available. Is that correct?

---

<div class="post-metadata">

**Author:** ![AEvar\_Arnfjord\_Bjarm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aevar_arnfjord_bjarm/32/2746_2.png) [@AEvar\_Arnfjord\_Bjarm](https://discuss.elastic.co/u/AEvar_Arnfjord_Bjarm)\
**Post date:** [July 9, 2012, 10:24pm UTC](https://discuss.elastic.co/t/multi-datacenter-deployments/5039/4 "2012-07-09T22:24:58Z")

</div>

I think you can do what you want with the cluster allocation API:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

But it doesn't look very robust, e.g. you can specify that certain  
nodes shouldn't have primary shards, but if a primary is missing a  
replica (in your other DC) might still be promoted to the primary.

Maybe someone here has more experience with this.

On Mon, Jul 9, 2012 at 11:42 PM, Loren [loren@siebert.org](mailto:loren@siebert.org) wrote:

> I am looking to do something very similar here with two datacenters that are  
> about 20ms apart. All of the indexing happens in one DC, so I would prefer  
> all the primary shards to be in one DC and all the replication to happen  
> across the VPN to the second DC. Otherwise I could be sending data across  
> the interconnect twice.
> 
> It sounds like right now the WAL replication functionality is not yet  
> implemented, so this sort of master-slave cluster replication is not (yet)  
> available. Is that correct?

---

<div class="post-metadata">

**Author:** ![loren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loren/32/44942_2.png) [@loren](https://discuss.elastic.co/u/loren)\
**Post date:** [July 9, 2012, 11:13pm UTC](https://discuss.elastic.co/t/multi-datacenter-deployments/5039/5 "2012-07-09T23:13:32Z")

</div>

I did notice the cluster API and some of the other posts on this multi-datacenter topic.

As I just want the primary copies to be in the DC1 zone before the main indexing begins, I was thinking I could disable the nodes in the DC2 zone long enough for any replica DC1 nodes to be promoted to primary, and then re-enable DC2 so that they become replicas. Seems like there might be some sort of way to do this programmatically, but I am still learning about ES and am not sure what’s there, what’s not, and what’s in the development pipeline.

On Jul 9, 2012, at 3:24 PM, Ævar Arnfjörð Bjarmason wrote:

> I think you can do what you want with the cluster allocation API:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/cluster.html)
> 
> But it doesn't look very robust, e.g. you can specify that certain  
> nodes shouldn't have primary shards, but if a primary is missing a  
> replica (in your other DC) might still be promoted to the primary.
> 
> Maybe someone here has more experience with this.
> 
> On Mon, Jul 9, 2012 at 11:42 PM, Loren [loren@siebert.org](mailto:loren@siebert.org) wrote:
> 
> > I am looking to do something very similar here with two datacenters that are  
> > about 20ms apart. All of the indexing happens in one DC, so I would prefer  
> > all the primary shards to be in one DC and all the replication to happen  
> > across the VPN to the second DC. Otherwise I could be sending data across  
> > the interconnect twice.
> > 
> > It sounds like right now the WAL replication functionality is not yet  
> > implemented, so this sort of master-slave cluster replication is not (yet)  
> > available. Is that correct?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:20am UTC](https://discuss.elastic.co/t/multi-datacenter-deployments/5039/6 "2017-07-06T03:20:51Z")

</div>


