# Multi-cluster installation, cluster loss and recovery

**URL:** <https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858>\
**Category:** Elasticsearch\
**Created:** [July 20, 2023, 8:20am UTC](https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858 "2023-07-20T08:20:01Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Christophe\_DAME](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christophe_dame/32/123684_2.png) [@Christophe\_DAME](https://discuss.elastic.co/u/Christophe_DAME)\
**Post date:** [July 20, 2023, 8:20am UTC](https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858/1 "2023-07-20T08:20:01Z")

</div>

Hi there,

As I'm new, I'll introduce myself quickly : I'm Christophe Dame and I'm working for Camunda. Our solution embarks an Elasticsearch 7.17 and I'm currently experimenting a dual cluster setup.  
Ideally, my goal would be to have an Elasticsearch spanning over 2 clusters. My first naive attempt is to install 2 ES nodes in cluster 1 and 2 ES nodes in cluster2 (using helm charts). I've added nodeGroup to each of them. I've added seed\_hosts to each pointing to the headless services. And only the first cluster I starts as the initial\_master\_nodes set. I forced an empty value in the second cluster.  
Result is good as I then have a 4 nodes healthy cluster. I run some operations that create indices and data is available in my web applications reading locally in both clusters. Youpi!  
Then I destroy a cluster and recreate it... and the recreated nodes wont join. I get such an error : ["org.elasticsearch.cluster.block.ClusterBlockException: blocked by: [SERVICE\_UNAVAILABLE/1/state not recovered / initialized];

Do you have any advice ? Something I'm doing wrong or that I missed ?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 20, 2023, 8:53am UTC](https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858/2 "2023-07-20T08:53:35Z")

</div>

> [@Christophe\_DAME](#):
>
> blocked by: [SERVICE\_UNAVAILABLE/1/state not recovered / initialized];

That looks like a cluster formation problem, see [these docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/discovery-troubleshooting.html) for information about how to troubleshoot it effectively.

If I had to guess, I think you probably destroyed too many master nodes at once. See [these docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/high-availability.html) for information about setting up a resilient cluster, particularly the [section on two-zone clusters](https://www.elastic.co/guide/en/elasticsearch/reference/current/high-availability-cluster-design-large-clusters.html#high-availability-cluster-design-two-zones):

> You cannot configure a two-zone cluster so that it can tolerate the loss of _either_ zone because this is theoretically impossible.

(terminology note, the whole Elasticsearch installation is called a "cluster" in these docs, and what you are calling a "cluster" is closer to a "zone")

---

<div class="post-metadata">

**Author:** ![Christophe\_DAME](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christophe_dame/32/123684_2.png) [@Christophe\_DAME](https://discuss.elastic.co/u/Christophe_DAME)\
**Post date:** [July 20, 2023, 9:04am UTC](https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858/3 "2023-07-20T09:04:47Z")

</div>

Hi David,

Many thanks for your quick response. You're right about the impossibility to survive the loss of half nodes. But my assumption would be that the data is replicated across the nodes and that I should not loose any data.  
By recreating empty nodes (named as the lost ones), I sould be able to redistribute the data and have my cluster healthy again, or ?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 20, 2023, 9:13am UTC](https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858/4 "2023-07-20T09:13:06Z")

</div>

> By recreating empty nodes (named as the lost ones), I sould be able to redistribute the data and have my cluster healthy again, or ?

No, that's not safe. As the docs say, what you are trying to do is not even theoretically possible.

---

<div class="post-metadata">

**Author:** ![Christophe\_DAME](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christophe_dame/32/123684_2.png) [@Christophe\_DAME](https://discuss.elastic.co/u/Christophe_DAME)\
**Post date:** [July 20, 2023, 10:00am UTC](https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858/5 "2023-07-20T10:00:13Z")

</div>

Thanks David, I think I get your point. I'll try to think about a workaround for that particular case. Thanks again 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 17, 2023, 10:00am UTC](https://discuss.elastic.co/t/multi-cluster-installation-cluster-loss-and-recovery/338858/6 "2023-08-17T10:00:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
