# Data loss when old master node dead and startup again

**URL:** <https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506>\
**Category:** Elasticsearch\
**Created:** [February 22, 2019, 1:43am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506 "2019-02-22T01:43:54Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 22, 2019, 1:43am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/1 "2019-02-22T01:43:54Z")

</div>

**Describe the feature** :

**Elasticsearch version** (`bin/elasticsearch --version`): elasticsearch/elasticsearch-oss:6.3.2

**docker version** : 18.03.1-ce

**OS version** (`uname -a` if on a Unix-like system): centos7.5

**Description of the problem including expected versus actual behavior** :

**Steps to reproduce** :

1. we have two node cluster, A (master), B(slave)
2. A node dead, and B become new master (standalone)
3. Post new data to B
4. A startup, and join cluster, B is still as a master, But new data lost, all data become old.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 22, 2019, 2:29am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/2 "2019-02-22T02:29:43Z")

</div>

Are you using an external volume to store the Elasticsearch data directory in?

---

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 22, 2019, 2:34am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/3 "2019-02-22T02:34:04Z")

</div>

Yes, we store data on external volume.  
The question is new data lost, seems be recoveried.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 22, 2019, 3:04am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/4 "2019-02-22T03:04:49Z")

</div>

> [@jffree](#):
>
> A node dead, and B become new master (standalone)

This means you have not set [`discovery.zen.minimum_master_nodes`](https://www.elastic.co/guide/en/elasticsearch/reference/6.3/discovery-settings.html#minimum_master_nodes) correctly. It must be 2 on each node. As the manual says:

> To prevent data loss, it is vital to configure the `discovery.zen.minimum_master_nodes` setting

---

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 22, 2019, 3:08am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/5 "2019-02-22T03:08:08Z")

</div>

When A dead, we have delete `discovery.zen.ping.unicast.hosts` config in `elasticsearch.yml` in B node, so B node will run as standalone mode.

---

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 22, 2019, 5:20am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/6 "2019-02-22T05:20:49Z")

</div>

I have already set `discovery.zen.minimum_master_nodes:1` in `elasticsearch.yml`, but it did not work.

---

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 22, 2019, 5:22am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/7 "2019-02-22T05:22:57Z")

</div>

If I set `discovery.zen.minimum_master_nodes:2`, if one node dead, so the cluster will not work. I do not want this happened.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 22, 2019, 7:02am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/8 "2019-02-22T07:02:42Z")

</div>

Elasticsearch operates in a clustered mode, not master-slave. If you want a highly available cluster you therefore need a minimum of three master-eligible nodes.

---

<div class="post-metadata">

**Author:** ![vigyas](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vigyas/32/40226_2.png) [@vigyas](https://discuss.elastic.co/u/vigyas)\
**Post date:** [February 22, 2019, 10:41am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/9 "2019-02-22T10:41:36Z")

</div>

While setting `minimum_master_nodes` to 2 will avoid this, curious to know if this expected behavior or a replication edge case?

Quoting steps to repro from [issues/39282](https://github.com/elastic/elasticsearch/issues/39282)

> 1. I have two nodes in cluster (A and B)
> 2. old cluster A is master, I set `node.master: true` in `elasticsearch.yml` , B is set `node.master: false` in  
> `elasticsearch.yml`
> 3. Stop A, update B config( remove `node.master: false` in `elasticsearch.yml` ), so B can run standalone
> 4. POST new data to B node
> 5. Startup A node and (set `node.master: false` in `elasticsearch.yml` ), B is set `node.master: true` in  
> `elasticsearch.yml` , so B is master in new cluster
> 6. But the new data loss!

This does not look like a split brain, at a time there was only one master. First A is master, then it is stopped and node B is made master (single node cluster), then A is made _master ineligible_ and added back to cluster. But it seems that node A's shards override node B's shards?

- Could this be an issue around allotting [primary terms](https://github.com/elastic/elasticsearch/pull/14062) for shards when master B promoted its replica to primary?

- How are the cluster state details from node A and node B reconciled when A joins back?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 22, 2019, 2:05pm UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/10 "2019-02-22T14:05:16Z")

</div>

The data loss starts here:

1. Stop A, update B config( remove `node.master: false` in `elasticsearch.yml` ), so B can run standalone

Since B was not a master-eligible node it doesn't have a full copy of the cluster metadata on disk, so it starts up empty. It will have _some_ index metadata, but maybe not all of it, and what it has could also be stale. It imports any indices it finds as dangling indices, and blindly trusts the corresponding index metadata even though this could be stale (and that includes primary terms). Re-using a primary term like this breaks all sorts of assumptions on which we rely, so from that point on the behaviour of Elasticsearch is undefined.

---

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 22, 2019, 2:28pm UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/11 "2019-02-22T14:28:11Z")

</div>

If I delete `node.master` setting and set `minimum_master_nodes` 1 in both A and B. Then I try the same steps, it will also loss data.  
But if I set `minimum_master_nodes` 2 when A and B both alive, and set `minimum_master_nodes` 1 when only one node alive, it will works.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 22, 2019, 2:47pm UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/12 "2019-02-22T14:47:40Z")

</div>

If you want high-availability and avoid data loss you need at least three master-eligible nodes in your cluster. Two is not sufficient.

---

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 22, 2019, 3:08pm UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/13 "2019-02-22T15:08:43Z")

</div>

Because there is only two nodes in our production, so I have to try every method to avoid data loss and provide high-availability service.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 22, 2019, 3:19pm UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/14 "2019-02-22T15:19:52Z")

</div>

I would recommend trying to add a small dedicated master node somewhere. This does not require a lot of resources and would give you three master-eligible nodes even if it does not hold data.

---

<div class="post-metadata">

**Author:** ![jffree](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jffree/32/41252_2.png) [@jffree](https://discuss.elastic.co/u/jffree)\
**Post date:** [February 25, 2019, 2:13am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/15 "2019-02-25T02:13:14Z")

</div>

I want to avoid one node dead and es can not support service. This is where the problem lies.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 25, 2019, 7:08am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/16 "2019-02-25T07:08:55Z")

</div>

If you want to handle that automatically without manual intervention or risk of data loss that requires a minimum of 3 master eligible nodes. There is no way around this.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 25, 2019, 7:18am UTC](https://discuss.elastic.co/t/data-loss-when-old-master-node-dead-and-startup-again/169506/17 "2019-03-25T07:18:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
