# Nodes fail to join cluster after full cluster restart (cluster uuid mismatch?)

**URL:** <https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290>\
**Category:** Elasticsearch\
**Created:** [November 18, 2019, 10:56am UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290 "2019-11-18T10:56:17Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![tomhe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomhe/32/120065_2.png) [@tomhe](https://discuss.elastic.co/u/tomhe)\
**Post date:** [November 18, 2019, 10:56am UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/1 "2019-11-18T10:56:17Z")

</div>

I was forced to stop all nodes in our cluster and now I can't bring the cluster back up. Looks like the there is a cluster uuid mismatch, but I don't know why this has happened or how to fix it.

```
"type": "server", "timestamp": "2019-11-18T10:46:02,609Z", "level": "WARN", "component": "o.e.c.c.Coordinator", "cluster.name": "docker-cluster", "node.name": "node-002", "message": "failed to validate incoming join request from node [{node-012}{GrpvmVyVSOm2UpZQIUa3pg}{guMNX7HRT0q8-Lx7HL0Cnw}{10.33.9.82}{10.33.9.82:9300}{dil}{ml.machine_memory=67388260352, ml.max_open_jobs=20, xpack.installed=true}]", "cluster.uuid": "fQo4028sSN-QWcCaG2w_ZA", "node.id": "v9st6CCkQyioc6YFMbj3Mg" , 
"stacktrace": ["org.elasticsearch.transport.RemoteTransportException: [node-012][172.19.0.2:9300][internal:cluster/coordination/join/validate]",
"Caused by: org.elasticsearch.cluster.coordination.CoordinationStateRejectedException: join validation on cluster state with a different cluster uuid fQo4028sSN-QWcCaG2w_ZA than local cluster uuid i44zLmaER4ipQYj-F9QVDw, rejecting",

```

The data folder on each node is unchanged. What can I do to bring up the cluster again?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 18, 2019, 11:04am UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/2 "2019-11-18T11:04:24Z")

</div>

The cluster UUID is stored on disk on all master-eligible nodes and on all data nodes, and must match to prevent nodes from joining a different cluster since this is a good way to lose data. The usual way to get to this exception is to be using ephemeral storage on the master-eligible nodes. If the cluster UUID is missing on the master-eligible nodes then they will invent a new one, but this indicates that they have lost the cluster metadata too which means the data on your data nodes cannot be read correctly. If so, the safest way to proceed is to fix the storage on the master nodes to persist across restarts and then restore your data from a recent snapshot.

---

<div class="post-metadata">

**Author:** ![tomhe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomhe/32/120065_2.png) [@tomhe](https://discuss.elastic.co/u/tomhe)\
**Post date:** [November 18, 2019, 11:16am UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/3 "2019-11-18T11:16:01Z")

</div>

We’re using persistent storage on all nodes including the master-eligible nodes.

Is a full restore my only option? This is our production cluster and a full restore will take too long time.

Maybe also worth noting: I’ve done a successful full cluster restart previously.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 18, 2019, 11:20am UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/4 "2019-11-18T11:20:53Z")

</div>

What exact version are you using?

Can you grep all your logs on all nodes for `INFO` messages containing the string `cluster UUID` going back as far as possible, and share those logs here (or on [https://gist.github.com](https://gist.github.com) if they don't fit here).

---

<div class="post-metadata">

**Author:** ![tomhe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomhe/32/120065_2.png) [@tomhe](https://discuss.elastic.co/u/tomhe)\
**Post date:** [November 18, 2019, 12:29pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/5 "2019-11-18T12:29:24Z")

</div>

OK, so your initial guess was right: The data folder was missing.

During the time that the cluster was offline, a cron job (that I was unaware of, running `docker system prune --volumes --force`) removed the Docker volume containing `/usr/share/elasticsearch/data` on four of our nodes, including the eligible master nodes.

We're looking into the possibility of restoring /var/lib/docker (and hopefully the volumes) from backup. Would this be a bad idea? Will this leave our cluster in an inconsistent state? Four of the nodes would be using old data folders.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 18, 2019, 12:38pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/6 "2019-11-18T12:38:19Z")

</div>

Yep that'd do it.

It is risky to try and restore from a filesystem backup and I can't recommend it in good conscience, since it will take some of your nodes "back in time" and the effects of this are undefined. It may result in lost data (possibly silently) or may render some of your indices unreadable.

---

<div class="post-metadata">

**Author:** ![tomhe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomhe/32/120065_2.png) [@tomhe](https://discuss.elastic.co/u/tomhe)\
**Post date:** [November 18, 2019, 12:39pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/7 "2019-11-18T12:39:42Z")

</div>

We have nightly snapshots of our data. Is there a guide on how to restore a full cluster (including security data) from scratch from a snapshot?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 18, 2019, 1:51pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/8 "2019-11-18T13:51:19Z")

</div>

I don't know of anything more specific than [the restore docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-snapshots.html#restore-snapshot). You may need to disable some components (e.g. Kibana, monitoring, watcher, rollups, ...) for the duration of the restore since they may otherwise create indices that block the restore, and you'll need to use a security realm other than `native` since the native realm uses the `.security` index that you'll be restoring.

---

<div class="post-metadata">

**Author:** ![tomhe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomhe/32/120065_2.png) [@tomhe](https://discuss.elastic.co/u/tomhe)\
**Post date:** [November 18, 2019, 2:55pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/9 "2019-11-18T14:55:09Z")

</div>

Thanks!

---

<div class="post-metadata">

**Author:** ![code-chris](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/code-chris/32/26035_2.png) [@code-chris](https://discuss.elastic.co/u/code-chris)\
**Post date:** [November 25, 2019, 5:36pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/10 "2019-11-25T17:36:00Z")

</div>

In which path is this UUID and the cluster metadata stored? In the data-directory of the node?  
If not, then this means, that a full cluster restart would always fail in a K8s environment cause the filesystem of pods is always ephemeral (if not mounted by volumes)...

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 25, 2019, 7:06pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/11 "2019-11-25T19:06:11Z")

</div>

> [@code-chris](#):
>
> In which path is this UUID and the cluster metadata stored? In the data-directory of the node?

Yes.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 23, 2019, 7:08pm UTC](https://discuss.elastic.co/t/nodes-fail-to-join-cluster-after-full-cluster-restart-cluster-uuid-mismatch/208290/12 "2019-12-23T19:08:09Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
