# Elasticsearch in K8s cannot restart because of recovery failure (multiple alias write indexes)

**URL:** <https://discuss.elastic.co/t/elasticsearch-in-k8s-cannot-restart-because-of-recovery-failure-multiple-alias-write-indexes/214100>\
**Category:** Elasticsearch\
**Created:** [January 7, 2020, 4:37pm UTC](https://discuss.elastic.co/t/elasticsearch-in-k8s-cannot-restart-because-of-recovery-failure-multiple-alias-write-indexes/214100 "2020-01-07T16:37:22Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![MF57](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mf57/32/72394_2.png) [@MF57](https://discuss.elastic.co/u/MF57)\
**Post date:** [January 7, 2020, 4:37pm UTC](https://discuss.elastic.co/t/elasticsearch-in-k8s-cannot-restart-because-of-recovery-failure-multiple-alias-write-indexes/214100/1 "2020-01-07T16:37:22Z")

</div>

Hello,

I have an elastic search cluster deployed in the k8s cluster:

```
"version" : {
"number" : "6.7.1",
"build_flavor" : "default",
"build_type" : "docker",
"build_hash" : "2f32220",
"build_date" : "2019-04-02T15:59:27.961366Z",
"build_snapshot" : false,
"lucene_version" : "7.7.0",
"minimum_wire_compatibility_version" : "5.6.0",
"minimum_index_compatibility_version" : "5.0.0"
 },

```

I had a problem, probably solved in here: [CircuitBreakingException: [parent] Data too large, data for [\<transport\_request\>]](https://discuss.elastic.co/t/circuitbreakingexception-parent-data-too-large-data-for-transport-request/143486), so I've performed a full restart of the es-cluster (I've removed all of the podes and k8s deployments have restarted them).

However now the cluster cannot get up, because of the java.lang.IllegalStateException: alias [mongo] has more than one write index [mongo-2020.01.03-000033,mongo-2019.08.29-000015]

Full log of the master node here: [https://pastebin.com/raw/qYpkjn9B](https://pastebin.com/raw/qYpkjn9B)

It seems obvious to delete one of the write indexes, however I don't know how to do it and what type of index it is. [https://github.com/elastic/elasticsearch/blob/6.7/server/src/main/java/org/elasticsearch/cluster/metadata/MetaData.java](https://github.com/elastic/elasticsearch/blob/6.7/server/src/main/java/org/elasticsearch/cluster/metadata/MetaData.java) suggests that it is some kind of the metadata, however my cluster state says: [https://pastebin.com/zipTQ6Tk](https://pastebin.com/zipTQ6Tk)

so it seems that i have a block on the metadata write/read.

How can i fix it? I cannot remove the data, because it is a production cluster

Thanks for the answer

---

<div class="post-metadata">

**Author:** ![HenningAndersen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/henningandersen/32/48188_2.png) [@HenningAndersen](https://discuss.elastic.co/u/HenningAndersen)\
**Post date:** [January 7, 2020, 6:56pm UTC](https://discuss.elastic.co/t/elasticsearch-in-k8s-cannot-restart-because-of-recovery-failure-multiple-alias-write-indexes/214100/2 "2020-01-07T18:56:45Z")

</div>

Hi @MF57,

the cluster state dump has two nodes named "es-master-\*", I wonder if you have had a split-brain situation due to this? Is `discovery.zen.minimum_master_nodes` set correctly to 2 for this setup or was it ever wrong?

It might look like one version of the cluster state has the mongo alias with write index for mongo-2020.01.03-000033 and the other has the mongo alias with write index for mongo-2019.08.29-000015. Looking at the date for the index name, maybe an old master has been resurrected and joined the cluster after the full restart?

Would be good to also see the log file from the other master node as well as settings (in particular minimum\_master\_nodes).

---

<div class="post-metadata">

**Author:** ![MF57](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mf57/32/72394_2.png) [@MF57](https://discuss.elastic.co/u/MF57)\
**Post date:** [January 8, 2020, 7:13am UTC](https://discuss.elastic.co/t/elasticsearch-in-k8s-cannot-restart-because-of-recovery-failure-multiple-alias-write-indexes/214100/3 "2020-01-08T07:13:26Z")

</div>

Hi @HenningAndersen

Thank you for your answer. The discovery.zen.minimum\_master\_nodes is currently set to 2, however i have no knowledge if it was always set to 2 in the past.

Also since the full restart was done by removing the k8s pods I don't believe that an old master could be resurrected, because I can't imagine how - maybe I am wrong though.

I am providing full logs and configurations of the cluster (I've hidden cluster name but its the same everywhere):

```
GET cluster/_state https://pastebin.com/ktsDxKCy

GET /_nodes https://pastebin.com/utRrzb7U

GET cluster/_settings?include_defaults=true https://pastebin.com/yR3RWGuv

```

Nodes:  
All of the elasticsearch.yml are the same, but there different env values (provided in the comments of the config file)

es-master-6f6bf6f789-jn4rd

Config: [https://pastebin.com/3uNH3mRT](https://pastebin.com/3uNH3mRT)  
Logs: [https://pastebin.com/TX5nF8xY](https://pastebin.com/TX5nF8xY)

es-master-6f6bf6f789-zwwpk

Config: [https://pastebin.com/mweMTEQm](https://pastebin.com/mweMTEQm)  
Logs: [https://pastebin.com/i6Vz1UBy](https://pastebin.com/i6Vz1UBy)

es-data-0

Config: [https://pastebin.com/LFwsy7Ka](https://pastebin.com/LFwsy7Ka)  
Logs: [https://pastebin.com/NKX21BKX](https://pastebin.com/NKX21BKX)

es-data-1

Config: [https://pastebin.com/vJGj1Biw](https://pastebin.com/vJGj1Biw)  
Logs: [https://pastebin.com/WnVj0WkU](https://pastebin.com/WnVj0WkU)

es-client-86655db574-xlbxv

Config: [https://pastebin.com/7bCyeV26](https://pastebin.com/7bCyeV26)  
Logs: [https://pastebin.com/yzZhdNpB](https://pastebin.com/yzZhdNpB)

---

<div class="post-metadata">

**Author:** ![HenningAndersen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/henningandersen/32/48188_2.png) [@HenningAndersen](https://discuss.elastic.co/u/HenningAndersen)\
**Post date:** [January 8, 2020, 9:41am UTC](https://discuss.elastic.co/t/elasticsearch-in-k8s-cannot-restart-because-of-recovery-failure-multiple-alias-write-indexes/214100/4 "2020-01-08T09:41:26Z")

</div>

Hi @MF57,

I believe we need to manually fix this to get it running again. We should be able to find the UUID of the offending index by enabling trace logging (either globally or for `org.elasticsearch.gateway`).

Feel free to PM me the resulting log files.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 5, 2020, 9:41am UTC](https://discuss.elastic.co/t/elasticsearch-in-k8s-cannot-restart-because-of-recovery-failure-multiple-alias-write-indexes/214100/5 "2020-02-05T09:41:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
