# Cluster "green" - but shards not in sync !? 😯

**URL:** <https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-monitoring\
**Created:** [March 25, 2021, 1:04pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344 "2021-03-25T13:04:57Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![some\_one](https://avatars.discourse-cdn.com/v4/letter/s/dfb087/32.png) [@some\_one](https://discuss.elastic.co/u/some_one)\
**Post date:** [March 25, 2021, 1:04pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/1 "2021-03-25T13:04:57Z")

</div>

Hello all,

we have been using ES productively for quite some time now but today we came across a new problem that we have never seen before:

We found two indices in our ES 6.8.14 cluster, each with one shard where primary and replica show different sync\_ids and vastly different document counts. I reckon that means they are totally out of sync. I am having a hard time understanding how the cluster health can be "green" under these circumstances, though.

```auto
curl -H 'Content-Type: application/json' -XGET "prod-db01:9200/_cat/shards/archive1599113539"
archive1599113539 2 r STARTED 6720 40.4mb 10.0.82.232 prod-db01
archive1599113539 2 p STARTED 656 4.3mb 10.0.82.233 prod-db02

```

As you can see, the document count is far greater on the replica. Since it looks like the problem has gone undetected for weeks at least (since the cluster is "green") backups are probably unusable. Also, the usually recommended way of setting replicas to 0 and then back to 1 will probably result in significant data loss because there are more documents in the replica than in the primary.

I am sort of clueless how to deal with the situation and ask you:

1. How can the cluster be _green_? Should this not be a bug?
2. Is there ANY thinkable way of evaluating or dumping the data in primary and replica separately for a merge attempt to recover otherwise potentially lost data sets?
3. How do I prevent this from happening again and do you have tips for detecting this type of failure?

Any help or tips would be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 25, 2021, 1:22pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/2 "2021-03-25T13:22:59Z")

</div>

Can you share `GET /archive1599113539/_stats?level=shards` please?

---

<div class="post-metadata">

**Author:** ![some\_one](https://avatars.discourse-cdn.com/v4/letter/s/dfb087/32.png) [@some\_one](https://discuss.elastic.co/u/some_one)\
**Post date:** [March 25, 2021, 1:46pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/3 "2021-03-25T13:46:27Z")

</div>

Sorry but the output you requested exceeds the max size for a posting here.

You can find the data you requested here: [shard stats - Pastebin.com](https://pastebin.com/NjTXJyw3)

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 25, 2021, 2:28pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/4 "2021-03-25T14:28:08Z")

</div>

Thanks, yes, that looks broken. How many master nodes are there in your cluster? How is `discovery.zen.minimum_master_nodes` configured?

---

<div class="post-metadata">

**Author:** ![some\_one](https://avatars.discourse-cdn.com/v4/letter/s/dfb087/32.png) [@some\_one](https://discuss.elastic.co/u/some_one)\
**Post date:** [March 25, 2021, 2:57pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/5 "2021-03-25T14:57:40Z")

</div>

> [@DavidTurner](#):
>
> discovery.zen.minimum\_master\_nodes

There seems to be no such setting at all. It is a three node cluster. All nodes have roles 'master', 'data' and 'ingest' set.  
This is the contents of the discovery setting:

```auto
        "discovery" : {
          "zen" : {
            "ping" : {
              "unicast" : {
                "hosts" : [
                  "10.0.82.232",
                  "10.0.82.233",
                  "10.0.82.234"
                ]
              }
            }
          }
        },
```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 25, 2021, 3:07pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/6 "2021-03-25T15:07:21Z")

</div>

Ok if you have not set `discovery.zen.minimum_master_nodes` then that would explain it. There should be warnings about it in your logs, looking like this:

> `value for setting "discovery.zen.minimum_master_nodes" is too low. This can result in data loss!`

---

<div class="post-metadata">

**Author:** ![some\_one](https://avatars.discourse-cdn.com/v4/letter/s/dfb087/32.png) [@some\_one](https://discuss.elastic.co/u/some_one)\
**Post date:** [March 25, 2021, 3:41pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/7 "2021-03-25T15:41:36Z")

</div>

Thank you, set it to 2 now. It was commented out for reasons we could not reconstruct.  
I found one such warning in a log backup from 2020-06 🙂

Any more ideas about what can be done to recover as much data as possible? 🙂

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 25, 2021, 4:30pm UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/8 "2021-03-25T16:30:36Z")

</div>

> [@some\_one](#):
>
> Any more ideas about what can be done to recover as much data as possible?

Nothing very easy or robust, sorry. You could try using [search preference](https://www.elastic.co/guide/en/elasticsearch/reference/6.8/search-request-preference.html) to extract the contents of each shard so you can compare them. You'll need to do that for every shard, even the ones with matching doc counts, just to check that they really have the same docs in them and nothing got messed up in their mappings either.

---

<div class="post-metadata">

**Author:** ![some\_one](https://avatars.discourse-cdn.com/v4/letter/s/dfb087/32.png) [@some\_one](https://discuss.elastic.co/u/some_one)\
**Post date:** [March 26, 2021, 11:23am UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/9 "2021-03-26T11:23:26Z")

</div>

Thanks again! It looks like we are going to be able to recover most of the data.

I'd like to get back to the original question, though: I still fail to comprehend how inconsistent replicas are not a sufficient condition to trigger a health warning?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 26, 2021, 11:40am UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/10 "2021-03-26T11:40:50Z")

</div>

> [@some\_one](#):
>
> I still fail to comprehend how inconsistent replicas are not a sufficient condition to trigger a health warning?

It sort of does. The cluster health depends only on whether the shards are assigned or not, and the assignment process includes checks to make sure that [all the copies are in sync](https://www.elastic.co/blog/tracking-in-sync-shard-copies). Unfortunately by configuring `discovery.zen.minimum_master_nodes` wrongly you end up with the information about which copies are in sync itself being out of sync, but there's not really a way to address that in general.

This is fixed in 7.x, in the sense that it is no longer possible to misconfigure Elasticsearch to lose data in this fashion.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 23, 2021, 11:41am UTC](https://discuss.elastic.co/t/cluster-green-but-shards-not-in-sync/268344/11 "2021-04-23T11:41:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
