# GET api by doc\_id returns different result whenever i try

**URL:** <https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293>\
**Category:** Elasticsearch\
**Created:** [September 5, 2023, 1:34am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293 "2023-09-05T01:34:12Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![ycice](https://avatars.discourse-cdn.com/v4/letter/y/ed655f/32.png) [@ycice](https://discuss.elastic.co/u/ycice)\
**Post date:** [September 5, 2023, 1:34am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/1 "2023-09-05T01:34:12Z")

</div>

Hi, i manage more than 100 ES clusters in my company for 3 years  
But at last week, I faced very strange issue. I think it is not possible... Could you carefully check this?

ES version : 6.8.2  
Cluster health : Green

GET api by doc\_id returns different result whenever i try  
As you can see in the picture, sometimes it says that there isn't such a document. And sometimes it says the result

I know score search can show different result since primary and replica can have different merge timing. I already read [Getting consistent scoring | Elasticsearch Guide [6.8] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/6.8/consistent-scoring.html#consistent-scoring)  
But this is GET api and it should show consistent result, shouldn't it?

And this symptom is not transient. It is on-going for 1 week. It is still happening  
This ES cluster has more than 100K ops/s indexing rate so merge would happen quite regularly

 ![스크린샷 2023-09-04 오후 9.50.59](https://us1.discourse-cdn.com/elastic/original/3X/0/c/0cf6afbfdb6154544351c7f85bbbd58b3b057cc9.png)  
 ![스크린샷 2023-09-04 오후 9.51.11](https://us1.discourse-cdn.com/elastic/original/3X/6/a/6afcbfdcadf797ddd4dcf874cb7a396122c22948.jpeg)

And when i try scoring search with preference,  
when i use \_primary, there is result  
when i use \_replica, there isn't result  
(Check below pictures also) Replica count is 1  
So it seems like primary shard has the document and replica doesn't have.

 ![스크린샷 2023-09-05 오전 10.23.50](https://us1.discourse-cdn.com/elastic/original/3X/2/4/244af347cf2a2d587b5df49daad1cf19af900d2c.jpeg)  
 ![스크린샷 2023-09-05 오전 10.24.03](https://us1.discourse-cdn.com/elastic/original/3X/0/1/01c09861cf7ce4ee4aded7aee3607c50d36d99f9.jpeg)

I tried using "realtime=false" but same symptom happened

As i said, cluster health is Green. When i checked master log, there isn't error log for now

We did DR test, make 1 AZ in AWS region down and see if our system is okay, at last week and this symptom happened after DR test  
During the DR test, our ES cluster succeeded in failover. There was 3 min downtime but after that, remained nodes worked well. At that time, cluster state was yellow, not red

So i also suspected that this DR situation made corruption between primary and replica...But if so, it shouldn't be fixed automatically? ES seems not to detect this unsync problem.

If you provide your opinion, it would be thankful

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 5, 2023, 4:26am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/2 "2023-09-05T04:26:21Z")

</div>

I suspect you cluster is incorrectly configured and may be suffering from a split brain scenario.

How many master eligible nodes does the cluster have?

What is `discovery.zen.minimum_master_nodes` set to in your configutation?

For the cluster to be correctly configured this parameter should be set to the number of master eligible nodes required to for a strict majority in the cluster, e.g. `2` if you have 2 or 3 master eligible nodes, `3` if you have 4 or 5 master eligible nodes etc. This has always been a common area of misconfiguration, and can cause data loss and inconsistencies (similar to what you are seeing). This setting was removed as resiliency was improved in version 7.0 so is not an issue in newer versions. I would recommend you upgrade any older clusters.

---

<div class="post-metadata">

**Author:** ![ycice](https://avatars.discourse-cdn.com/v4/letter/y/ed655f/32.png) [@ycice](https://discuss.elastic.co/u/ycice)\
**Post date:** [September 5, 2023, 5:29am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/3 "2023-09-05T05:29:50Z")

</div>

No, i don't think it is split brain issue  
We have 3 master node and discovery.zen.minimum\_master\_nodes is 2  
Each master is in A, B, C zone. So when DR situation, only 1 master in A was down and there wan't any split brain issue  
For now, 3 master joins well

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 5, 2023, 6:03am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/4 "2023-09-05T06:03:49Z")

</div>

> [@ycice](#):
>
> So i also suspected that this DR situation made corruption between primary and replica...But if so, it shouldn't be fixed automatically? ES seems not to detect this unsync problem.

If you do not have `minimum_master_nodes` configured incorrectly I would expect Elasticsearch to recover correctly. Do you have any custom settings around recovery or [translog durability](https://www.elastic.co/guide/en/elasticsearch/reference/6.8/index-modules-translog.html#_translog_settings) that could impact this?

I recall there were issues with resiliency in Elsticsearch 6 and earlier, which is why [significant improvements to resiliency and stability were made in version 7](https://www.elastic.co/blog/a-new-era-for-cluster-coordination-in-elasticsearch).

To resolve this you should be able to drop the replica and then recreate it so it is guaranteed to be a copy of the primary. To avoid this in the future I would again recommend upgrading.

---

<div class="post-metadata">

**Author:** ![ycice](https://avatars.discourse-cdn.com/v4/letter/y/ed655f/32.png) [@ycice](https://discuss.elastic.co/u/ycice)\
**Post date:** [September 5, 2023, 7:15am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/5 "2023-09-05T07:15:26Z")

</div>

Here is our zen.discovery cluster settings. Most of them are just default

 ![스크린샷 2023-09-05 오후 3.26.55](https://us1.discourse-cdn.com/elastic/original/3X/2/b/2bbf3a7dfc139cb17a48c15d9c32f01fac60a4df.png)

And when i checked index setting, we don't use any custom translog setting

Okay, i understand what you mean. I want to ask several more questions and get confirm

1. While we are testing DR situation(network down and recovery), if translog itself or replay had issue, then consistency between primary and replica can be broken? Do you suspect that part?

2. There is no way to do manual sync between primary and replica? Only reduce replica and recreate ?

Thanks for your help. If you have related issue link, please share it. Thanks

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 5, 2023, 7:21am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/6 "2023-09-05T07:21:08Z")

</div>

> [@ycice](#):
>
> While we are testing DR situation(network down and recovery), if translog itself or replay had issue, then consistency between primary and replica can be broken? Do you suspect that part?

If the replica that comes back is out of date it should be replaced by a copy of the primary, so I would not expect this unles you are hitting some bug or have some custom setting that reduces reliability. I have not troubleshot this type of issues in version 6 for yaers so do not remember details.

> [@ycice](#):
>
> There is no way to do manual sync between primary and replica? Only reduce replica and recreate ?

No, not as far as I know. The addition of sequence numbers in version 7 made this a lot better and more efficient, but in version 6 I think that is the only way.

---

<div class="post-metadata">

**Author:** ![ycice](https://avatars.discourse-cdn.com/v4/letter/y/ed655f/32.png) [@ycice](https://discuss.elastic.co/u/ycice)\
**Post date:** [September 5, 2023, 7:30am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/7 "2023-09-05T07:30:42Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> If the replica that comes back is out of date it should be replaced by a copy of the primary, so I would not expect this unles you are hitting some bug or have some custom setting that reduces reliability. I have not troubleshot this type of issues in version 6 for yaers so do not remember details.

Network disconnect time was short, about 20 minutes  
So existing replica seemed not to be replaced. According to our metric, there seemed no index recovery which usually happens when shard moves  
So existing shard seemed to get translog replay from primary

---

<div class="post-metadata">

**Author:** ![ycice](https://avatars.discourse-cdn.com/v4/letter/y/ed655f/32.png) [@ycice](https://discuss.elastic.co/u/ycice)\
**Post date:** [September 12, 2023, 1:53am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/8 "2023-09-12T01:53:49Z")

</div>

I want to ask if there is any ticket related to this kind of issue?  
Or where can i search related tickets? Github or Jira?

And i also want to ask if we upgrade to version 7, we can guarantee that this kind of issue doesn't happen.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 10, 2023, 1:54am UTC](https://discuss.elastic.co/t/get-api-by-doc-id-returns-different-result-whenever-i-try/342293/9 "2023-10-10T01:54:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
