# Replicas out of sync

**URL:** <https://discuss.elastic.co/t/replicas-out-of-sync/117128>\
**Category:** Elasticsearch\
**Created:** [January 25, 2018, 10:16pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128 "2018-01-25T22:16:24Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 25, 2018, 10:16pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/1 "2018-01-25T22:16:24Z")

</div>

For one particular index, I've been having issues with the primary and replicas repeatedly getting out of sync.

![image](https://us1.discourse-cdn.com/elastic/original/3X/8/1/8153b20fd5ed5f259db5b307fc5fcc5a213416e9.png)

When updating this index, I make a delete by query request to delete all values with a particular property, followed immediately by a bulk insert to re-add the updated values.

ES is version 5.6 running in a 5-node cluster.

I haven't been able to consistently reproduce it, and I can fix it temporarily by switching the replicas to 0 and back to 1 to get ES to rebuild them, but the issue seems to crop back up after a day or so.

Anyone run into this before?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 26, 2018, 8:21am UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/2 "2018-01-26T08:21:04Z")

</div>

Hi there,

We have seen very occasional reports of this, and have been investigating, but it has proved _extremely_ tricky for us to reproduce. We need help from someone like you who sees this problem regularly enough to be useful in diagnosis.

Please could you tell us more about this cluster and the environment in which it lives? For instance: what version are you running exactly? What is it running on? How frequently are you doing the bulk-delete-and-insert that you describe? What other activity does the cluster see?

Would you be willing to run the [support diagnostics tool](https://github.com/elastic/elasticsearch-support-diagnostics/releases) on your cluster and share the results? Don't post them here: I'll get you an email address to use if you can run this.

Would you be able to enable the following **very verbose** logging, and toggle the replica count to 0 and then back to 1 to make sure everything is in sync? I say again that this is **very verbose** so it will cause extra I/O and may fill up your disks, so proceed with caution here.

```
, "logger.org.elasticsearch.action.bulk": "TRACE"
, "logger.org.elasticsearch.cluster.service": "DEBUG"
, "logger.org.elasticsearch.indices.recovery": "TRACE"
, "logger.org.elasticsearch.index.shard": "TRACE"

```

In case it helps, we've only so far been able to reproduce anything like this by simulating some very strange networking failures that coincide with shards being reallocated, and even then it's very sporadic.

---

<div class="post-metadata">

**Author:** ![bleskes](https://avatars.discourse-cdn.com/v4/letter/b/71c47a/32.png) [@bleskes](https://discuss.elastic.co/u/bleskes)\
**Post date:** [January 26, 2018, 10:30am UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/3 "2018-01-26T10:30:31Z")

</div>

On top of what David suggested, can you share the exact version you use?

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 26, 2018, 3:02pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/4 "2018-01-26T15:02:16Z")

</div>

The cluster is a 5 node cluster, all running ES 5.4.0 as master/client/data, all Centos7.3 VMs with no plugins. Cluster has about 5 indexes in it, the largest one of which has about 1.6M documents with a fair amount of churn which has never gotten out of sync. It updates by diffing and just performing bulk update/deletes on individual documents which aside from # of documents is the only significant difference between it and the index that is getting out of sync. The index which is causing trouble is pretty new and just has a couple thousand documents, but is updated by just deleting (via delete\_by\_query) and re-adding groups of documents.

Version details:  
"version": {  
"number": "5.4.0",  
"build\_hash": "780f8c4",  
"build\_date": "2017-04-28T17:43:27.229Z",  
"build\_snapshot": false,  
"lucene\_version": "6.5.0"  
}

I'll see if I can enable verbose logging. Unfortunately I can only reproduce this at the moment on our production cluster so will have to check about running the diagnostics tool.

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 26, 2018, 3:05pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/5 "2018-01-26T15:05:37Z")

</div>

The indexing on the client is being done with NEST on .net. Code looks about like this:

```
      var actions = terms
    .Select(term => new BulkIndexOperation<SearchTerm>(term)
    {
      Routing = term.ClientId
    })
    .Cast<IBulkOperation>();
 
  Client.Instance.DeleteByQuery(new DeleteByQueryRequest("inventory_suggestions", typeof(SearchTerm))
  {
    Query = new TermQuery
    {
      Field = typeof(SearchTerm).GetProperty(nameof(SearchTerm.ClientId)),
      Value = _client.ClientId.ToString()
    }
  });

  Client.Instance.Bulk(new BulkRequest("inventory_suggestions")
  {
    Operations = actions.ToList()
  });
```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 26, 2018, 3:43pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/6 "2018-01-26T15:43:32Z")

</div>

Thanks, Eric, much appreciated.

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 26, 2018, 5:54pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/7 "2018-01-26T17:54:29Z")

</div>

We turned on the trace logging and were able to reproduce it getting out of sync on one of the shards. Seems to only be ~15 documents off at the moment. What's the best way to share the logs?

![image](https://us1.discourse-cdn.com/elastic/original/3X/1/4/14adc1cf7e0a1dc84a2393ae9e083c3b8fc8a6cc.png)

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 26, 2018, 6:07pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/8 "2018-01-26T18:07:32Z")

</div>

Please could you zip them up and send them to me at [david.turner@elastic.co](mailto:david.turner@elastic.co)? I'm unlikely to look at them before 0900 UTC Monday now, so don't promise an immediate response.

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 26, 2018, 7:02pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/9 "2018-01-26T19:02:31Z")

</div>

Logs sent. Thank you very much for your help.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 29, 2018, 2:20pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/10 "2018-01-29T14:20:05Z")

</div>

Hi Eric,

Thanks for the logs, they're much appreciated. We have come up with one hypothesis about what might possibly be happening here, but unfortunately cannot test it from those logs alone. Could you possibly repeat the period of trace logging with the same settings, starting from a point where the shards are all in sync, wait for them to fall out of sync, and then grab a list of all the document IDs on both primary and replica as well as the logs? Ideally we'd like the indexing process to be stopped and for you to perform a refresh before querying for the doc IDs to make sure that we get everything.

Many thanks,

David

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 29, 2018, 3:40pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/11 "2018-01-29T15:40:24Z")

</div>

Awesome. We've got the trace logging enabled as before. I'll send you those results once it starts getting out of sync again.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 29, 2018, 3:46pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/12 "2018-01-29T15:46:02Z")

</div>

Thanks Eric. Could you also confirm that, in this index at least, you're using auto-generated IDs, and not using external versioning at all?

Many thanks,

David

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 29, 2018, 3:47pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/13 "2018-01-29T15:47:29Z")

</div>

Yeah, that's correct.

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 29, 2018, 6:09pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/14 "2018-01-29T18:09:24Z")

</div>

Sent you logs along with the ids on primary/replica. I stopped all indexing and did a refresh of the index prior to pulling the IDs.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 30, 2018, 9:26am UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/15 "2018-01-30T09:26:50Z")

</div>

Awesome. We didn't find exactly what we expected, but we weren't far off. It seems there are occasions where you index a document and delete it very soon afterwards (before the indexing operation has even returned to the client), and the indexing and deletion operations are arriving in the wrong order at the replica, and for some reason (still under investigation) they're not being put back in the right order. We can now reproduce this with a single document.

As a workaround for you for now, I think it'd be sufficient to avoid running concurrent deletion and indexing operations on your `inventory_suggestions_v1` index. Could you try that?

---

<div class="post-metadata">

**Author:** ![edamtoft](https://avatars.discourse-cdn.com/v4/letter/e/2acd7d/32.png) [@edamtoft](https://discuss.elastic.co/u/edamtoft)\
**Post date:** [January 30, 2018, 10:55pm UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/16 "2018-01-30T22:55:34Z")

</div>

I put some stuff in place to try to prevent concurrent indexing and it hasn't gotten out of sync since. For a longer term workaround, I'm also changing around the indexing strategy some to do a bit more targeted updated/deletes since I think that would decrease the churn considerably for my use case.

---

<div class="post-metadata">

**Author:** ![eebbrr](https://avatars.discourse-cdn.com/v4/letter/e/5f9b8f/32.png) [@eebbrr](https://discuss.elastic.co/u/eebbrr)\
**Post date:** [January 31, 2018, 5:50am UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/17 "2018-01-31T05:50:18Z")

</div>

It's an old blog post, but I've never seen it documented anywhere else. Read this: [https://www.elastic.co/blog/elasticsearch-versioning-support](https://www.elastic.co/blog/elasticsearch-versioning-support)

EDIT: As it seems this kind of "documentation" goes, you want to just skip ahead to the last section titled "Some final words about deletes.". What you need to know is never at the start!

I'd have to guess that you're re-using \_id values within the window of ES' deleted document garbage collection process -- which would be unrelated to concurrent deleting/indexing.

There's likely a few solutions, but always using an (index-wide) increasing version number for each new doc is one way to fix this. Probably using ES' auto-generated \_ids is another but I haven't confirmed that approach.

Good luck!

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 31, 2018, 10:46am UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/18 "2018-01-31T10:46:47Z")

</div>

> [@eebbrr](#):
>
> I've never seen it documented anywhere else

The documentation on the relationship between deletes and versioning is indeed quite scarce, and I agree that this should be properly spelled out in the reference manual. It is, however, not relevant in this case.

> [@eebbrr](#):
>
> I'd have to guess that you're re-using \_id values

The OP is not re-using document IDs.

> [@eebbrr](#):
>
> There's likely a few solutions, but always using an (index-wide) increasing version number for each new doc is one way to fix this. Probably using ES' auto-generated \_ids is another but I haven't confirmed that approach.

Assigning document IDs based on an external counter is certainly possible, but it's quite tricky to make it robust to all the things that might go wrong in your system, particularly network partitions and GC pauses. Auto-generated IDs allow Elasticsearch to do this for you, so I'd say to use that functionality unless you have a very compelling reason to use externally-assigned IDs.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [January 31, 2018, 10:54am UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/19 "2018-01-31T10:54:48Z")

</div>

> [@edamtoft](#):
>
> I put some stuff in place to try to prevent concurrent indexing and it hasn't gotten out of sync since.

Great news. Thanks for letting us know.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 28, 2018, 10:54am UTC](https://discuss.elastic.co/t/replicas-out-of-sync/117128/20 "2018-02-28T10:54:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
