# Confusions about how globlal checkpoint advances in es 7.6

**URL:** <https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698>\
**Category:** Elasticsearch\
**Created:** [March 2, 2020, 2:25pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698 "2020-03-02T14:25:58Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![iamorchid](https://avatars.discourse-cdn.com/v4/letter/i/bbce88/32.png) [@iamorchid](https://discuss.elastic.co/u/iamorchid)\
**Post date:** [March 2, 2020, 2:25pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/1 "2020-03-02T14:25:58Z")

</div>

In es 7.6, I see that we introduced persisted local checkpoint and processed local checkpoint. After each update replication, replicas return their **persisted** local checkpoint to primary shard. And based on these checkpoints from in-sync replicas, primary would advance the global check point. However, since **persisted** local checkpoint could lag behind the op seq that's really processed(namely **processed** local checkpoint)，I wonder is it possible for the global checkpoint to catch up to MAX\_SEQ.

Here I given an example to illustrate this. Suppose we have 3 copies for one shard (suppose all copies have inital seq no 6 for both **persisted** and **processed** local checkpoint). And after replicating an update, we could have:

> Primary : **persisted** local checkpoint is 6, **processed** local checkpoint 7  
> Replica1: **persisted** local checkpoint is 6, **processed** local checkpoint 7  
> Replica2: **persisted** local checkpoint is 6, **processed** local checkpoint 7

For the moment, the global checkpoint doesn't advance and it lags behind MAX\_SEQ 7。I don't see when the global checkpoint would catch up to MAX\_SEQ.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 2, 2020, 4:28pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/2 "2020-03-02T16:28:23Z")

</div>

If there are more operations in flight then we know that a future operation will eventually advance the global checkpoint. If there are no more operations in flight then the `GlobalCheckpointSyncAction` brings everything up to date instead.

---

<div class="post-metadata">

**Author:** ![HenningAndersen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/henningandersen/32/48188_2.png) [@HenningAndersen](https://discuss.elastic.co/u/HenningAndersen)\
**Post date:** [March 2, 2020, 4:35pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/3 "2020-03-02T16:35:16Z")

</div>

Hi @iamorchid,

I believe you are are referring to [this change](https://github.com/elastic/elasticsearch/pull/43205) which went into 7.3.

The persisted local checkpoint should eventually catch up unless there is something serious wrong (bug or inability to persist translog).

Before the change mentioned above, the global checkpoint could advance based on non-durably stored information from replicas. While data were safe, we rely on all data below global checkpoint to be durably stored to ensure that replicas and primary have the same data (as well as for CCR).

Is this based on a theoretical exercise/question or have you run into an issue where the global checkpoint no longer advances?

---

<div class="post-metadata">

**Author:** ![iamorchid](https://avatars.discourse-cdn.com/v4/letter/i/bbce88/32.png) [@iamorchid](https://discuss.elastic.co/u/iamorchid)\
**Post date:** [March 2, 2020, 4:43pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/4 "2020-03-02T16:43:40Z")

</div>

@HenningAndersen, thanks for your reply. @DavidTurner mentioned **GlobalCheckpointSyncAction** , I noticed that we have the following logic in **[IndexShard.java](https://github.com/elastic/elasticsearch/blob/7.6/server/src/main/java/org/elasticsearch/index/shard/IndexShard.java#L2338)** to trigger the action. However, if we are not using asyncDurability, the action would not be triggered, right (suppose global checkpoint is behind MAX\_SEQ)? Or do we have other chances to tirgger that action?

 ![maybeSyncGlobalCheckpoint](https://us1.discourse-cdn.com/elastic/original/3X/9/2/92778ab1ae1889613265fae6b9e510ff08798b17.png)

The question is based on my analysis about the code (I'm a new learner and hope understand better about how ES works ).

Thanks

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 2, 2020, 5:15pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/5 "2020-03-02T17:15:19Z")

</div>

`maybeSyncGlobalCheckpoint` runs periodically and regardless of the durability setting it syncs the global checkpoint if `stats.getMaxSeqNo() == stats.getGlobalCheckpoint()`, which is a way of expressing "no operations in flight".

---

<div class="post-metadata">

**Author:** ![iamorchid](https://avatars.discourse-cdn.com/v4/letter/i/bbce88/32.png) [@iamorchid](https://discuss.elastic.co/u/iamorchid)\
**Post date:** [March 2, 2020, 5:20pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/6 "2020-03-02T17:20:22Z")

</div>

But based on my analysis above, we could have **stats.getMaxSeqNo() \> stats.getGlobalCheckpoint()**. And to make global checkpoint to cache pu to MAX\_SEQ, we have to trigger **GlobalCheckpointSyncAction**. However, since **stats.getMaxSeqNo() \> stats.getGlobalCheckpoint()**, the action won't be triggered.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 2, 2020, 5:37pm UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/7 "2020-03-02T17:37:49Z")

</div>

With request-level durability every operation is properly fsynced before responding to the primary. This means that the last operation to complete on each copy will advance that copy's (persisted) local checkpoint up to the maximum sequence number and then tell the primary about it. Note that these things don't necessarily happen in sequence-number order, so the last operation isn't necessarily the one with the highest sequence number, but that doesn't affect the argument. Once all copies have completed all operations this means that the global checkpoint (on the primary) should indeed equal the maximum sequence number.

In some sense this is why we have to handle async durability differently in `maybeSyncGlobalCheckpoint`: with async durability it isn't true that "no operations in flight" implies `stats.getMaxSeqNo() == stats.getGlobalCheckpoint()` so we must be conservative and do some extra work to make sure.

---

<div class="post-metadata">

**Author:** ![iamorchid](https://avatars.discourse-cdn.com/v4/letter/i/bbce88/32.png) [@iamorchid](https://discuss.elastic.co/u/iamorchid)\
**Post date:** [March 3, 2020, 1:18am UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/8 "2020-03-03T01:18:36Z")

</div>

I see what you mean. For some requests, it doesn't need to wait for **GlobalCheckpointSyncAction** and they can trigger sync before replica replies the primary, right ? I would check more details.

Thanks again for all your detailed explanations @DavidTurner @HenningAndersen !

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [March 3, 2020, 7:47am UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/9 "2020-03-03T07:47:19Z")

</div>

Yes that's correct, most of the time there's ongoing indexing activity in which the primary sends replication requests to the replicas and they respond with a `TransportReplicationAction#ReplicaResponse` which includes the replica's new persisted local checkpoint and its last-synced global checkpoint. We only fall back on an explicit sync when the indexing activity finishes.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 31, 2020, 7:47am UTC](https://discuss.elastic.co/t/confusions-about-how-globlal-checkpoint-advances-in-es-7-6/221698/10 "2020-03-31T07:47:22Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
