# GlobalCheckpoint syncer only syncs translog when the translog durability is REQUEST

**URL:** <https://discuss.elastic.co/t/globalcheckpoint-syncer-only-syncs-translog-when-the-translog-durability-is-request/272468>\
**Category:** Elasticsearch\
**Created:** [May 8, 2021, 10:06am UTC](https://discuss.elastic.co/t/globalcheckpoint-syncer-only-syncs-translog-when-the-translog-durability-is-request/272468 "2021-05-08T10:06:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![asce0705](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/asce0705/32/81317_2.png) [@asce0705](https://discuss.elastic.co/u/asce0705)\
**Post date:** [May 8, 2021, 10:06am UTC](https://discuss.elastic.co/t/globalcheckpoint-syncer-only-syncs-translog-when-the-translog-durability-is-request/272468/1 "2021-05-08T10:06:56Z")

</div>

Elasticsearch version: 7.10  
Currently, `IndexShard.maybeSyncGlobalCheckpoint` is called in two places:

1. In the `AsyncGlobalCheckpointTask` of `IndexService`;
2. In the `TransportReplicationAction` after a write action.  
In `IndexShard.maybeSyncGlobalCheckpoint`, it runs the `globalCheckpointSyncer` according to the following conditions:

```auto
    // only sync if there are no operations in flight, or when using async durability
    final SeqNoStats stats = getEngine().getSeqNoStats(replicationTracker.getGlobalCheckpoint());
    final boolean asyncDurability = indexSettings().getTranslogDurability() == Translog.Durability.ASYNC;
        if (stats.getMaxSeqNo() == stats.getGlobalCheckpoint() || asyncDurability) {
            final ObjectLongMap<String> globalCheckpoints = getInSyncGlobalCheckpoints();
            final long globalCheckpoint = replicationTracker.getGlobalCheckpoint();
            // async durability means that the local checkpoint might lag (as it is only advanced on fsync)
            // periodically ask for the newest local checkpoint by syncing the global checkpoint, so that ultimately the global
            // checkpoint can be synced. Also take into account that a shard might be pending sync, which means that it isn't
            // in the in-sync set just yet but might be blocked on waiting for its persisted local checkpoint to catch up to
            // the global checkpoint.
            final boolean syncNeeded =
                (asyncDurability && (stats.getGlobalCheckpoint() < stats.getMaxSeqNo() || replicationTracker.pendingInSync()))
                    // check if the persisted global checkpoint
                    || StreamSupport
                            .stream(globalCheckpoints.values().spliterator(), false)
                            .anyMatch(v -> v.value < globalCheckpoint);
            // only sync if index is not closed and there is a shard lagging the primary
            if (syncNeeded && indexSettings.getIndexMetadata().getState() == IndexMetadata.State.OPEN) {
                logger.trace("syncing global checkpoint for [{}]", reason);
                globalCheckpointSyncer.run();
            }
        }

```

One of the condition checks the translog durability, which should be `ASYNC`, and if the local checkpoint lags, it runs the `GlobalCheckpointSyncer`, which will then execute `GlobalCheckpointSyncAction`. This action syncs the translog of the given indexShard when the translog durability is `REQUEST`.

```auto
    private void maybeSyncTranslog(final IndexShard indexShard) throws IOException {
        if (indexShard.getTranslogDurability() == Translog.Durability.REQUEST &&
            indexShard.getLastSyncedGlobalCheckpoint() < indexShard.getLastKnownGlobalCheckpoint()) {
            indexShard.sync();
        }
    }

```

This condition and the action behavior conflicts. Should we remove the translog durability check int `GlobalCheckpointSyncAction`?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [May 8, 2021, 11:41am UTC](https://discuss.elastic.co/t/globalcheckpoint-syncer-only-syncs-translog-when-the-translog-durability-is-request/272468/2 "2021-05-08T11:41:55Z")

</div>

> [@asce0705](#):
>
> Should we remove the translog durability check int `GlobalCheckpointSyncAction` ?

I don't think so, we don't want to actually sync the translog here if the durability is `ASYNC`. In that case we just want to run a no-op `GlobalCheckpointSyncAction` to ensure that the primary's `ReplicationTracker` is kept up to date.

---

<div class="post-metadata">

**Author:** ![asce0705](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/asce0705/32/81317_2.png) [@asce0705](https://discuss.elastic.co/u/asce0705)\
**Post date:** [May 14, 2021, 6:44am UTC](https://discuss.elastic.co/t/globalcheckpoint-syncer-only-syncs-translog-when-the-translog-durability-is-request/272468/4 "2021-05-14T06:44:15Z")

</div>

Indeed as you say. By running a no-op `GlobalCheckpointSyncAction` , the replicas learn about current `globalCheckpoint` from primary and the primary collects the `localCheckpoints` from replicas, which may result in `globalCheckpoint` advance and `checkpoints` update in `ReplicationTracker` .

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 11, 2021, 6:45am UTC](https://discuss.elastic.co/t/globalcheckpoint-syncer-only-syncs-translog-when-the-translog-durability-is-request/272468/5 "2021-06-11T06:45:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
