# How to ensure unique sorting with search\_after in Elasticsearch 8?

**URL:** <https://discuss.elastic.co/t/how-to-ensure-unique-sorting-with-search-after-in-elasticsearch-8/375233>\
**Category:** Elasticsearch\
**Created:** [February 28, 2025, 5:50pm UTC](https://discuss.elastic.co/t/how-to-ensure-unique-sorting-with-search-after-in-elasticsearch-8/375233 "2025-02-28T17:50:18Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![jiel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jiel/32/141720_2.png) [@jiel](https://discuss.elastic.co/u/jiel)\
**Post date:** [February 28, 2025, 5:50pm UTC](https://discuss.elastic.co/t/how-to-ensure-unique-sorting-with-search-after-in-elasticsearch-8/375233/1 "2025-02-28T17:50:18Z")

</div>

Hello,

I am working on a tool that needs to retrieve large batches of records from an Elasticsearch index. The recommended method used to be the **Scroll API** , but it is now deprecated in favor of **search\_after** as stated in the [documentation](https://www.elastic.co/guide/en/elasticsearch/reference/current/scroll-api.html).

I want to use sorting criteria that are agnostic to the document content. Sorting by **@timestamp** and **\_id** seems appropriate, but sorting on **\_id** is [now disabled by default in Elasticsearch 8](https://www.elastic.co/guide/en/elasticsearch/reference/current/migrating-8.0.html).

If I sort only by **@timestamp** , this value is not unique, which means I could miss some records.

So is there a way to efficiently retrieve large volumes of data (\>10'000) while ensuring no records are skipped, using sorting criteria independent of document content?

---

<div class="post-metadata">

**Author:** ![strawgate](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/strawgate/32/131008_2.png) [@strawgate](https://discuss.elastic.co/u/strawgate)\
**Post date:** [March 1, 2025, 2:14pm UTC](https://discuss.elastic.co/t/how-to-ensure-unique-sorting-with-search-after-in-elasticsearch-8/375233/2 "2025-03-01T14:14:28Z")

</div>

Can you use the point in time API with search after? It adds a unique tiebreaker using the shard doc value.

If not, and you don't have a unique value then the flow that I use is:

1. Get a search\_after batch of 10k documents
2. Grab the timestamp from the last document (aka max\_timestamp)
3. Iterate through the batch `while timestamp < max_timestamp` doing whatever processing is required
4. On the last iteration grab the sort values for your next search\_after call

This effectively means you're excluding the documents with max timestamp from processing and ensuring that all documents with max\_timestamp will be present in your next search\_after call

---

<div class="post-metadata">

**Author:** ![jiel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jiel/32/141720_2.png) [@jiel](https://discuss.elastic.co/u/jiel)\
**Post date:** [March 5, 2025, 9:52am UTC](https://discuss.elastic.co/t/how-to-ensure-unique-sorting-with-search-after-in-elasticsearch-8/375233/3 "2025-03-05T09:52:29Z")

</div>

> [@strawgate](#):
>
> it adds a unique tiebreaker using the shard doc value.

Thanks William. I tried to sort on `tie_breaker_id` as the example from the documentation but the query no longer returned any results (It is theoretically available, my instance is in version 8.17).

Anyway, I got a working solution by sorting on the \_shard\_doc field.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 2, 2025, 9:52am UTC](https://discuss.elastic.co/t/how-to-ensure-unique-sorting-with-search-after-in-elasticsearch-8/375233/4 "2025-04-02T09:52:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
