# Inconsistent results when search\_after is used for pagination sorted by score and id

**URL:** <https://discuss.elastic.co/t/inconsistent-results-when-search-after-is-used-for-pagination-sorted-by-score-and-id/286554>\
**Category:** Elasticsearch\
**Created:** [October 12, 2021, 10:54pm UTC](https://discuss.elastic.co/t/inconsistent-results-when-search-after-is-used-for-pagination-sorted-by-score-and-id/286554 "2021-10-12T22:54:07Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![vjgorla](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vjgorla/32/95775_2.png) [@vjgorla](https://discuss.elastic.co/u/vjgorla)\
**Post date:** [October 12, 2021, 10:54pm UTC](https://discuss.elastic.co/t/inconsistent-results-when-search-after-is-used-for-pagination-sorted-by-score-and-id/286554/1 "2021-10-12T22:54:08Z")

</div>

We have implemented pagination using search\_after and sorting the results by \_score and a unique id field as a tie-breaker. However sometimes we are getting duplicate results across pages, and other times matches do not appear in any of the pages.

For example, when there are 65 total hits and paginated using page size of 10, the last page has 6 results instead of 5. A document appears in both page 5 and page 6.

Another example, query matches 13,552 documents with the exact same score but total results from all pages is only 5,900.

This is happening consistently even when the index hasn't been updated between requests. Curiously this happens when a replica shard has a different size than the primary shard but with the same number of documents. Recreating the replica fixes the problem and result count matches up.

This post suggests score can be used in search\_after

> [@Using search\_after for pagination where you sort by score](https://discuss.elastic.co/t/using-search-after-for-pagination-where-you-sort-by-score/146233):
>
> In my query I sort by score, so the most relevant results are first. I'd like to be able to paginate this, and search\_after is apparently the best way for my use case. But search\_after doesn't seem to let you use \_score as a sort field. I'd need to sort by something like \_id or whatever which is not good, because then in the initial search the results are ordered by id and not relevance. What's the correct way to use search\_after for pagination while keeping the results ordered by relevance?

We are using elastic 7.13.2 in a 3 node cluster with 5 shards and a replication factor of 1.

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [October 12, 2021, 11:56pm UTC](https://discuss.elastic.co/t/inconsistent-results-when-search-after-is-used-for-pagination-sorted-by-score-and-id/286554/2 "2021-10-12T23:56:27Z")

</div>

If your unique field is indeed unique it should not happen. It means that there were index updates in between search requests, or possibly they were not propagated to the replica.

The recommended way to do `search_after` in newer versions (including 7.13) is to use [point in time](https://www.elastic.co/guide/en/elasticsearch/reference/current/paginate-search-results.html#search-after) search. In this case you don't need to provide your unique tie-breaker field, it will be provided automatically and is called `_shard_doc`.

---

<div class="post-metadata">

**Author:** ![vjgorla](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vjgorla/32/95775_2.png) [@vjgorla](https://discuss.elastic.co/u/vjgorla)\
**Post date:** [October 14, 2021, 7:34am UTC](https://discuss.elastic.co/t/inconsistent-results-when-search-after-is-used-for-pagination-sorted-by-score-and-id/286554/3 "2021-10-14T07:34:58Z")

</div>

Thanks for your answer. We are integrating search\_after query to our UI pagination. If we use PIT and keep the search context alive between UI page fetches, which could be minutes depending on user think time, would that scale well?. We have to support at least a few hundred concurrent users.

Also how would that affect other searches and background ingestion process?. Documentation says lucene segment merging is impacted by open PIT contexts. We have a fairly rapidly changing large index.

---

<div class="post-metadata">

**Author:** ![mayya](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mayya/32/83147_2.png) [@mayya](https://discuss.elastic.co/u/mayya)\
**Post date:** [October 26, 2021, 7:14pm UTC](https://discuss.elastic.co/t/inconsistent-results-when-search-after-is-used-for-pagination-sorted-by-score-and-id/286554/4 "2021-10-26T19:14:28Z")

</div>

For scroll requests we have a limitation for the max number of open scroll context of 500, because PIT contexts are much more lightweight, we don’t have any limit on the number of PIT contexts, so you can open as many PIT contexts as possible. We probably need to introduce some limitation though. In the worst case scenario, when you constantly open PIT contexts with a very long `keep_alive` parameter and constantly update your indices, you may ran out of file descriptors or heap memory, because as you rightly noticed segments used by PIT contexts are being kept and not being deleted by merge.

On other hand, if you use relatively small `keep_alive` , say 10-15 mins, use high enough `refresh_interval` not to create many segments, and regularly monitor the number of PIT contexts with `GET /_nodes/stats/indices/search` than probably it will work fine.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 23, 2021, 7:15pm UTC](https://discuss.elastic.co/t/inconsistent-results-when-search-after-is-used-for-pagination-sorted-by-score-and-id/286554/5 "2021-11-23T19:15:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
