# Get documents starting at a high from

**URL:** <https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714>\
**Category:** Elasticsearch\
**Created:** [March 3, 2022, 8:42am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714 "2022-03-03T08:42:36Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![NominaSumpta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nominasumpta/32/52412_2.png) [@NominaSumpta](https://discuss.elastic.co/u/NominaSumpta)\
**Post date:** [March 3, 2022, 8:42am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/1 "2022-03-03T08:42:37Z")

</div>

Hi,

I have an with ~ 1.500.000 (million) documents or so. I only want to get 1.000 results from it, but I want to start counting backwards. So, I want to retrieve documents 1.499.000 (- 1.000) through 1.500.000. I've set 'from' to 1499000 and 'size' to 1000. from + size is therefore the total amount of documents: 1.500.000. This, expectedly, causes:

```auto
elasticsearch.exceptions.RequestError: RequestError(400, 'search_phase_execution_exception', 'Result window is too large, from + size must be less than or equal to: [10000] but was [1474810]. See the scroll api for a more efficient way to request large data sets. This limit can be set by changing the [index.max_result_window] index level setting.')

```

The scroll API reference @ [Scroll API | Elasticsearch Guide [7.17] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/scroll-api.html) says:

> We no longer recommend using the scroll API for deep pagination. If you need to preserve the index state while paging through more than 10,000 hits, use the search\_after parameter with a point in time (PIT).

Should I also be using PIT for retrieving just a few results, but starting at a high from?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [March 3, 2022, 9:50am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/2 "2022-03-03T09:50:26Z")

</div>

my gut feeling here is, that even though you could switch from a regular query to scroll search/PIT/search\_after ,maybe the query itself could be improved? If you tell more about the use-case that might help.

Could you change the sorting strategy or filtering to retrieve the required documents instead of paginating through them?

---

<div class="post-metadata">

**Author:** ![NominaSumpta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nominasumpta/32/52412_2.png) [@NominaSumpta](https://discuss.elastic.co/u/NominaSumpta)\
**Post date:** [March 3, 2022, 10:05am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/3 "2022-03-03T10:05:25Z")

</div>

Thanks for your reply.

The use case is as follows: I have a sort order and a limit. When the sort order is ascending, I want to inverse the limit. So:

- If I have 10.000 documents, the sort order set to ascending and the limit set to 1, I will get document 9.999
- If I have 10.000 documents, the sort order set to descending and the limit set to 1, I will get document 1

In other words: when the sort order is ascending, `from` is set to document count minus limit and size is set to the limit.

---

<div class="post-metadata">

**Author:** ![NominaSumpta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nominasumpta/32/52412_2.png) [@NominaSumpta](https://discuss.elastic.co/u/NominaSumpta)\
**Post date:** [March 3, 2022, 11:37am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/4 "2022-03-03T11:37:07Z")

</div>

FWIW: when I refer to 'limit', I actually mean the Elasticsearch concept of 'size'.

---

<div class="post-metadata">

**Author:** ![casterQ](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/casterq/32/93257_2.png) [@casterQ](https://discuss.elastic.co/u/casterQ)\
**Post date:** [March 4, 2022, 7:37am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/5 "2022-03-04T07:37:55Z")

</div>

Your question is not very clear，Suppose you have 10000 documents：  
if you want get last 100 docs(9900~10000) sort with asc，  
can you use desc to sort and get Top100(1~100)？

---

<div class="post-metadata">

**Author:** ![NominaSumpta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nominasumpta/32/52412_2.png) [@NominaSumpta](https://discuss.elastic.co/u/NominaSumpta)\
**Post date:** [March 4, 2022, 9:17am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/6 "2022-03-04T09:17:22Z")

</div>

Yes, I could, but they'd be in the wrong order.

E.g. when I want documents 7 and 8 (in that order):

DESC: 8, 7, 6, 5  
ASC: 5, 6, 7, 8

DESC will give me the documents I need, 8 and 7, but in the wrong order.

I could of course drop the `from`, sort DESC and `reverse()` the results in Python ...

---

<div class="post-metadata">

**Author:** ![casterQ](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/casterq/32/93257_2.png) [@casterQ](https://discuss.elastic.co/u/casterQ)\
**Post date:** [March 4, 2022, 9:27am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/7 "2022-03-04T09:27:31Z")

</div>

Oh，I see，I may choose to get the docs and reverse it by myself

---

<div class="post-metadata">

**Author:** ![casterQ](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/casterq/32/93257_2.png) [@casterQ](https://discuss.elastic.co/u/casterQ)\
**Post date:** [March 4, 2022, 9:29am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/8 "2022-03-04T09:29:26Z")

</div>

Because it's expensive to use from+size ，and it's not appropriate to use scroll and PIT for your needs......

---

<div class="post-metadata">

**Author:** ![NominaSumpta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nominasumpta/32/52412_2.png) [@NominaSumpta](https://discuss.elastic.co/u/NominaSumpta)\
**Post date:** [March 4, 2022, 9:52am UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/9 "2022-03-04T09:52:45Z")

</div>

Thanks, that's what I thought (and why I asked the question 🙂 ). I'll stick around for a bit to see if anyone else has any ideas.

---

<div class="post-metadata">

**Author:** ![NominaSumpta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nominasumpta/32/52412_2.png) [@NominaSumpta](https://discuss.elastic.co/u/NominaSumpta)\
**Post date:** [March 5, 2022, 6:21pm UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/10 "2022-03-05T18:21:26Z")

</div>

> I could of course drop the from, sort DESC and reverse() the results in Python ...

I've solved it this way. Thanks for your replies, @casterQ and @spinscale!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 2, 2022, 6:21pm UTC](https://discuss.elastic.co/t/get-documents-starting-at-a-high-from/298714/11 "2022-04-02T18:21:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
