# Handle big result set?

**URL:** <https://discuss.elastic.co/t/handle-big-result-set/25010>\
**Category:** Elasticsearch\
**Created:** [July 7, 2015, 6:16am UTC](https://discuss.elastic.co/t/handle-big-result-set/25010 "2015-07-07T06:16:16Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![linlma](https://avatars.discourse-cdn.com/v4/letter/l/ec9cab/32.png) [@linlma](https://discuss.elastic.co/u/linlma)\
**Post date:** [July 7, 2015, 6:16am UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/1 "2015-07-07T06:16:16Z")

</div>

Hello Elastic experts,

Suppose a query matches a large volumes of records (saying a few million records), what is the best way to handle (I want to store the results on local disk) the results? Is there a way to streaming big result set as I have the concern the local box memory may not be able to hold all result set?

thanks in advance,  
Lin

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [July 7, 2015, 8:18am UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/2 "2015-07-07T08:18:56Z")

</div>

Deep pagination can be achieved using the Scan-Scroll feature: [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#scroll-scan](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#scroll-scan)

---

<div class="post-metadata">

**Author:** ![linlma](https://avatars.discourse-cdn.com/v4/letter/l/ec9cab/32.png) [@linlma](https://discuss.elastic.co/u/linlma)\
**Post date:** [July 7, 2015, 8:44am UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/3 "2015-07-07T08:44:57Z")

</div>

Hi Colin,

Good sharing. Looked through the document you referred and find sorting may have cost. My use case is, I just need to find top N results, and do not care the order of results in top N. Wondering in my case, what is the most efficient way to write the query?

regards,  
Lin

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [July 7, 2015, 8:54am UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/4 "2015-07-07T08:54:14Z")

</div>

> [@linlma](#):
>
> I just need to find top N results, and do not care the order of results in top N

What do you mean by this? to define the top N of something there has to be some kind of sorting. How are you defining the top N? do you want the top N scoring documents?

---

<div class="post-metadata">

**Author:** ![linlma](https://avatars.discourse-cdn.com/v4/letter/l/ec9cab/32.png) [@linlma](https://discuss.elastic.co/u/linlma)\
**Post date:** [July 7, 2015, 5:21pm UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/5 "2015-07-07T17:21:09Z")

</div>

Yes, Colin, yes, I need top N scored documents, you are correct. I mean I do not need to strict ascending/descending order sort inside the top N documents, as long as top N documents are returned. Any efficient way to implement? Thanks.

regards,  
Lin

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [July 7, 2015, 9:06pm UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/6 "2015-07-07T21:06:22Z")

</div>

Then the scan-scroll feature is what you want, Just don't set any explicit sorting in the scan request

---

<div class="post-metadata">

**Author:** ![linlma](https://avatars.discourse-cdn.com/v4/letter/l/ec9cab/32.png) [@linlma](https://discuss.elastic.co/u/linlma)\
**Post date:** [July 7, 2015, 9:13pm UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/7 "2015-07-07T21:13:52Z")

</div>

Thanks Colin,

I read the document ([https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#scroll-scan](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html#scroll-scan)), for statements, "Deep pagination with from and size — e.g. ?size=10&from=10000 — is very inefficient as (in this example) 100,000 sorted results have to be retrieved from each shard and resorted in order to return just 10 results.", I am confused, should it be 10,000 sorted results? Other than 100,000? Which maps to from =10000 parameter?

Please feel free to correct me if I am wrong.

BTW, another quick question is, if I want to use scroll only without scan ([https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html)), could I combine sorting with scroll? And why using scroll is more efficient than ordinary queries?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:03am UTC](https://discuss.elastic.co/t/handle-big-result-set/25010/8 "2017-07-06T00:03:01Z")

</div>


