# Query on aggregation and scroll

**URL:** <https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416>\
**Category:** Elasticsearch\
**Created:** [December 11, 2018, 5:25pm UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416 "2018-12-11T17:25:20Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![sanjeebkdeka](https://avatars.discourse-cdn.com/v4/letter/s/b782af/32.png) [@sanjeebkdeka](https://discuss.elastic.co/u/sanjeebkdeka)\
**Post date:** [December 11, 2018, 5:25pm UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416/1 "2018-12-11T17:25:21Z")

</div>

Hi,

I have a requirement to retrieve all instances of records having unique value for a particular column. All the records having that same unique value must appear in one cluster. The number of records in the index could be in billions.

Should i be using scroll with aggregation? I somewhere read aggregation is not the best solution for this one.

The other approach could be to scroll over those records and sort on that particular column. For this approach, i wanted to know whether the sorting will be over the 10000 records to be presented or all the matching records will be sorted first and then 10000 records will presented.

Please suggest.

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [December 13, 2018, 9:46am UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416/2 "2018-12-13T09:46:18Z")

</div>

Hi,

Can you explain more the requirement ?

> [@sanjeebkdeka](#):
>
> must appear in one cluster

Do you need to scroll in order to extract those data ?

bye  
Xavier

---

<div class="post-metadata">

**Author:** ![sanjeebkdeka](https://avatars.discourse-cdn.com/v4/letter/s/b782af/32.png) [@sanjeebkdeka](https://discuss.elastic.co/u/sanjeebkdeka)\
**Post date:** [December 13, 2018, 10:17am UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416/3 "2018-12-13T10:17:42Z")

</div>

Yes, I need to scroll to extract those data.

Thanks

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [December 13, 2018, 10:30am UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416/4 "2018-12-13T10:30:08Z")

</div>

> [@sanjeebkdeka](#):
>
> The number of records in the index could be in billions.

You can run a filtered query on the particular column and then scroll results, but it can be long and heavy, depending on the number of "selected documents".  
Note that sorting sorts over all records.

Extra question: Do you need to aggregate selected documents or not ?

---

<div class="post-metadata">

**Author:** ![sanjeebkdeka](https://avatars.discourse-cdn.com/v4/letter/s/b782af/32.png) [@sanjeebkdeka](https://discuss.elastic.co/u/sanjeebkdeka)\
**Post date:** [December 13, 2018, 11:16am UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416/5 "2018-12-13T11:16:53Z")

</div>

Yes. I need aggregation over records matching my query.

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [December 13, 2018, 1:38pm UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416/6 "2018-12-13T13:38:46Z")

</div>

Recently I "solved" a search dilemn with the following trick: We have complex searches with various parameters and they can return lot of records or not. My trick is to run the query with size = 0 to get the totalHits and then run a query with complex aggregations or to scroll and doing aggregations in our code.  
The limit is fixed to 30000 records. Above this limit I let ES do aggregations, under this limit it's very quicker to do it ourself. Running queries with size=0 is super fast (\< 100ms), but running the same query with complex aggregrations can take over 6/7s !

Note that we are running ES 2.4.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 10, 2019, 1:49pm UTC](https://discuss.elastic.co/t/query-on-aggregation-and-scroll/160416/7 "2019-01-10T13:49:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
