# Result Sets of Aggregation Partitioning

**URL:** <https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117>\
**Category:** Elasticsearch\
**Created:** [May 25, 2017, 11:34am UTC](https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117 "2017-05-25T11:34:58Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![blaineparker](https://avatars.discourse-cdn.com/v4/letter/b/8dc957/32.png) [@blaineparker](https://discuss.elastic.co/u/blaineparker)\
**Post date:** [May 25, 2017, 11:34am UTC](https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117/1 "2017-05-25T11:34:59Z")

</div>

How is partitioning in aggregations achieved? How is the data selected for each partition and what guarantees can be made surrounding the ordering of results in each partition and across partitions?

How can partitioning be used to paginate results if I want to guarantee some ordering of the aggregation results?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 25, 2017, 9:24pm UTC](https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117/2 "2017-05-25T21:24:28Z")

</div>

Are you talking about the partitioning feature provided in the 'include' clause of the 'terms' aggregation?

---

<div class="post-metadata">

**Author:** ![blaineparker](https://avatars.discourse-cdn.com/v4/letter/b/8dc957/32.png) [@blaineparker](https://discuss.elastic.co/u/blaineparker)\
**Post date:** [May 26, 2017, 7:03am UTC](https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117/3 "2017-05-26T07:03:21Z")

</div>

Yeah.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 26, 2017, 8:30am UTC](https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117/4 "2017-05-26T08:30:32Z")

</div>

> [@blaineparker](#):
>
> How is the data selected for each partition

The value of each term e.g. ipaddress `146.204.187.221` is hashed and then we take the modulo of the number of paritions you picked e.g. 20. This means all terms will be evenly assigned between 20 partitions. We use the same technique to route documents evenly to shards based on their doc IDs.  
For each aggregation request you make you pick one of your partitions e.g. partition 7 of 20. This means each shard will be performing analysis on the same subset of terms.

> [@blaineparker](#):
>
> what guarantees can be made surrounding the ordering of results in each partition and across partitions?

The terms within one request (e.g. partition 1 of 20) will be ordered by whatever criteria you pick.  
There is no guarantee of order across partitions because partitions by design are a randomised division of the data.

> [@blaineparker](#):
>
> How can partitioning be used to paginate results if I want to guarantee some ordering of the aggregation results?

You can't. "Some ordering" could be expanded to include something as tricky as userIds sorted by their last logged access date. Given very large numbers of user IDs and a distributed index it is impossible to compute this total ordering without resorting to map/reduce style streaming of masses of data across the network and creating temporary data files etc.  
Using term partitioning however you can reduce the amount of data streamed but have to live with in-partition ordering only. This would be adequate for the scenario of finding userIDs that have expired but will not solve all problems. Another option for the complex scenarios like "last access dates" is to opt for entity-centric indexes based around user.

---

<div class="post-metadata">

**Author:** ![blaineparker](https://avatars.discourse-cdn.com/v4/letter/b/8dc957/32.png) [@blaineparker](https://discuss.elastic.co/u/blaineparker)\
**Post date:** [May 26, 2017, 9:18am UTC](https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117/5 "2017-05-26T09:18:51Z")

</div>

Thank so much. This gives me a lot of clarity.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 23, 2017, 9:19am UTC](https://discuss.elastic.co/t/result-sets-of-aggregation-partitioning/87117/6 "2017-06-23T09:19:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
