# How are the documents selected by the Sampler Aggregation

**URL:** <https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420>\
**Category:** Elasticsearch\
**Created:** [April 5, 2016, 3:22pm UTC](https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420 "2016-04-05T15:22:27Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dominik\_Stadler](https://avatars.discourse-cdn.com/v4/letter/d/8e8cbc/32.png) [@Dominik\_Stadler](https://discuss.elastic.co/u/Dominik_Stadler)\
**Post date:** [April 5, 2016, 3:22pm UTC](https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420/1 "2016-04-05T15:22:27Z")

</div>

Hi,

we are looking at using the [https://www.elastic.co/guide/en/elasticsearch/reference/2.1/search-aggregations-bucket-sampler-aggregation.html](https://www.elastic.co/guide/en/elasticsearch/reference/2.1/search-aggregations-bucket-sampler-aggregation.html) for limiting the run-time of aggregations on large data sets.

What we could not find in the documentation is how the documents are selected? Are they sorted in some way before being sampled or is the order of documents and thus the actual documents used for the aggregation deterministic in any way?

We would like to ensure that the samples are selected in a good random order so that things like a TopX aggregation still returns useful information with high probability.

Thanks... Dominik.

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [April 6, 2016, 7:23am UTC](https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420/2 "2016-04-06T07:23:51Z")

</div>

For the `sampler` aggregation the documents selected for the sample are the N top scoring documents (where score is defined by the query) from each shard (where N is the sample size). The `diversified_sampler` aggregation also selects the N top scoring documents but limits the sample to only X documents for each value of the field you select for diversification.

Hope that helps

---

<div class="post-metadata">

**Author:** ![Dominik\_Stadler](https://avatars.discourse-cdn.com/v4/letter/d/8e8cbc/32.png) [@Dominik\_Stadler](https://discuss.elastic.co/u/Dominik_Stadler)\
**Post date:** [April 6, 2016, 10:59am UTC](https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420/3 "2016-04-06T10:59:27Z")

</div>

Thanks for the explanation, does that mean I could use something like [https://www.elastic.co/guide/en/elasticsearch/guide/current/random-scoring.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/random-scoring.html) to ensure a pseudo-random selection based on the seed that I pass in?

I.e. with the same seed I get the same set sampled and with pure-random seed I get a good random distribution of sampled documents?

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [April 6, 2016, 12:12pm UTC](https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420/4 "2016-04-06T12:12:26Z")

</div>

Yes you should be able to do that, although I haven't tested that out myself

---

<div class="post-metadata">

**Author:** ![Dominik\_Stadler](https://avatars.discourse-cdn.com/v4/letter/d/8e8cbc/32.png) [@Dominik\_Stadler](https://discuss.elastic.co/u/Dominik_Stadler)\
**Post date:** [April 6, 2016, 12:24pm UTC](https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420/5 "2016-04-06T12:24:14Z")

</div>

I quickly tested it, seems to work, however it has quite a performance impact and so invalidates the benefit that we tried to get from the SamplingAggregation in the first place as it then takes at least as long as doing a full aggregation without sampling anyway ☹

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:01pm UTC](https://discuss.elastic.co/t/how-are-the-documents-selected-by-the-sampler-aggregation/46420/6 "2017-07-05T23:01:54Z")

</div>


