# Easy way to get a terms bucket count stats?

**URL:** <https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523>\
**Category:** Elasticsearch\
**Created:** [December 4, 2019, 11:39am UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523 "2019-12-04T11:39:26Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![djmcgreal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/djmcgreal/32/52956_2.png) [@djmcgreal](https://discuss.elastic.co/u/djmcgreal)\
**Post date:** [December 4, 2019, 11:39am UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523/1 "2019-12-04T11:39:27Z")

</div>

Hi,

We're a need to get some extended\_stats on the bucket counts of a terms aggregation, i.e. we want to know the stats (avg, stddev etc) of the frequency of terms. So we do: `terms {...} extended_stats { path:"terms._count" }`, which works. The problem is that there a fairly high number of terms, and we also want them wrapped in a `date_histogram`, so we hit the max buckets limit.

Since I don't actually need the terms in 'buckets', is there a shortcut that avoids hitting the barrier?

Thanks!  
Dan.

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [December 4, 2019, 12:56pm UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523/2 "2019-12-04T12:56:11Z")

</div>

Have you considered using the [cardinality aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-metrics-cardinality-aggregation.html)? This aggregation returns the number of unique values, which should be the same as the total number of buckets from a terms aggregation.

One thing to be aware of is that the cardinality aggregation [returns an approximation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-metrics-cardinality-aggregation.html#_counts_are_approximate) of the unique count. Depending on your use case that may or not may be a problem.

---

<div class="post-metadata">

**Author:** ![djmcgreal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/djmcgreal/32/52956_2.png) [@djmcgreal](https://discuss.elastic.co/u/djmcgreal)\
**Post date:** [December 4, 2019, 12:59pm UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523/3 "2019-12-04T12:59:59Z")

</div>

Hi, thanks Abdon,

I want stats on how many times each term is duplicated. E.g. if at one time bucket, the value 'A' occurs 1000 times, that's different to another time bucket where there are 500 terms, each occurring twice.

Thanks, Dan.

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [December 4, 2019, 1:20pm UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523/4 "2019-12-04T13:20:41Z")

</div>

In that case you may want to look at the [composite aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html). This aggregation allows you to paginate through all the buckets, without hitting the limit.

The maximum number of buckets is a "soft limit" by the way. You could change it with the `search.max_buckets` cluster setting - but be aware that this may cause Elasticsearch to run out of memory.

---

<div class="post-metadata">

**Author:** ![djmcgreal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/djmcgreal/32/52956_2.png) [@djmcgreal](https://discuss.elastic.co/u/djmcgreal)\
**Post date:** [December 4, 2019, 2:59pm UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523/5 "2019-12-04T14:59:57Z")

</div>

Thanks, that leads to a new problem, does my app need to combine the time buckets? Or can that also be done with e.g. a pipeline aggregation in ES?  
Also, I noticed that my auto\_date\_histogram targets 10 buckets, and my terms aggregation is limited to 100 terms (e.g. 10\*100=1000 buckets), but I still get the max buckets problem. Am I missing something?

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [December 6, 2019, 11:29am UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523/6 "2019-12-06T11:29:30Z")

</div>

Unfortunately pipeline aggregations are not supported with composite aggregations. That may change in the future. You can [follow the conversation on this topic here](https://github.com/elastic/elasticsearch/issues/32692), if you're interested.

I'm not sure why you're hitting that max buckets limits with only 1000 buckets.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 3, 2020, 11:37am UTC](https://discuss.elastic.co/t/easy-way-to-get-a-terms-bucket-count-stats/210523/7 "2020-01-03T11:37:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
