# Terms aggregation with a limit

**URL:** <https://discuss.elastic.co/t/terms-aggregation-with-a-limit/17773>\
**Category:** Elasticsearch\
**Created:** [May 28, 2014, 9:54am UTC](https://discuss.elastic.co/t/terms-aggregation-with-a-limit/17773 "2014-05-28T09:54:21Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Guillermo\_Arias\_del\_](https://avatars.discourse-cdn.com/v4/letter/g/cab0a1/32.png) [@Guillermo\_Arias\_del\_](https://discuss.elastic.co/u/Guillermo_Arias_del_)\
**Post date:** [May 28, 2014, 9:54am UTC](https://discuss.elastic.co/t/terms-aggregation-with-a-limit/17773/1 "2014-05-28T09:54:21Z")

</div>

Hi!

I am using a terms aggregation to get the 10 best terms that match a query.  
The problem is that since I am performing a query that returns a lot of  
documents, the number of distinct terms is very big, as it is the number of  
documents per bucket. An example would be:

{  
"aggregations" : {  
... // filters  
"not\_exact" : {  
"doc\_count" : 2257428,  
"text" : {  
"buckets" : [ {  
"key" : "abb",  
"doc\_count" : 135686  
}, {  
"key" : "ansprache",  
"doc\_count" : 118570  
}, {  
"key" : "aus",  
"doc\_count" : 106023  
}, {  
"key" : "auf",  
"doc\_count" : 74338  
}, {  
"key" : "archiv",  
"doc\_count" : 54315  
}, {  
"key" : "außen",  
"doc\_count" : 52444  
}, {  
"key" : "am",  
"doc\_count" : 52178  
}, {  
"key" : "ab",  
"doc\_count" : 45723  
}, {  
"key" : "an",  
"doc\_count" : 44656  
}, {  
"key" : "athen",  
"doc\_count" : 32070  
} ]  
},  
...  
}

I am not interested in the actual number of documents, and I would even be  
willing to sacrifice precision if I can speed up the query (which now takes  
6 seconds), so my question is: is there a way to tell the terms aggregation  
to stop counting at a certain limit? Imagine, for instance I could specify  
this value to be 50 000. I could get the top buckets in the wrong order,  
but I could live with that. Elasticsearch would take less time, I suppose.  
I would be even happy if the limit was set to 10 000 and I would end up  
with different keys, because as the user specifies more, values above 10  
000 become less and less probable.

And if that is possible, the next question would be: is there a way of  
making this value dependable on the number of total matches (in that case 2  
257 428)?

For those who are interested: the background of this request is an  
autocompletion request that matches against different fields, with ngram or  
edge\_ngram depending on the field type, and returns the best matches. In  
the example above, the user types "a" and gets those results. If you are  
thinking about caching results, this is generally not possible, since the  
query is constrained by document types and fields; data chages; and  
finally, each user can have a different view on it (so, a user may not see  
a any documents containing "athen" and then that match would be incorrect).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/1278cc4c-83b6-4019-b7bf-17a1aae45e0a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/1278cc4c-83b6-4019-b7bf-17a1aae45e0a%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [May 28, 2014, 3:21pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-a-limit/17773/2 "2014-05-28T15:21:15Z")

</div>

Unfortunately, I don't think stopping incrementing counts after a certain  
limit would improve response times, given that most time is not spent  
incrementing this counter but reading the hash table to figure out whether  
the current term is new or has already been seen. Maybe you can specify a  
timeout (  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/search-request-body.html#_parameters_4))  
on your requests? This will return partial results in case the timeout is  
exceeded.

It is not possible to configure aggregations depending on the number of  
matches because the number of matches is computed in parallel with  
aggregations.

On Wed, May 28, 2014 at 11:54 AM, Guillermo Arias del Río \<  
[ariasdelrio@gmail.com](mailto:ariasdelrio@gmail.com)\> wrote:

> Hi!
> 
> I am using a terms aggregation to get the 10 best terms that match a  
> query. The problem is that since I am performing a query that returns a lot  
> of documents, the number of distinct terms is very big, as it is the number  
> of documents per bucket. An example would be:
> 
> {  
> "aggregations" : {  
> ... // filters  
> "not\_exact" : {  
> "doc\_count" : 2257428,  
> "text" : {  
> "buckets" : [ {  
> "key" : "abb",  
> "doc\_count" : 135686  
> }, {  
> "key" : "ansprache",  
> "doc\_count" : 118570  
> }, {  
> "key" : "aus",  
> "doc\_count" : 106023  
> }, {  
> "key" : "auf",  
> "doc\_count" : 74338  
> }, {  
> "key" : "archiv",  
> "doc\_count" : 54315  
> }, {  
> "key" : "außen",  
> "doc\_count" : 52444  
> }, {  
> "key" : "am",  
> "doc\_count" : 52178  
> }, {  
> "key" : "ab",  
> "doc\_count" : 45723  
> }, {  
> "key" : "an",  
> "doc\_count" : 44656  
> }, {  
> "key" : "athen",  
> "doc\_count" : 32070  
> } ]  
> },  
> ...  
> }
> 
> I am not interested in the actual number of documents, and I would even be  
> willing to sacrifice precision if I can speed up the query (which now takes  
> 6 seconds), so my question is: is there a way to tell the terms aggregation  
> to stop counting at a certain limit? Imagine, for instance I could specify  
> this value to be 50 000. I could get the top buckets in the wrong order,  
> but I could live with that. Elasticsearch would take less time, I suppose.  
> I would be even happy if the limit was set to 10 000 and I would end up  
> with different keys, because as the user specifies more, values above 10  
> 000 become less and less probable.
> 
> And if that is possible, the next question would be: is there a way of  
> making this value dependable on the number of total matches (in that case 2  
> 257 428)?
> 
> For those who are interested: the background of this request is an  
> autocompletion request that matches against different fields, with ngram or  
> edge\_ngram depending on the field type, and returns the best matches. In  
> the example above, the user types "a" and gets those results. If you are  
> thinking about caching results, this is generally not possible, since the  
> query is constrained by document types and fields; data chages; and  
> finally, each user can have a different view on it (so, a user may not see  
> a any documents containing "athen" and then that match would be incorrect).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/1278cc4c-83b6-4019-b7bf-17a1aae45e0a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/1278cc4c-83b6-4019-b7bf-17a1aae45e0a%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/1278cc4c-83b6-4019-b7bf-17a1aae45e0a%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/1278cc4c-83b6-4019-b7bf-17a1aae45e0a%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAL6Z4j6jbmNsi\_%2BvQdvtvY%2B9p8K7EjNN5URqf0qi3u4eEQkUPQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAL6Z4j6jbmNsi_%2BvQdvtvY%2B9p8K7EjNN5URqf0qi3u4eEQkUPQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:26am UTC](https://discuss.elastic.co/t/terms-aggregation-with-a-limit/17773/3 "2017-07-06T01:26:17Z")

</div>


