# Large aggregate (too\_many\_buckets\_exception)

**URL:** <https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091>\
**Category:** Elasticsearch\
**Created:** [July 5, 2019, 11:51am UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091 "2019-07-05T11:51:17Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 5, 2019, 11:51am UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/1 "2019-07-05T11:51:17Z")

</div>

I have an index with millions of rows, most of the rows contain a hash value (md5)  
I want to group by the hashed value and calculate the count of documents per hash and then sum the total count. This only for buckets with at least 2 documents.

I do this using Kibana and Elasticsearch (7.1). I got this working but for this particular set I have more then 800K of group by results (buckets) so Elasticsearch runs into a too\_many\_buckets\_exception.

I know I can increase the max\_bucket value but as far as I found out this is something you shouldn't do. Also in the future the 800K may easily become 2 MIL buckets or higher.

How can I get this metric witouth having to increase the max\_bucket value? For me, used to SQL, this seems like a relatively easy question.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 5, 2019, 12:38pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/2 "2019-07-05T12:38:37Z")

</div>

Hi Roel,

> [@rvanegmond](#):
>
> For me, used to SQL, this seems like a relatively easy question.

If you're used to putting everything on one machine then life certainly is easy. In a distributed system like most elasticsearch deployments, you have to deal with what we call the [FAB conundrum](https://discuss.elastic.co/t/background-count-in-significant-terms-not-consistent/55824/8) which means you have to pick a trade-off.

The trade offs come from physical limits (speeds of networks, RAM limitations etc) and there are various options. Try run [this wizard](https://plnkr.co/edit/iJSFP8eRrhC7l7Hx2XOL?p=preview) and see which option it leads you to and we can discuss further here.

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 5, 2019, 12:52pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/3 "2019-07-05T12:52:51Z")

</div>

Hi Mark,

Thanks for your quick response and the great wizard. I seemed that I was already looking towards the right direction as I have read about partitioning and composite aggregation. However I can't figure out how do this in Kibana. Btw, this wizard pointed me to partitioning.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 5, 2019, 1:18pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/4 "2019-07-05T13:18:24Z")

</div>

Composite will only ever give you buckets in the order of a key (but the key can be made from multiple values in a single doc).  
If your required sort order is unrelated to the key - eg a different value like sum, average or count computed from multiple docs then life gets more complex because the related docs may be on many different machines and there may be many grouping keys to consider (too many to fit comfortably into a response's RAM). At this point your options become:

1. Look at a common subset of the distributed keys in each request (term partitioning) or
2. Put related data close to each other:  
a) Use "routing" to ensure all the same-keyed docs end up on the same machine  
b) Put all related info into the same doc ("entity centric indexing")

Can you say more about why you need to page through _all_ of the results? As you can guess, it's not a straight-forward problem for any system.

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 5, 2019, 2:00pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/5 "2019-07-05T14:00:32Z")

</div>

I'm not entirely sure about what you mean, but I'll try to explain.

I have a set of documens with an md5 hash for each document. I want to calculate the total amount of documents which match another document. So I do a terms aggregation on the hash, count the documents per hash and sum on the count to get the total. But this results in the to many buckets exception.

So I guess to awnser your question to through all is this case is required to get a total..?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 5, 2019, 2:03pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/6 "2019-07-05T14:03:23Z")

</div>

It might help to "zoom out" to the business problem you're trying to solve.

For example - if you were just trying to avoid duplicate docs maybe using the MD5 as the doc ID would be a straightforward way to avoid duplicates?

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 5, 2019, 2:11pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/7 "2019-07-05T14:11:24Z")

</div>

Good point. Well in this case it is actually to find and report on those duplicates. So hence the difficulty. But the partitioning seems like the awnser because it would mean to keep summing until I have the total number or am I wrong?

And would this be possible in Kibana?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 5, 2019, 3:15pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/8 "2019-07-05T15:15:05Z")

</div>

Term partitioning is essentially a way to limit RAM when glueing results together from many machines to derive things like counts or averages etc for some keys.

If the set of keys to consider is too large for one single request then you need to find a way to break that set of keys into multiple partitions that can be considered in different requests.  
You can think of it like processing just the MD5s that start with the letter "A", then running another request to do the ones that start with "B" and so on. Each result should have a sufficiently small subset of keys that you can gather and all the related data from machines and not run out of RAM when fusing a response. Rather than using the first letter of a key to partition into groups though we compute the hash of a key and modulo the number of required partitions to see if the key lands in the required partition.

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 5, 2019, 3:28pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/9 "2019-07-05T15:28:00Z")

</div>

Thanks again Mark makes sense, it is like a paging meganisme right? But is there anyway to make Kibana do this. To fetch the first X buckets, calculate the metrics and next X buckets calculate the buckets and so on and in the end sum the totals of the requests?

A little additional info:

I just tried answering the same question using the same set it Mongo. This also generates a very large set but I can enable allowDiskUsage to calculate the number in the end.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 5, 2019, 3:36pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/10 "2019-07-05T15:36:32Z")

</div>

> [@rvanegmond](#):
>
> Thanks again Mark makes sense, it is like a paging meganisme right?

KInd of - within a partition (like the MD5's starting with "A") you get the required arbitrary sort order.  
However, there's no global sort order to the results. Partitions aren't sorted. You just have N arbitrarily decided subdivisions of your data and within each of them they have a logically ordered subset of results.

> [@rvanegmond](#):
>
> is there anyway to make Kibana do this.

As far as I know, no. If you fused the data at index time ie created entity-centric single docs, one for each unique MD5 with a count of related docs on them Kibana could work happily with those. The new "dataframes" feature in 7.2 might be able to help with that.

> [@rvanegmond](#):
>
> I just tried answering the same question using the same set it Mongo. This also generates a very large set but I can enable allowDiskUsage to calculate the number in the end.

Interesting. Spilling intermediate results to disk is not a strategy we've reached for yet - the emphasis to date has been on doing things fast in the constraints of RAM.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 5, 2019, 4:04pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/11 "2019-07-05T16:04:49Z")

</div>

As you have multiple documents having the same hash it could possibly be a lot faster and easier for you if you could reindex the data into a new index and use the hash as a routing key. This will make sure that all documents with a specific hash end up in the same shard, which means you can perform the aggregation accurately at the shard level.

---

<div class="post-metadata">

**Author:** ![amyc](https://avatars.discourse-cdn.com/v4/letter/a/ea5d25/32.png) [@amyc](https://discuss.elastic.co/u/amyc)\
**Post date:** [July 8, 2019, 2:57pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/12 "2019-07-08T14:57:22Z")

</div>

You may try this in Kibana:  
GET /{index\_name}/\_search  
{  
"query": {  
"match\_all": {}  
},  
"track\_total\_hits" : true,  
"collapse": {  
"field": "{hash\_value\_field\_name}"  
},  
"aggregations": {  
"unique\_hash\_value\_list": {  
"cardinality": {  
"field": "{hash\_value\_field\_name}"  
}  
}  
}  
}

The result of total/hits/value and aggregations/unique\_hash\_value\_list/value are total of all and total of unique hash.

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 9, 2019, 7:14am UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/13 "2019-07-09T07:14:45Z")

</div>

At this point, because there are not so many documents I'm running with the out of the box settings. Meaning (I thought) one shard. Even when changing this I think I run into the issue of the 10K limit of buckets. Or is it ok to change this if hardware resources are sufficiënt?

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 9, 2019, 7:57am UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/14 "2019-07-09T07:57:59Z")

</div>

hi Yifeng,

Thanks for your response.

I just run this query but it doesn't give the expacted results. It gives the count of the unique hashes as where I need the sum of the count per hash.

Also I don't think there is a way to collapse using Kibana, but I might be wrong.

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 10, 2019, 6:59am UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/15 "2019-07-10T06:59:39Z")

</div>

Hi Mark,

Thanks again. It would be nice if Kibana would support this some how. At this point I think data fusion is one of the best options.

In addition to this I do wonder why the query below gives an to many buckets error as it limited to 10000 which is equal to the max.

```
{
"aggs": {
"2": {
  "terms": {
    "field": "name",
    "order": {
        "_count": "desc"
      },
        "size": 10000,
        "min_doc_count": 2
   }
  }
 }
}
```

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 10, 2019, 8:07am UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/16 "2019-07-10T08:07:09Z")

</div>

Check out [shard\_size](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#_shard_size_3)  
By default each shard is asked for more than `size` number of terms in order to improve accuracy

---

<div class="post-metadata">

**Author:** ![rvanegmond](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@rvanegmond](https://discuss.elastic.co/u/rvanegmond)\
**Post date:** [July 12, 2019, 2:31pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/17 "2019-07-12T14:31:15Z")

</div>

Hi Mark,

I just tried the data frame feature and this gives the expected results. I does make me wonder why this doesn't run in to the bucket exception. I guess it splits things up but that might be a possiblity as well when using a datatable visual you don't see all the data at once so you use the pages anyways.

It might be something to implement?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 12, 2019, 5:10pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/18 "2019-07-12T17:10:33Z")

</div>

> [@rvanegmond](#):
>
> I does make me wonder why this doesn't run in to the bucket exception

As you suggest, it does multiple requests using the ‘composite’ aggregation to group data under a key. Because it works through all keys in their natural sort order it does not have to worry about sorting by some value derived from multiple docs eg a cardinality count of some other field.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 9, 2019, 5:10pm UTC](https://discuss.elastic.co/t/large-aggregate-too-many-buckets-exception/189091/19 "2019-08-09T17:10:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
