# Control number of buckets created in an aggregation

**URL:** <https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360>\
**Category:** Elasticsearch\
**Created:** [August 8, 2019, 6:32am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360 "2019-08-08T06:32:35Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ahmad\_Mozafarnia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ahmad_mozafarnia/32/45836_2.png) [@Ahmad\_Mozafarnia](https://discuss.elastic.co/u/Ahmad_Mozafarnia)\
**Post date:** [August 8, 2019, 6:32am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/1 "2019-08-08T06:32:35Z")

</div>

This question was originally posted on StackOverflow:

> <https://stackoverflow.com/questions/57393548/control-number-of-buckets-created-in-an-aggregation>

I'm having trouble with controlling the number of buckets an aggregation is going to create.

This is a simple aggregation containing no nested or sub aggregations on an index having roughly `80000` documents:

```
GET /my_index/_search
{
   "size":0,
   "query":{
      "match_all":{}
   },
   "aggregations":{
      "unique":{
         "terms":{
            "field":"_id",
            "size":<NUM_TERM_BUCKETS>
         }
      }
   }
}

```

If I set the `<NUM_TERM_BUCKETS>` to `7000`, I get this error response in **ES 7.3** :

```
{
   "error":{
      "root_cause":[
         {
            "type":"too_many_buckets_exception",
            "reason":"Trying to create too many buckets. Must be less than or equal to: [10000] but was [10001]. This limit can be set by changing the [search.max_buckets] cluster level setting.",
            "max_buckets":10000
         }
      ],
      "type":"search_phase_execution_exception",
      "reason":"all shards failed",
      "phase":"query",
      "grouped":true,
      "failed_shards":[
         {
            "shard":0,
            "index":"my_index",
            "node":"XYZ",
            "reason":{
               "type":"too_many_buckets_exception",
               "reason":"Trying to create too many buckets. Must be less than or equal to: [10000] but was [10001]. This limit can be set by changing the [search.max_buckets] cluster level setting.",
               "max_buckets":10000
            }
         }
      ]
   },
   "status":503
}

```

And it runs successfully if I decrease the `<NUM_TERM_BUCKETS>` to `6000`.

I'm really confused. how on earth this aggregation creates more than `10000` buckets?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [August 8, 2019, 7:08am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/2 "2019-08-08T07:08:24Z")

</div>

> [@Ahmad\_Mozafarnia](#):
>
> how on earth this aggregation creates more than `10000` buckets?

To address issues of accuracy in a distributed system elasticsearch asks for a number higher than ‘size’ from each shard (see ‘shard\_size’ setting).  
If you have a lot of unique values, making multiple requests using the ‘composite’ aggregation is probably a better way to go

---

<div class="post-metadata">

**Author:** ![Ahmad\_Mozafarnia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ahmad_mozafarnia/32/45836_2.png) [@Ahmad\_Mozafarnia](https://discuss.elastic.co/u/Ahmad_Mozafarnia)\
**Post date:** [August 8, 2019, 7:23am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/3 "2019-08-08T07:23:59Z")

</div>

Is this a documented behavior?

---

<div class="post-metadata">

**Author:** ![Ahmad\_Mozafarnia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ahmad_mozafarnia/32/45836_2.png) [@Ahmad\_Mozafarnia](https://discuss.elastic.co/u/Ahmad_Mozafarnia)\
**Post date:** [August 8, 2019, 7:35am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/4 "2019-08-08T07:35:03Z")

</div>

> [@Mark\_Harwood](#):
>
> To address issues of accuracy in a distributed system elasticsearch asks for a number higher than ‘size’ from each shard (see ‘shard\_size’ setting).

Here's the response of `GET /my_index/_settings`:

```
{
  "my_index" : {
    "settings" : {
      "index" : {
        "creation_date" : "1564578971559",
        "number_of_shards" : "1",
        "number_of_replicas" : "1",
        "uuid" : "KXlMMbZpT-yX8AaQK1FI7w",
        "version" : {
          "created" : "6080299",
          "upgraded" : "7030099"
        },
        "provided_name" : "my_index"
      }
    }
  }
}

```

i.e. `my_index` has only one shard.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [August 8, 2019, 8:07am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/5 "2019-08-08T08:07:46Z")

</div>

Works OK on 6.x - I'll need to dig deeper into what changed in 7.x.  
Looks like the `shard_size` multiplier is still in effect for a single-sharded system.  
The good news is if you also set `shard_size` to 8000 it should work OK (does for me here).

FYI - not sure if your example with "unique" agg on "\_id" field is representative but if you just want to count docs this will be much cheaper:

```
GET my_index/_count
```

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [August 8, 2019, 8:17am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/6 "2019-08-08T08:17:47Z")

</div>

Replying to myself - this `shard_size = size` optimisation for single-sharded systems changed when we introduced cross-cluster search. A cluster could no longer know the total number of shards in a request so this RAM-saving optimisation [was removed](https://github.com/elastic/elasticsearch/commit/42ea6449033d82902a86285ae6db84412c8c7f35#diff-df6fb503e931ead1f7a026975e9d9797)

---

<div class="post-metadata">

**Author:** ![Ahmad\_Mozafarnia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ahmad_mozafarnia/32/45836_2.png) [@Ahmad\_Mozafarnia](https://discuss.elastic.co/u/Ahmad_Mozafarnia)\
**Post date:** [August 8, 2019, 8:32am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/7 "2019-08-08T08:32:05Z")

</div>

Thanks. it was helpful.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 5, 2019, 8:32am UTC](https://discuss.elastic.co/t/control-number-of-buckets-created-in-an-aggregation/194360/8 "2019-09-05T08:32:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
