# How to avoid or ease terms aggregation for fields with high cardinality?

**URL:** <https://discuss.elastic.co/t/how-to-avoid-or-ease-terms-aggregation-for-fields-with-high-cardinality/296853>\
**Category:** Elasticsearch\
**Created:** [February 10, 2022, 2:12pm UTC](https://discuss.elastic.co/t/how-to-avoid-or-ease-terms-aggregation-for-fields-with-high-cardinality/296853 "2022-02-10T14:12:30Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![asupranovich](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/asupranovich/32/101652_2.png) [@asupranovich](https://discuss.elastic.co/u/asupranovich)\
**Post date:** [February 10, 2022, 2:12pm UTC](https://discuss.elastic.co/t/how-to-avoid-or-ease-terms-aggregation-for-fields-with-high-cardinality/296853/1 "2022-02-10T14:12:30Z")

</div>

Hi community!

My Elasticsearch version is 7.7.0. I'm facing a `too many buckets` issue for my aggregation query.  
**Task description:**  
Let's say we have a car index, car data model looks like below:

```auto
{
    "vin": "08c00679-8a58-11ec-9745-1aed518f8637",
    "make": "Lexus",
    "model": "GS",
    "country": "US",
    "state": "CA",
    "city": "San-Diego",
    "color": "black",
    // some other fields
    "dates": [
      {
        "type": "manufacture",
        "date": "2010-01-01"
      },
      {
        "type": "lastRepair",
        "date": "2022-01-02"
      }
    ]
  }

```

User can search and group cars data by several fields (make, country, state, etc)  
As a result of search and grouping user can get some table like below:

 ![Screen Shot 2022-02-10 at 16.52.59](https://us1.discourse-cdn.com/elastic/original/3X/9/6/96fc13ee65467b3331e455cc8d5b813968b3c2b8.png)  
Basically, the rule for table cell is when the values count is less than 5 - all the values are shown, otherwise Many(count) is displayed.

**The issue:**  
For each column I have two aggregations: cardinality and terms. The rough query I use is below:

```auto
GET cars.read/_search
{
  "size": 0,
  "query": {
    ...
  },
  "aggregations": {
    "root": {
      "composite": {
        "size": 50,
        "sources": [{
          "countries": {
            "terms": {
              "field": "country"
            }
          }
        }, {
          "states": {
            "terms": {
              "field": "state"
            }
          }
        }]
      },
      "aggregations": {
        "cities": {
          "terms": {
            "field": "city",
            "size": 5
          }
        },
        "citiesCount": {
          "cardinality": {
            "field": "city"
          }
        },
        "makes": {
          "terms": {
            "field": "make",
            "size": 5
          }
        },
        "makesCount": {
          "cardinality": {
            "field": "make"
          }
        },
        "models": {
          "terms": {
            "field": "model",
            "size": 5
          }
        },
        "modelsCount": {
          "cardinality": {
            "field": "model"
          }
        },
        "colors": {
          "terms": {
            "field": "color",
            "size": 5
          }
        },
        "colorsCount": {
          "cardinality": {
            "field": "color"
          }
        },
       // other fields aggregations
        "datesNested": {
          "nested": {
            "path": "dates"
          },
          "aggregations": {
            "dates": {
              "terms": {
                "field": "dates.type",
                "size": 10
              },
              "aggregations": {
                "value": {
                  "terms": {
                    "field": "dates.date",
                    "missing": 0,
                    "size": 5
                  }
                },
                "valueCount": {
                  "cardinality": {
                    "field": "dates.date",
                    "missing": 0
                  }
                }
              }
            }
          }
        }
      }
    }
  }
}

```

Under some conditions I'm getting `too many buckets` exception. As far as I understand even though I add `"size" : 5` line to term aggregations, Elastic will create buckets for all the unique values within a shard, so hitting the 10k bucket limit is no surprise.  
In fact, I don't really need terms aggregation when field cardinality is greater than 5, so in most cases I'm just wasting Elasticsearch resources.  
So I'm wondering:

1. Is there a way to have some kind of conditional terms aggregation, e.g. do the terms aggregation when cardinality is low?
2. I tried sampler aggregation, but it didn't seem to help. Sampler aggregation was added as a sub-aggregation to root, not sure if it's a good place to use sampler.
3. Does it make sense to play around with term's `shard_size` field? I don't need 5 top scored values, any 5 values are fine.
4. Will `significant_terms` (or any other) aggregation help me here?

Would be grateful for any reply.

Best regards,  
Alex

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [February 11, 2022, 9:14am UTC](https://discuss.elastic.co/t/how-to-avoid-or-ease-terms-aggregation-for-fields-with-high-cardinality/296853/2 "2022-02-11T09:14:30Z")

</div>

> [@asupranovich](#):
>
> I don't really need terms aggregation when field cardinality is greater than 5,

Sounds like the [scripted metric aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-metrics-scripted-metric-aggregation.html) would be the way to implement this special logic

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [February 11, 2022, 10:59am UTC](https://discuss.elastic.co/t/how-to-avoid-or-ease-terms-aggregation-for-fields-with-high-cardinality/296853/3 "2022-02-11T10:59:00Z")

</div>

> [@asupranovich](#):
>
> As far as I understand even though I add `"size" : 5` line to term aggregations, Elastic will create buckets for all the unique values within a shard,

I suppose not. I changed `search.max_buckets` and `size` for terms aggregation, terms aggregation with `size` the same as or less than `search.max_buckets` worked well.

The number of combinations would cause the extreme number of buckets. There are 50 composite buckets and for example 10\*5 on datesNested aggregation. They come up to 50 \* 10 \* 5 = 2500 buckets. If there are some other fields, it could be above the limit.

As [pagination](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html#_pagination) is possible with composite aggregation, it could be a simple solution in this case.

---

<div class="post-metadata">

**Author:** ![asupranovich](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/asupranovich/32/101652_2.png) [@asupranovich](https://discuss.elastic.co/u/asupranovich)\
**Post date:** [February 12, 2022, 11:34am UTC](https://discuss.elastic.co/t/how-to-avoid-or-ease-terms-aggregation-for-fields-with-high-cardinality/296853/4 "2022-02-12T11:34:22Z")

</div>

@Mark_Harwood thanks for reply, did some tests - looks promising!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 12, 2022, 11:35am UTC](https://discuss.elastic.co/t/how-to-avoid-or-ease-terms-aggregation-for-fields-with-high-cardinality/296853/5 "2022-03-12T11:35:22Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
