# Count emails that have more than 10 documents (buckets count)

**URL:** <https://discuss.elastic.co/t/count-emails-that-have-more-than-10-documents-buckets-count/307477>\
**Category:** Elasticsearch\
**Created:** [June 17, 2022, 6:42am UTC](https://discuss.elastic.co/t/count-emails-that-have-more-than-10-documents-buckets-count/307477 "2022-06-17T06:42:03Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jack\_daniels00](https://avatars.discourse-cdn.com/v4/letter/j/3bc359/32.png) [@jack\_daniels00](https://discuss.elastic.co/u/jack_daniels00)\
**Post date:** [June 17, 2022, 6:42am UTC](https://discuss.elastic.co/t/count-emails-that-have-more-than-10-documents-buckets-count/307477/1 "2022-06-17T06:42:03Z")

</div>

Hi!  
I have documents in elastic like:

```auto
{
    "email": "user@example.com",
    "subject": "Email subject",
    "body": "Email body"
}

```

And I'm trying to count all unique emails that have more than 10 documents.  
I can retrieve all such emails by the aggregation query:

```auto
GET emails/_search
{
  "size": 0,
  "aggs": {
    "email_with_more_than_10_docs": {
      "terms": {
        "field": "email",
        "min_doc_count": 10
      }
    }
  }
}

```

And I get the output:

```auto
{
  "took" : 3,
  "timed_out" : false,
  "_shards" : {
    "total" : 3,
    "successful" : 3,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 74,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : []
  },
  "aggregations" : {
    "email_with_more_than_10_docs" : {
      "doc_count_error_upper_bound" : 0,
      "sum_other_doc_count" : 0,
      "buckets" : [
        {
          "key" : "user1@example.com",
          "doc_count" : 40
        },
        {
          "key" : "user2@example.com",
          "doc_count" : 14
        },
        {
          "key" : "user3@example.com",
          "doc_count" : 13
        }
      ]
    }
  }
}

```

But how can I get the number of the unique emails (number of buckets)? In the example of output above that number is 3.

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [June 20, 2022, 6:50am UTC](https://discuss.elastic.co/t/count-emails-that-have-more-than-10-documents-buckets-count/307477/2 "2022-06-20T06:50:08Z")

</div>

As long as the `size` parameter of aggregation is larger than the number of emails which meet the criteria, count the length of the `buckets` array on the client side is a viable option to implement.

If you really need the count in the response, you may use stats\_bucket pipeline aggregation. Keep in mind set "size" larger than the number of terms. The "`count`" value is what you want.

```auto
GET kibana_sample_data_flights/_search?filter_path=aggregations.stats
{
  "size":0,
  "aggs": {
    "dest": {
      "terms": {
        "field": "DestAirportID",
        "size": 10
      }
    },
    "stats":{
      "stats_bucket": {
        "buckets_path": "dest>_count"
      }
    }
  }
}

```

```auto
{
  "aggregations" : {
    "stats" : {
      "count" : 10,
      "min" : 305.0,
      "max" : 691.0,
      "avg" : 416.1,
      "sum" : 4161.0
    }
  }
}

```

```auto
GET kibana_sample_data_flights/_search?filter_path=aggregations.stats
{
  "size":0,
  "aggs": {
    "dest": {
      "terms": {
        "field": "DestAirportID",
        "size": 1000
      }
    },
    "stats":{
      "stats_bucket": {
        "buckets_path": "dest>_count"
      }
    }
  }
}

```

```auto
{
  "aggregations" : {
    "stats" : {
      "count" : 156,
      "min" : 1.0,
      "max" : 691.0,
      "avg" : 83.71153846153847,
      "sum" : 13059.0
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![jack\_daniels00](https://avatars.discourse-cdn.com/v4/letter/j/3bc359/32.png) [@jack\_daniels00](https://discuss.elastic.co/u/jack_daniels00)\
**Post date:** [July 13, 2022, 6:42am UTC](https://discuss.elastic.co/t/count-emails-that-have-more-than-10-documents-buckets-count/307477/3 "2022-07-13T06:42:09Z")

</div>

But what if there are a lot of documents in elstic? Is there any other solution that takes it into account?

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [July 13, 2022, 8:58am UTC](https://discuss.elastic.co/t/count-emails-that-have-more-than-10-documents-buckets-count/307477/4 "2022-07-13T08:58:07Z")

</div>

For pagination of aggregation results, [Composite aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html#_pagination) may help you.

Using [transform](https://www.elastic.co/guide/en/elasticsearch/reference/current/transform-overview.html) to perform aggregation beforehand might be another option.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 10, 2022, 8:58am UTC](https://discuss.elastic.co/t/count-emails-that-have-more-than-10-documents-buckets-count/307477/5 "2022-08-10T08:58:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
