# Filter based on the doc\_count with aggregations

**URL:** <https://discuss.elastic.co/t/filter-based-on-the-doc-count-with-aggregations/62677>\
**Category:** Elasticsearch\
**Created:** [October 11, 2016, 2:09am UTC](https://discuss.elastic.co/t/filter-based-on-the-doc-count-with-aggregations/62677 "2016-10-11T02:09:01Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![yeikel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yeikel/32/125434_2.png) [@yeikel](https://discuss.elastic.co/u/yeikel)\
**Post date:** [October 11, 2016, 2:09am UTC](https://discuss.elastic.co/t/filter-based-on-the-doc-count-with-aggregations/62677/1 "2016-10-11T02:09:02Z")

</div>

I have the following index :

```
POST /cars/transactions/_bulk
    { "index": {}}
    { "price" : 10000, "color" : "red", "make" : "honda", "sold" : "2014-10-28" }
    { "index": {}}
    { "price" : 20000, "color" : "red", "make" : "honda", "sold" : "2014-11-05" }
    { "index": {}}
    { "price" : 30000, "color" : "green", "make" : "ford", "sold" : "2014-05-18" }
    { "index": {}}
    { "price" : 15000, "color" : "blue", "make" : "toyota", "sold" : "2014-07-02" }
    { "index": {}}
    { "price" : 12000, "color" : "green", "make" : "toyota", "sold" : "2014-08-19" }
    { "index": {}}
    { "price" : 20000, "color" : "red", "make" : "honda", "sold" : "2014-11-05" }
    { "index": {}}
    { "price" : 80000, "color" : "red", "make" : "bmw", "sold" : "2014-01-01" }
    { "index": {}}
    { "price" : 25000, "color" : "blue", "make" : "ford", "sold" : "2014-02-12" }

```

And I am performing the following search :

```
    GET /cars/transactions/_search
    {
        "size" : 0,
        "aggs" : { 
            "popular_colors" : { 
                "terms" : { 
                  "field" : "color"
                }
            }
        }
    }

```

The response that I receive is the following :

```
    {
      "took": 2,
      "timed_out": false,
      "_shards": {
        "total": 5,
        "successful": 5,
        "failed": 0
      },
      "hits": {
        "total": 8,
        "max_score": 0,
        "hits": []
      },
      "aggregations": {
        "popular_colors": {
          "doc_count_error_upper_bound": 0,
          "sum_other_doc_count": 0,
          "buckets": [
            {
              "key": "red",
              "doc_count": 4
            },
            {
              "key": "blue",
              "doc_count": 2
            },
            {
              "key": "green",
              "doc_count": 2
            }
          ]
        }
      }
    }

```

My question is , **how can I filter by doc\_count?**

For example "return only the documents where **doc\_count is equal to 1**"

I am using Elasticsearch 2.3.5

Thank you

---

<div class="post-metadata">

**Author:** ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)\
**Post date:** [October 11, 2016, 12:43pm UTC](https://discuss.elastic.co/t/filter-based-on-the-doc-count-with-aggregations/62677/2 "2016-10-11T12:43:35Z")

</div>

Hi,

you can use the [bucket\_selector](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-pipeline-bucket-selector-aggregation.html) pipeline aggregation for this kind of filtering. In your case, the following query:

```auto
GET /cars/transactions/_search
{
   "size": 0,
   "aggs": {
      "popular_colors": {
         "terms": {
            "field": "color"
         },
         "aggs": {
            "my_filter": {
               "bucket_selector": {
                  "buckets_path": {
                     "the_doc_count": "_count"
                  },
                  "script": "the_doc_count == 2"
               }
            }
         }
      }
   }
}

```

Should only filter out the buckets with `"doc_count" : 2`. However, be aware that Pipeline aggregations work on the outputs produced from other aggregations, so the overall amount of work that needs to be done to calculate the initial `doc_counts` will be the same. Since the script parts needs to be executed for each input bucket, the opetation might potentially be slow for high cardinality fields (as in thousands of thousands of terms), but it should work well for relatively low cardinality fields (like colors, as in this case).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:13pm UTC](https://discuss.elastic.co/t/filter-based-on-the-doc-count-with-aggregations/62677/3 "2017-07-05T22:13:25Z")

</div>


