# Pipeline aggregation for selecting the documents

**URL:** <https://discuss.elastic.co/t/pipeline-aggregation-for-selecting-the-documents/80450>\
**Category:** Elasticsearch\
**Created:** [March 29, 2017, 9:27am UTC](https://discuss.elastic.co/t/pipeline-aggregation-for-selecting-the-documents/80450 "2017-03-29T09:27:13Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![pgrigorenko\_zt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pgrigorenko_zt/32/13232_2.png) [@pgrigorenko\_zt](https://discuss.elastic.co/u/pgrigorenko_zt)\
**Post date:** [March 29, 2017, 9:27am UTC](https://discuss.elastic.co/t/pipeline-aggregation-for-selecting-the-documents/80450/1 "2017-03-29T09:27:13Z")

</div>

Hello,

I have the documents in the following form:

```
  {
    "_type": "...",
    "_id": "...",
    "_source": {
      "duration": 1000,
      "version": "<some version>",
      "desc": "<some text>",
      "timestamp": "<date>"
    }
  }

```

In a query docs are aggregated by type and desc, filtered by version and in a sub aggregation percentiles are calculated for the 'duration' field. In YAML it looks like this:

```
aggs:
    byType:
        terms: { size: 10000, field: "_type" }
        aggs:
            byDesc:
                terms: { size: 10000, field: "desc" }
                aggs:
                    a:
                        filter : { term: { version: "<version a>" } }
                        aggs:
                            p: { percentiles: { field: "duration", percents: [..., 99] } }
                    b:
                        filter : { term: { version: "<version b>" } }
                        aggs:
                            p: { percentiles: { field: "duration", percents: [..., 99] } }

```

now, I would like to select some documents for both versions which fall into the 99% bucket, in other words, find a doc which has 'duration' value greater or equal to the 99th percentile. Looking at the documentation I failed to find such option. I would really like to avoid doing N+1 requests to fetch the data I need.  
Thanks!

---

<div class="post-metadata">

**Author:** ![pgrigorenko\_zt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pgrigorenko_zt/32/13232_2.png) [@pgrigorenko\_zt](https://discuss.elastic.co/u/pgrigorenko_zt)\
**Post date:** [March 30, 2017, 9:05am UTC](https://discuss.elastic.co/t/pipeline-aggregation-for-selecting-the-documents/80450/2 "2017-03-30T09:05:39Z")

</div>

up, anyone?!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 27, 2017, 9:05am UTC](https://discuss.elastic.co/t/pipeline-aggregation-for-selecting-the-documents/80450/3 "2017-04-27T09:05:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
