# Bucket sort aggregation with actual documents

**URL:** <https://discuss.elastic.co/t/bucket-sort-aggregation-with-actual-documents/249751>\
**Category:** Elasticsearch\
**Created:** [September 24, 2020, 7:18am UTC](https://discuss.elastic.co/t/bucket-sort-aggregation-with-actual-documents/249751 "2020-09-24T07:18:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![het](https://avatars.discourse-cdn.com/v4/letter/h/c67d28/32.png) [@het](https://discuss.elastic.co/u/het)\
**Post date:** [September 24, 2020, 7:18am UTC](https://discuss.elastic.co/t/bucket-sort-aggregation-with-actual-documents/249751/1 "2020-09-24T07:18:29Z")

</div>

Hi team,

I wanted to get aggegatated data with pagination. So I have tried to use Bucket-Sort. I have referred this URL.

> **[Bucket Sort Aggregation | Elasticsearch Reference \[7.7\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.7/search-aggregations-pipeline-bucket-sort-aggregation.html)**

Now, I am able to get aggegatated data, but the problem is, I am able to get the only the key,aggegation and document count field in the Search response.

My query :

```
{
    "from": 0,
    "size": 0,
    "query": {
        "bool": {
            "must": [
                {
                    "query_string": {
                        "query": "name:hello*^10000.0",
                    }
                }
            ],
            "filter": [
                {
                    "terms": {
                        "locale": [
                            "XYZ",
                            "ABC"
                        ],
                        "boost": 1.0
                    }
                }
            ],
            "adjust_pure_negative": true,
            "boost": 1.0
        }
    },
    "aggregations": {
        "groupbyid": {
            "terms": {
                "field": "groupbyid.raw",
                "size": 10000
            },
            "aggs": {
                "test_bucket_sort": {
                    "bucket_sort": {
                        "size": 2,
                        "from": 3
                    }
                }
            }
        }
    }
}

```

The response which I get:

```
{
    "took": 16,
    "timed_out": false,
    "_shards": {
        "total": 4,
        "successful": 4,
        "skipped": 0,
        "failed": 0
    },
    "hits": {
        "total": {
            "value": 1311,
            "relation": "eq"
        },
        "max_score": null,
        "hits": []
    },
    "aggregations": {
        "groupbyid": {
            "doc_count_error_upper_bound": 0,
            "sum_other_doc_count": 0,
            "buckets": [
                {
                    "key": "id1",
                    "doc_count": 3
                },
                {
                    "key": "id3",
                    "doc_count": 3
                }
            ]
        }
    }
}

```

I wanted in each bucket all the 3 document source also. I have tried couple of things but seems it did not work.

I got this error when I have tried to use `top_hits` with include parameter

`[bucket_sort] unknown field [_source]`

I am on elasticsearch 7.7 version.

Please help me out of this.

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [September 24, 2020, 2:46pm UTC](https://discuss.elastic.co/t/bucket-sort-aggregation-with-actual-documents/249751/2 "2020-09-24T14:46:25Z")

</div>

I suggest to look into [composite aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html), you can implement `groupbyid` as values source. I am not sure what you need bucket sort for, if you need sorting, you can sort in the composite aggregation. For retrieving the source of all documents in the bucket you can use `scripted_metric`, e.g.

```auto
"all_docs": {
  "scripted_metric": {
    "init_script": "state.docs = []",
    "map_script": "state.docs.add(new HashMap(params['_source']))",
    "combine_script": "return state.docs",
    "reduce_script": "def docs = []; for (s in states) {for (d in s) { docs.add(d);}}return docs"
  }
}

```

But be careful, if you have a lot of docs in a bucket, this can cause memory explosion.

---

<div class="post-metadata">

**Author:** ![het](https://avatars.discourse-cdn.com/v4/letter/h/c67d28/32.png) [@het](https://discuss.elastic.co/u/het)\
**Post date:** [September 29, 2020, 11:27am UTC](https://discuss.elastic.co/t/bucket-sort-aggregation-with-actual-documents/249751/3 "2020-09-29T11:27:54Z")

</div>

Hi @Hendrik_Muhs,

Thank you for your reply. I wanted to use pagination + aggregation. So have used bucket sort.

Right now, was thinking to use composite aggregation as this using script, could be an expensive operation as you said.

Thanks

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 27, 2020, 11:27am UTC](https://discuss.elastic.co/t/bucket-sort-aggregation-with-actual-documents/249751/4 "2020-10-27T11:27:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
