# Top N documents with multiple buckets

**URL:** <https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354>\
**Category:** Elasticsearch\
**Created:** [August 31, 2015, 1:20pm UTC](https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354 "2015-08-31T13:20:09Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pieter\_Agenbag](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pieter_agenbag/32/4562_2.png) [@Pieter\_Agenbag](https://discuss.elastic.co/u/Pieter_Agenbag)\
**Post date:** [August 31, 2015, 1:20pm UTC](https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354/1 "2015-08-31T13:20:09Z")

</div>

Hi.

I'm trying to get the top N documents for a specific aggregation , but with multiple term buckets - or any other way to get the results.

Basically , I have a documents with from.ip,from.hostname , to.ip,to.hostname, from.bytes,to.bytes, total.bytes (amongst other) fields.

I would like to just get the top 10 documents in terms of sum total.bytes.

In sql terms...

> select from.ip,from.hostname,to.ip,to.hostname,sum(total.bytes) from myTable group by from.ip,from.hostname,to.ip,to.hostname

I have the following query request, but it returns the top 10 for Each bucket , which obviously results in waaay too many results.  
P.s. I know my query is wrong in that it sorts on \_count and not sum(total.bytes) ...

```
{
    "query": {
        "match_all": {}
    },
    "size": 10,
    "aggs": {
        "group": {
            "terms": {
                "field": "from.ip.raw",
                "size": 10,
                "order": {
                    "_count": "desc"
                }
            },
            "aggs": {
                "group": {
                    "terms": {
                        "field": "from.hostname.raw",
                        "size": 10,
                        "order": {
                            "_count": "desc"
                        }
                    },
                    "aggs": {
                        "group": {
                            "terms": {
                                "field": "to.ip.raw",
                                "size": 10,
                                "order": {
                                    "_count": "desc"
                                }
                            },
                            "aggs": {
                                "group": {
                                    "terms": {
                                        "field": "to.hostname.raw",
                                        "size": 10,
                                        "order": {
                                            "aggval": "desc"
                                        }
                                    },
                                    "aggs": {
                                        "aggval": {
                                            "sum": {
                                                "field": "total.bytes"
                                            }
                                        }
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}

```

---

<div class="post-metadata">

**Author:** ![Pieter\_Agenbag](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pieter_agenbag/32/4562_2.png) [@Pieter\_Agenbag](https://discuss.elastic.co/u/Pieter_Agenbag)\
**Post date:** [September 1, 2015, 6:45am UTC](https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354/2 "2015-09-01T06:45:01Z")

</div>

Did some more searching and seems a lot of people (probably people coming from relational databases) , are asking this or similar questions.

I seem to have solved my problem using a single (script) terms bucket , combining the four "group by" fields into one... I'll then just have to split the field values out again when processing the result.

I hope this is the right way to do it...or at least an acceptable way 😄

```
{
    "query": {
        "match_all": {}
    },
    "size": 10,
    "aggs": {
        "flow": {
            "terms": {
                "script": "doc['from.ip'].value + ',' + doc['from.hostname'].value + ',' + doc['to.ip'].value + ',' + doc['to.hostname'].value",
                "size": 10,
                "order": {
                    "total": "desc"
                }
            },
			"aggs": {
				"total": {
					"sum": {
						"field": "total.bytes"
					}
				}
				,"from": {
					"sum": {
						"field": "from.bytes"
					}
				}
				,"to": {
					"sum": {
						"field": "to.bytes"
					}
				}
			}
		}
	}
}
```

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [September 1, 2015, 11:47am UTC](https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354/3 "2015-09-01T11:47:58Z")

</div>

I think in this case the `top_hits` aggregation is something you would use in this case. It allows you to return the top matching documents per bucket (depending on the level you define it).

[https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-metrics-top-hits-aggregation.html?q=top\_hits#search-aggregations-metrics-top-hits-aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-metrics-top-hits-aggregation.html?q=top_hits#search-aggregations-metrics-top-hits-aggregation)

---

<div class="post-metadata">

**Author:** ![samthurston](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/samthurston/32/4812_2.png) [@samthurston](https://discuss.elastic.co/u/samthurston)\
**Post date:** [September 16, 2015, 3:35pm UTC](https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354/4 "2015-09-16T15:35:56Z")

</div>

`top_hits` doesn't allow for sub-aggregation so getting stats on the results seems problematic

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [September 17, 2015, 1:14pm UTC](https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354/6 "2015-09-17T13:14:18Z")

</div>

On what kind of property do you like to aggregate under the top\_hits agg?

top\_hits is a metric agg and the idea is that it should be a leaf agg (just  
like other metric aggs), If you like to aggregate deeper then you'll need  
to add an additional bucket agg on the same level that the top\_hits agg is  
and then place another top\_hits agg under that new bucket agg.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:49pm UTC](https://discuss.elastic.co/t/top-n-documents-with-multiple-buckets/28354/7 "2017-07-05T23:49:45Z")

</div>


