# Per bucket scoring in aggregations

**URL:** <https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488>\
**Category:** Elasticsearch\
**Created:** [July 29, 2015, 1:43pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488 "2015-07-29T13:43:56Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![sHenzen](https://avatars.discourse-cdn.com/v4/letter/s/df705f/32.png) [@sHenzen](https://discuss.elastic.co/u/sHenzen)\
**Post date:** [July 29, 2015, 1:43pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488/1 "2015-07-29T13:43:56Z")

</div>

Hi all,

I have documents that are, very simplified, like this:

```
{ 
  id: 20
  base_quality: 10
  application: [
    { id: 2, quality: 10 }
    { id: 3, quality: 20 }
  ]
}

```

What I want to do is:

- Bucket them by `application.id`
- Calculate a score for every document in every bucket based on `base_quality + application.quality` (`application.quality` where `application.id` = `id` of bucket)
- Get the best scoring document for every bucket

It's easy to bucket documents by `application.id` and to get the best quality for every bucket:

```
  {
    query: { match_all: {} }
    aggs: {
      nested1: { 
        nested: { path: 'applications' },
        aggs: {
          terms1: {
            terms: { field: 'applications.id'},
            aggs: { min_price: { 
                min: { script: "doc['quality'].value + _source.base_quality" }
              }
            }
          }
        }
      }
    }
  }

```

But I want is _the document that creates this quality_. Is that possible? What I need is something like top hits aggregation, but then with custom scoring. Maybe with a scripted metric aggregation?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [July 29, 2015, 2:10pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488/2 "2015-07-29T14:10:41Z")

</div>

Why not use the [`function_score` query](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-function-score-query.html) in the query section to score the document based on your criteria and then use the `top_hits` aggregation to get the top doc for each bucket (the top doc will have a score based on your `function_score` query)?

---

<div class="post-metadata">

**Author:** ![sHenzen](https://avatars.discourse-cdn.com/v4/letter/s/df705f/32.png) [@sHenzen](https://discuss.elastic.co/u/sHenzen)\
**Post date:** [July 29, 2015, 2:23pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488/3 "2015-07-29T14:23:27Z")

</div>

@colings86 the problem is that the score of a document can be different in every bucket where the document appears (based on the nested document that caused it to be in that bucket).

You can get a score per nested document in the query, but the combined score for the top-level document is used by the `top_hits` aggregation.

---

<div class="post-metadata">

**Author:** ![sHenzen](https://avatars.discourse-cdn.com/v4/letter/s/df705f/32.png) [@sHenzen](https://discuss.elastic.co/u/sHenzen)\
**Post date:** [July 29, 2015, 2:27pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488/4 "2015-07-29T14:27:15Z")

</div>

Ow and thanks for the response!

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [July 29, 2015, 2:39pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488/5 "2015-07-29T14:39:06Z")

</div>

Ok, I had missed the nested agg in there. However, you should be able to just use the [`sort`](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-metrics-top-hits-aggregation.html#search-aggregations-metrics-top-hits-aggregation) in the `top_hits` agg to order the documents by ascending `quality` field since the `base_quality` will be the same for all the documents in the same bucket?

---

<div class="post-metadata">

**Author:** ![sHenzen](https://avatars.discourse-cdn.com/v4/letter/s/df705f/32.png) [@sHenzen](https://discuss.elastic.co/u/sHenzen)\
**Post date:** [July 29, 2015, 2:54pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488/6 "2015-07-29T14:54:20Z")

</div>

I've looked into that. The `base_quality` can be different for every document, and the `application.quality` can be different for every nested document. They are bucketed purely on `application.id`. If I were able to sort them by descending `quality + base_quality` and then just get the first one that would be great, but I don't know how.

Also, I asumed `sort` was only performed on the results actually returned by `top_hits`, and that those were always determined by `_score`. If it's not, it really is almost exactly what I need, but not quite 😁 .

What I think I need is something like:

```
query: { match_all: {} },
aggs: {
  nested1: { 
    nested: { path: 'applications' },
    aggs: {
      terms1: {
        terms: { field: 'applications.id'},
        aggs: { best_quality: { 
          scripted_metric: {
             init_script: "_agg['results'] = []",
             map_script: "_agg.results.add([source: _source, score: doc['quality'].value + _source.base_quality])",
             # This is psuedo code, don't know if it can actually be done
             reduce_script: "result = []; for (a in _aggs) { result.add(a.results.sort().first()) }; return result.sort().first()"
          }
        }}
      }
    }
  }
}

```

But I can't figure out how to do the sorting in the reduce script.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:58pm UTC](https://discuss.elastic.co/t/per-bucket-scoring-in-aggregations/26488/7 "2017-07-05T23:58:19Z")

</div>


