# Return only highest scoring document from a family of documents

**URL:** <https://discuss.elastic.co/t/return-only-highest-scoring-document-from-a-family-of-documents/151541>\
**Category:** Elasticsearch\
**Created:** [October 9, 2018, 2:10am UTC](https://discuss.elastic.co/t/return-only-highest-scoring-document-from-a-family-of-documents/151541 "2018-10-09T02:10:03Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jon\_Hourany](https://avatars.discourse-cdn.com/v4/letter/j/8baadc/32.png) [@Jon\_Hourany](https://discuss.elastic.co/u/Jon_Hourany)\
**Post date:** [October 9, 2018, 2:10am UTC](https://discuss.elastic.co/t/return-only-highest-scoring-document-from-a-family-of-documents/151541/1 "2018-10-09T02:10:03Z")

</div>

Let's say I have documents mapped such that:

```
PUT test_documents
{
  "mappings": {
    "doc": {
      "properties": {
        "parent_id": { "type": "keyword" },
        "body": { "type": "text" }
      }
    }
  }
}

```

Where `body` is some body of text and `parent_id` is the id of the parent document where that body of text came from

```
PUT test_documents/doc/1
{
  "parent_id": "ZOO BOOK",
  "body": "Zoo's are places where you can see animals"
}

PUT test_documents/doc/2
{
  "parent_id": "ZOO BOOK",
  "body": "Zoo's have lots of animals"
}

PUT test_documents/doc/3
{
  "parent_id": "VET BOOK",
  "body": "Vet's are doctors for animals"
}

```

When I do a search on this text for both "zoo's" and "animals" I'll get all three documents back as expected

```
GET test_documents/_search
{
  "query": {
    "bool": {
      "should": [
        {
          "match": {
            "body": "zoo's animals"
          }
        }
      ]
    }
  }
}

```

but what I'd like is for the return to only have the highest scoring member from each document that shares a `parent_id` so that in this case, the return would only have 2 documents: the highest scoring member from "ZOO BOOK" and the highest scoring from "VET BOOK" **in order of relevance** so that if the order of relevance was "ZOO BOOK", "VET BOOK", "ZOO BOOK" this distinct list would just be "ZOO BOOK", "VET BOOK".

I tried doing aggregation on the `parent_id` field but that didn't really do what I wanted.

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [October 10, 2018, 2:38pm UTC](https://discuss.elastic.co/t/return-only-highest-scoring-document-from-a-family-of-documents/151541/2 "2018-10-10T14:38:41Z")

</div>

Take a look at the [field collapsing](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-collapse.html) feature. It allows you to return the highest scoring document for unique values of a specific field.

To get to what you want to do, your request would look something like this:

```auto
GET test_documents/_search
{
  "query": {
    "bool": {
      "should": [
        {
          "match": {
            "body": "zoo's animals"
          }
        }
      ]
    }
  },
  "collapse": {
    "field": "parent_id"
  }
}

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 7, 2018, 2:47pm UTC](https://discuss.elastic.co/t/return-only-highest-scoring-document-from-a-family-of-documents/151541/3 "2018-11-07T14:47:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
