# What's the best approach to balance the text similarity with different fields weights?

**URL:** <https://discuss.elastic.co/t/whats-the-best-approach-to-balance-the-text-similarity-with-different-fields-weights/171898>\
**Category:** Elasticsearch\
**Created:** [March 12, 2019, 9:19am UTC](https://discuss.elastic.co/t/whats-the-best-approach-to-balance-the-text-similarity-with-different-fields-weights/171898 "2019-03-12T09:19:15Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [March 12, 2019, 9:19am UTC](https://discuss.elastic.co/t/whats-the-best-approach-to-balance-the-text-similarity-with-different-fields-weights/171898/1 "2019-03-12T09:19:15Z")

</div>

Say we have two docs

```json
{"_id": 1, "title": "James Harden wins the MVP", "content": "xxxxxxxxxxxxxxxxxxxxxx"}

{"_id": 2, "title": "The new 007 movie comes!", "content": "xxxx James Bond xxxxxxxxxxxxx"}

```

And when users searched query `James Bond`, we may construct es query like this

```auto
GET docs/_search
{
  "query": {
    "multi_match": {
      "query": "james bond",
      "fields": ["title^3", "content"]
    }
  }
}

```

For the overweight of title field, `doc 1` may score better than `doc 2`.

So my question is what the best approach to make sure `doc 2` scores better `doc 1`.

Thanks for help!

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [March 12, 2019, 10:09am UTC](https://discuss.elastic.co/t/whats-the-best-approach-to-balance-the-text-similarity-with-different-fields-weights/171898/2 "2019-03-12T10:09:36Z")

</div>

Generally, if you blend strict and sloppier interpretations of a user query the docs that match best (strict AND sloppy) will rank higher.

In declining order of strictness:

1. Phrase query (all terms must match and be next to each other in the text)
2. AND query (all terms must appear somewhere in the text)
3. OR query (at least one term must match)
4. fuzzy query (at least one vaguely reminiscent term must match).

These can all be assembled into a single `bool` query in the `should` property.  
The more clauses a document matches, the higher the score - the downside is it will be more costly to run.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 12, 2019, 10:19am UTC](https://discuss.elastic.co/t/whats-the-best-approach-to-balance-the-text-similarity-with-different-fields-weights/171898/3 "2019-03-12T10:19:07Z")

</div>

I wrote an example of this in the following gist:

> <https://gist.github.com/dadoonet/5179ee72ecbf08f12f53d4bda1b76bab#file-search_kibana_console-txt-L362-L457>

---

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [March 22, 2019, 9:18am UTC](https://discuss.elastic.co/t/whats-the-best-approach-to-balance-the-text-similarity-with-different-fields-weights/171898/4 "2019-03-22T09:18:36Z")

</div>

Thanks for help!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 19, 2019, 9:18am UTC](https://discuss.elastic.co/t/whats-the-best-approach-to-balance-the-text-similarity-with-different-fields-weights/171898/5 "2019-04-19T09:18:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
