# The fuzzier matching, the higher score?

**URL:** <https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009>\
**Category:** Elasticsearch\
**Created:** [December 24, 2018, 3:35pm UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009 "2018-12-24T15:35:23Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jevgenij](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jevgenij/32/33543_2.png) [@Jevgenij](https://discuss.elastic.co/u/Jevgenij)\
**Post date:** [December 24, 2018, 3:35pm UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009/1 "2018-12-24T15:35:23Z")

</div>

Hello!

My template is:

```auto
        ...,
        "code": {
          "type": "keyword",
          "copy_to": "full_text"
        },
        ...,
        "full_text": {
          "type": "text"
        }

```

My query is:

```auto
{
    "bool": {
        "must": {
            "match": {
                "full_text": {
                    "query": "AD2480ME",
                    "operator": "and",
                    "fuzziness": "AUTO"
                }
            }
        }
    }
}

```

And response is:

```auto
        {
            ...,
            "_score": 2.7549238,
            "_source": {
                "code": "AD3440ME",
                ....
            }
        },
        {
            ...,
            "_score": 2.7438653,
            "_source": {
                "code": "AD2480ME",
                ...
            }
        }

```

So the question is why the exact matching record has lower score than fuzzy one?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 24, 2018, 4:02pm UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009/2 "2018-12-24T16:02:57Z")

</div>

The [explain API](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-explain.html) can help diagnose what's going on. I expect it may be to do with IDF (rarity) of the terms.  
What version of elasticsearch are you running and how many shards/docs per shard do you have?

---

<div class="post-metadata">

**Author:** ![Jevgenij](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jevgenij/32/33543_2.png) [@Jevgenij](https://discuss.elastic.co/u/Jevgenij)\
**Post date:** [December 25, 2018, 8:22pm UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009/3 "2018-12-25T20:22:16Z")

</div>

Elasticsearch version `6.5.2`

Actually I've met the issue on my development dataset:

```auto
$ curl -XGET 'http://localhost:9200/my_index/_count'

```

```auto
{
    "count": 18,
    "_shards": {
        "total": 5,
        "successful": 5,
        "skipped": 0,
        "failed": 0
    }
}

```

Mentioned `code` property is unique through index docs.

Not sure what do you mean by

> [@Mark\_Harwood](#):
>
> IDF (rarity) of the terms

But Explain API has brought unexpected results.

First of all I did:

```auto
GET /my_index/_search

```

```auto
{
    "query": {
        "bool": {
            "must": {
                "match": {
                    "full_text": {
                        "query": "AD2480ME",
                        "operator": "and",
                        "fuzziness": "AUTO"
                    }
                }
            }
        }
    }
}

```

And got:

```auto
{
    "took": 22,
    "timed_out": false,
    "_shards": {
        "total": 5,
        "successful": 5,
        "skipped": 0,
        "failed": 0
    },
    "hits": {
        "total": 2,
        "max_score": 0.7549239,
        "hits": [
            {
                "_index": "my_index",
                "_type": "doc",
                "_id": "31",
                "_score": 0.7549239,
                "_source": {
                    "code": "AD3440ME"
                }
            },
            {
                "_index": "my_index",
                "_type": "doc",
                "_id": "22",
                "_score": 0.7438652,
                "_source": {
                    "code": "AD2480ME"
                }
            }
        ]
    }
}

```

So the next thing I did was:

```auto
GET /my_index/_default_/31/_explain

```

and

```auto
GET /my_index/_default_/22/_explain

```

both with

```auto
{
    "query": {
        "bool": {
            "must": {
                "match": {
                    "full_text": {
                        "query": "AD2480ME",
                        "operator": "and",
                        "fuzziness": "AUTO"
                    }
                }
            }
        }
    }
}

```

And both returned me the same

```auto
{
    "_index": "my_index",
    "_type": "_default_",
    "_id": "22",
    "matched": false,
    "explanation": {
        "value": 0,
        "description": "Failure to meet condition(s) of required/prohibited clause(s)",
        "details": []
    }
}

```

with a lot of different details about score counting, but with the same `"matched": false`.

So, if I understand correctly, the exactly matching doc is not being seen as matching one.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 25, 2018, 11:51pm UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009/4 "2018-12-25T23:51:08Z")

</div>

> [@Jevgenij](#):
>
> GET /my\_index/_default_/31/\_explain

Use “doc” not “\_default” here

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 26, 2018, 10:10am UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009/5 "2018-12-26T10:10:17Z")

</div>

> [@Jevgenij](#):
>
> Mentioned `code` property is unique through index docs.

18 docs in 5 shards will have some funky scoring. The number of docs in each will be very uneven (maybe varying by 25%) so the IDF score of a unique term will be affected.

Choices are:

- Add more docs
- Use one shard or
- Use “[DFS](https://www.elastic.co/blog/understanding-query-then-fetch-vs-dfs-query-then-fetch)” style search

---

<div class="post-metadata">

**Author:** ![Jevgenij](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jevgenij/32/33543_2.png) [@Jevgenij](https://discuss.elastic.co/u/Jevgenij)\
**Post date:** [December 26, 2018, 11:39am UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009/6 "2018-12-26T11:39:45Z")

</div>

Thank you for your time, attention to detail, corrections and right URLs to read!

As soon as I switched the local development elasticsearch instance to one shard, the funky scoring disappeared and everything fell into place!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 23, 2019, 11:39am UTC](https://discuss.elastic.co/t/the-fuzzier-matching-the-higher-score/162009/7 "2019-01-23T11:39:46Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
