# \_score higher than suspected

**URL:** <https://discuss.elastic.co/t/-score-higher-than-suspected/27820>\
**Category:** Elasticsearch\
**Created:** [August 21, 2015, 7:25am UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820 "2015-08-21T07:25:15Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Wojciech\_Rola](https://avatars.discourse-cdn.com/v4/letter/w/74df32/32.png) [@Wojciech\_Rola](https://discuss.elastic.co/u/Wojciech_Rola)\
**Post date:** [August 21, 2015, 7:25am UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/1 "2015-08-21T07:25:16Z")

</div>

Hi @kimchy😄

I was waiting for the moment I will need your help and it's today 😃 .  
I have tags, with this mapping:

```
        "tag": {
        "_all": {
            "index_analyzer": "nGram_analyzer",
            "search_analyzer": "whitespace_analyzer"
        },
        "properties": {
            "name": {
                "type": "string",
                "index_analyzer": "nGram_analyzer",
                "search_analyzer": "whitespace_analyzer"
            },
        }
    }

```

and now. I'm searching in this way:

```
{
   "sort":[
      "_score"
   ],
   "query":{
      "bool":{
         "must":[
            {
               "match":{
                  "name":{
                     "operator":"and",
                     "query":"blue"
                  }
               }
            }
         ]
      }
   },
   "size":20
}

```

What I have in result? When searching 20 items I have 20 other words containing blue, for example bluesea, bluesky, but not blue. What could it be that blue hasn't higher score than words containing blue?

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [August 21, 2015, 11:01am UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/2 "2015-08-21T11:01:13Z")

</div>

I'm not Shay, but I might be able to help =)

Hard to say without seeing the documents, but there is more to scoring than just token matching. For example, the length of the field is taken into account, as well as the individual term and doc frequency. It's likely that some of those ngram fragments are matching other parts of the document and contributing to a higher score.

If you add `explain: true` to your query, you'll get a dump of how Lucene calculated the score. It is pretty verbose, but not too terrible to read. If you gist it up, I can take a look.

---

<div class="post-metadata">

**Author:** ![Wojciech\_Rola](https://avatars.discourse-cdn.com/v4/letter/w/74df32/32.png) [@Wojciech\_Rola](https://discuss.elastic.co/u/Wojciech_Rola)\
**Post date:** [August 21, 2015, 12:03pm UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/3 "2015-08-21T12:03:52Z")

</div>

@polyfractal Yes, you're right, now I see. Wondering how can I determine based on what \_score should be calculated?

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [August 21, 2015, 12:14pm UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/4 "2015-08-21T12:14:43Z")

</div>

Not sure I understand your question?

In general, the actual score generated by Lucene is relatively meaningless. One query may return results from 0-1, another may return results from 0-0.03, and another 0-100. You can't really compare scores. Just think of them as the relative ranking for documents returned by the search.

---

<div class="post-metadata">

**Author:** ![Wojciech\_Rola](https://avatars.discourse-cdn.com/v4/letter/w/74df32/32.png) [@Wojciech\_Rola](https://discuss.elastic.co/u/Wojciech_Rola)\
**Post date:** [August 21, 2015, 12:48pm UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/5 "2015-08-21T12:48:28Z")

</div>

I was thinking, is it possible to define how \_score should be calculated.

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [August 21, 2015, 1:08pm UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/6 "2015-08-21T13:08:54Z")

</div>

Ah, I see. Yes, you can modulate the score in a number of ways. I'd recommend reading through the [Controlling Relevancy](https://www.elastic.co/guide/en/elasticsearch/guide/current/controlling-relevance.html) portion of the Definitive Guide, which outlines a bunch of ways to modulate or override the score.

---

<div class="post-metadata">

**Author:** ![Wojciech\_Rola](https://avatars.discourse-cdn.com/v4/letter/w/74df32/32.png) [@Wojciech\_Rola](https://discuss.elastic.co/u/Wojciech_Rola)\
**Post date:** [September 1, 2015, 9:06am UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/7 "2015-09-01T09:06:35Z")

</div>

Thanks.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:52pm UTC](https://discuss.elastic.co/t/-score-higher-than-suspected/27820/8 "2017-07-05T23:52:57Z")

</div>


