# Ngram and score in query

**URL:** <https://discuss.elastic.co/t/ngram-and-score-in-query/93901>\
**Category:** Elasticsearch\
**Created:** [July 20, 2017, 9:50am UTC](https://discuss.elastic.co/t/ngram-and-score-in-query/93901 "2017-07-20T09:50:06Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![weibin.wu](https://avatars.discourse-cdn.com/v4/letter/w/4491bb/32.png) [@weibin.wu](https://discuss.elastic.co/u/weibin.wu)\
**Post date:** [July 20, 2017, 9:50am UTC](https://discuss.elastic.co/t/ngram-and-score-in-query/93901/1 "2017-07-20T09:50:06Z")

</div>

Hi Elasticsearch.

When we do a ngram tokenizer, we will get token with start\_offset.  
{  
"token": "vi",  
"start\_offset": 0,  
"end\_offset": 2,  
"type": "word",  
"position": 0  
},  
{  
"token": "iv",  
"start\_offset": 1,  
"end\_offset": 3,  
"type": "word",  
"position": 1  
}

Can we have like bigger "start\_offset" has lower score than smaller "start\_offset"?

In this case is when I have two document  
{"text" : "vivo"}  
{"text" : "ivov"}  
All I use ngram (min\_gram: 2, max\_gram:2) as tokenizer.  
When I search "iv", can I expect {"text" : "ivov"} has higher score than {"text" : "vivo"} because "start\_offset" is smaller?  
For now I see they have the same score in this case.

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [July 20, 2017, 8:40pm UTC](https://discuss.elastic.co/t/ngram-and-score-in-query/93901/2 "2017-07-20T20:40:21Z")

</div>

You're correct in that ngrams are only scored for how well they match, not the position. I don't think there is a way to weight the score based on their offset.

What's the use-case here? You want to weight matches at the start of the word higher than at the end? You could probably accomplish that manually using span queries but it'd be a huge pain. If you can describe the motivation I might be able to help work out an alternative method 🙂

---

<div class="post-metadata">

**Author:** ![weibin.wu](https://avatars.discourse-cdn.com/v4/letter/w/4491bb/32.png) [@weibin.wu](https://discuss.elastic.co/u/weibin.wu)\
**Post date:** [July 31, 2017, 1:28am UTC](https://discuss.elastic.co/t/ngram-and-score-in-query/93901/3 "2017-07-31T01:28:03Z")

</div>

Thanks Polyfractal,

I have a use case like this.  
Two words: [vivo] [ivid].  
Analyzer: ngram: min\_gram:2, max\_gram:2  
search: "match":{"text": "vi"}  
vivo doc\_id: 1  
ivid doc\_id: 2  
It supposes to give back "vivo" because searching "vi" is more likely as searching "vivo" rather than "ivid".  
However, ngram return the same score, and because of doc\_id, ivid will be the first document return.

Anyway to solve this?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 28, 2017, 1:28am UTC](https://discuss.elastic.co/t/ngram-and-score-in-query/93901/4 "2017-08-28T01:28:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
