# Scoring and boost

**URL:** <https://discuss.elastic.co/t/scoring-and-boost/6094>\
**Category:** Elasticsearch\
**Created:** [December 7, 2011, 11:43pm UTC](https://discuss.elastic.co/t/scoring-and-boost/6094 "2011-12-07T23:43:56Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ciryus](https://avatars.discourse-cdn.com/v4/letter/c/aeb1de/32.png) [@Ciryus](https://discuss.elastic.co/u/Ciryus)\
**Post date:** [December 7, 2011, 11:43pm UTC](https://discuss.elastic.co/t/scoring-and-boost/6094/1 "2011-12-07T23:43:56Z")

</div>

Hi,

I'm new to ElasticSearch, and I have been playing with boosting. Here is my  
test: [https://gist.github.com/1445264](https://gist.github.com/1445264).

I am a bit confused by the results. The boosting effects are those that I  
expect for a single word, but the two words search give scores much closer  
than I would expect. In particular, the phrase search gives the same score  
in all cases, which I don't understand (I expect the same scoring as in  
case 1: keyword, title, body).

Should I largely increase the boosting, or should I do my mapping and/or  
queries differently to achieve this result?

Thanks for any help.

---

<div class="post-metadata">

**Author:** ![alichi](https://avatars.discourse-cdn.com/v4/letter/a/bc79bd/32.png) [@alichi](https://discuss.elastic.co/u/alichi)\
**Post date:** [December 8, 2011, 9:01am UTC](https://discuss.elastic.co/t/scoring-and-boost/6094/2 "2011-12-08T09:01:57Z")

</div>

I believe this is a lucene thing more than anything. Try this URL:

- 

[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)  
\*[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)

This is the core of it:

```
  score(q,d) = coord(q,d)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_coord>

```

· queryNorm(q)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_queryNorm](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_queryNorm)  
· ∑ ( tf(t in d)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_tf](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_tf)  
· idf(t)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_idf](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_idf)  
2 · t.getBoost()[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_termBoost](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_termBoost)  
· norm(t,d)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_norm](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_norm)  
) t in q _Lucene Practical Scoring Function_

where

1. _tf(t in d)_ correlates to the term's _frequency_, defined as the  
number of times term _t_ appears in the currently scored document _d_.  
Documents that have more occurrences of a given term receive a higher  
score. Note that _tf(t in q)_ is assumed to be _1_ and therefore it does  
not appear in this equation, However if a query contains twice the same  
term, there will be two term-queries with that same term and hence the  
computation would still be correct (although not very efficient). The  
default computation for _tf(t in d)_ in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float))  
is:

```
 tf(t in d)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float)>
  = frequency½

```

1. _idf(t)_ stands for Inverse Document Frequency. This value  
correlates to the inverse of _docFreq_ (the number of documents in which  
the term _t_ appears). This means rarer terms give higher contribution  
to the total score. _idf(t)_ appears for _t_ in both the query and the  
document, hence it is squared in the equation. The default computation for  
_idf(t)_ in _DefaultSimilarity_\<[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf)(int,  
int)\> is:

```
 idf(t)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf(int, 

```

int)\> = 1 + log ( numDocs ––––––––– docFreq+1 )

1. _coord(q,d)_ is a score factor based on how many of the query terms  
are found in the specified document. Typically, a document that contains  
more of the query's terms will receive a higher score than another document  
with fewer query terms. This is a search time factor computed in \*  
coord(q,d)\*\<[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#coord](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#coord)(int,  
int)\> by the Similarity in effect at search time.

2. \*queryNorm(q) \*is a normalizing factor used to make scores between  
queries comparable. This factor does not affect document ranking (since all  
ranked documents are multiplied by the same factor), but rather just  
attempts to make scores from different queries (or even different indexes)  
comparable. This is a search time factor computed by the Similarity in  
effect at search time. The default computation in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float))  
produces a _Euclidean norm_[http://en.wikipedia.org/wiki/Euclidean\_norm#Euclidean\_norm](http://en.wikipedia.org/wiki/Euclidean_norm#Euclidean_norm)  
:

```
 queryNorm(q) = queryNorm(sumOfSquaredWeights)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float)>
  = 1 –––––––––––––– sumOfSquaredWeights½

```

The sum of squared weights (of the query terms) is computed by the query  
_Weight_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Weight.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Weight.html)  
object. For example, a _boolean query_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/BooleanQuery.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/BooleanQuery.html)  
computes this value as:

```
 sumOfSquaredWeights<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Weight.html#sumOfSquaredWeights()>
  = q.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#getBoost()>
2 · ∑ ( idf(t)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_idf>
 · t.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_termBoost>
) 2 t in q 

```

1. _t.getBoost()_ is a search time boost of term _t_ in the query _q_ as  
specified in the query text (see _query syntax_\<[http://lucene.apache.org/java/3\_0\_2/queryparsersyntax.html#Boosting](http://lucene.apache.org/java/3_0_2/queryparsersyntax.html#Boosting) a Term\>),  
or as set by application calls to _setBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Query.html#setBoost(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#setBoost(float)).  
Notice that there is really no direct API for accessing a boost of one term  
in a multi term query, but rather multi terms are represented in a query as  
multi _TermQuery_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/TermQuery.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/TermQuery.html)  
objects, and so the boost of a term in the query is accessible by  
calling the sub-query _getBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Query.html#getBoost()](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#getBoost())  
.

2. _norm(t,d)_ encapsulates a few (indexing time) boost and length  
factors:

```
  - *Document boost* - set by calling *doc.setBoost()*<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Document.html#setBoost(float)>
   before adding the document to the index. 
  - *Field boost* - set by calling *field.setBoost()*<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Fieldable.html#setBoost(float)>
   before adding the field to a document. 
  - *lengthNorm(field)*<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String, 
  int)> - computed when the document is added to the index in 
  accordance with the number of tokens of this field in the document, so that 
  shorter fields contribute more to the score. LengthNorm is computed by the 
  Similarity class in effect at indexing.

```

When a document is added to the index, all the above factors are  
multiplied. If the document has multiple fields with the same name, all  
their boosts are multiplied together:

```
 norm(t,d) = doc.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Document.html#getBoost()>
 · lengthNorm(field)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String, 

```

int)\> · ∏ f.getBoost[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/document/Fieldable.html#getBoost()](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Fieldable.html#getBoost())  
() field _f_ in _d_ named as _t_

However the resulted _norm_ value is _encoded_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#encodeNorm(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#encodeNorm(float))  
as a single byte before being stored. At search time, the norm byte  
value is read from the index _directory_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/store/Directory.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/store/Directory.html)  
and _decoded_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#decodeNorm(byte)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#decodeNorm(byte))  
back to a float _norm_ value. This encoding/decoding, while reducing  
index size, comes with the price of precision loss - it is not guaranteed  
that _decode(encode(x)) = x_. For instance, _decode(encode(0.89)) = 0.75_  
.

Compression of norm values to a single byte saves memory at search time,  
because once a field is referenced at search time, its norms - for all  
documents - are maintained in memory.

The rationale supporting such lossy compression of norm values is that  
given the difficulty (and inaccuracy) of users to express their true  
information need by a query, only big differences matter.

Last, note that search time is too late to modify this _norm_ part of  
scoring, e.g. by using a different _Similarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)  
for search.

---

<div class="post-metadata">

**Author:** ![Stefan\_Nguyen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stefan_nguyen/32/2806_2.png) [@Stefan\_Nguyen](https://discuss.elastic.co/u/Stefan_Nguyen)\
**Post date:** [July 14, 2012, 1:18pm UTC](https://discuss.elastic.co/t/scoring-and-boost/6094/3 "2012-07-14T13:18:15Z")

</div>

Great answer!

Would you pls help with my situation?  
I have documents with a 'sentence' field of type String. How can I boost  
the score of documents with shorter 'sentence' values?

On Thursday, December 8, 2011 4:01:57 PM UTC+7, Ali Loghmani wrote:

> I believe this is a lucene thing more than anything. Try this URL:
> 
> - 
> 
> [Similarity (Lucene 3.0.3 API)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)  
> \*[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)
> 
> This is the core of it:
> 
> ```
> score(q,d) = coord(q,d)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_coord>
> 
> ```
> 
> · queryNorm(q)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_queryNorm](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_queryNorm)  
> · ∑ ( tf(t in d)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_tf](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_tf)  
> · idf(t)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_idf](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_idf)  
> 2 · t.getBoost()[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_termBoost](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_termBoost)  
> · norm(t,d)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_norm](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_norm)  
> ) t in q _Lucene Practical Scoring Function_
> 
> where
> 
> 1. _tf(t in d)_ correlates to the term's _frequency_, defined as the  
> number of times term _t_ appears in the currently scored document _d_.  
> Documents that have more occurrences of a given term receive a higher  
> score. Note that _tf(t in q)_ is assumed to be _1_ and therefore it  
> does not appear in this equation, However if a query contains twice the  
> same term, there will be two term-queries with that same term and hence the  
> computation would still be correct (although not very efficient). The  
> default computation for _tf(t in d)_ in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float))  
> is:
> 
> ```
> tf(t in d)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float)>
> = frequency½
> 
> ```
> 
> 1. _idf(t)_ stands for Inverse Document Frequency. This value  
> correlates to the inverse of _docFreq_ (the number of documents in  
> which the term _t_ appears). This means rarer terms give higher  
> contribution to the total score. _idf(t)_ appears for _t_ in both the  
> query and the document, hence it is squared in the equation. The default  
> computation for _idf(t)_ in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf(int,+int)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf(int,+int))  
> is:
> 
> ```
> idf(t)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf(int,+int)>  
> 
> ```
> 
> = 1 + log ( numDocs ––––––––– docFreq+1 )
> 
> 1. _coord(q,d)_ is a score factor based on how many of the query  
> terms are found in the specified document. Typically, a document that  
> contains more of the query's terms will receive a higher score than another  
> document with fewer query terms. This is a search time factor computed in  
> _coord(q,d)_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#coord(int,+int)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#coord(int,+int))  
> by the Similarity in effect at search time.
> 
> 2. \*queryNorm(q) \*is a normalizing factor used to make scores between  
> queries comparable. This factor does not affect document ranking (since all  
> ranked documents are multiplied by the same factor), but rather just  
> attempts to make scores from different queries (or even different indexes)  
> comparable. This is a search time factor computed by the Similarity in  
> effect at search time. The default computation in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float))  
> produces a _Euclidean norm_[http://en.wikipedia.org/wiki/Euclidean\_norm#Euclidean\_norm](http://en.wikipedia.org/wiki/Euclidean_norm#Euclidean_norm)  
> :
> 
> ```
> queryNorm(q) = queryNorm(sumOfSquaredWeights)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float)>
> = 1 –––––––––––––– sumOfSquaredWeights½
> 
> ```
> 
> The sum of squared weights (of the query terms) is computed by the  
> query _Weight_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Weight.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Weight.html)  
> object. For example, a _boolean query_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/BooleanQuery.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/BooleanQuery.html)  
> computes this value as:
> 
> ```
> sumOfSquaredWeights<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Weight.html#sumOfSquaredWeights()>
> = q.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#getBoost()>
> 2 · ∑ ( idf(t)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_idf>
> · t.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_termBoost>
> ) 2 t in q 
> 
> ```
> 
> 1. _t.getBoost()_ is a search time boost of term _t_ in the query _q_ as  
> specified in the query text (see _query syntax_[http://lucene.apache.org/java/3\_0\_2/queryparsersyntax.html#Boosting+a+Term](http://lucene.apache.org/java/3_0_2/queryparsersyntax.html#Boosting+a+Term)),  
> or as set by application calls to _setBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Query.html#setBoost(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#setBoost(float)).  
> Notice that there is really no direct API for accessing a boost of one term  
> in a multi term query, but rather multi terms are represented in a query as  
> multi _TermQuery_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/TermQuery.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/TermQuery.html)  
> objects, and so the boost of a term in the query is accessible by  
> calling the sub-query _getBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Query.html#getBoost()](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#getBoost())  
> .
> 
> 2. _norm(t,d)_ encapsulates a few (indexing time) boost and length  
> factors:  
> - _Document boost_ - set by calling _doc.setBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/document/Document.html#setBoost(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Document.html#setBoost(float))  
> before adding the document to the index.  
> - _Field boost_ - set by calling _field.setBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/document/Fieldable.html#setBoost(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Fieldable.html#setBoost(float))  
> before adding the field to a document.  
> - _lengthNorm(field)_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String,+int)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String,+int))
> 
> When a document is added to the index, all the above factors are  
> multiplied. If the document has multiple fields with the same name, all  
> their boosts are multiplied together:
> 
> ```
> norm(t,d) = doc.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Document.html#getBoost()>
> · lengthNorm(field)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String,+int)>
> · ∏ f.getBoost<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Fieldable.html#getBoost()>
> 
> ```
> 
> () field _f_ in _d_ named as _t_
> 
> However the resulted _norm_ value is _encoded_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#encodeNorm(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#encodeNorm(float))  
> as a single byte before being stored. At search time, the norm byte  
> value is read from the index _directory_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/store/Directory.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/store/Directory.html)  
> and _decoded_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#decodeNorm(byte)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#decodeNorm(byte))  
> back to a float _norm_ value. This encoding/decoding, while reducing  
> index size, comes with the price of precision loss - it is not guaranteed  
> that _decode(encode(x)) = x_. For instance, _decode(encode(0.89)) =  
> 0.75_.
> 
> Compression of norm values to a single byte saves memory at search  
> time, because once a field is referenced at search time, its norms - for  
> all documents - are maintained in memory.
> 
> The rationale supporting such lossy compression of norm values is that  
> given the difficulty (and inaccuracy) of users to express their true  
> information need by a query, only big differences matter.
> 
> Last, note that search time is too late to modify this _norm_ part of  
> scoring, e.g. by using a different _Similarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)  
> for search.

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [July 18, 2012, 4:59pm UTC](https://discuss.elastic.co/t/scoring-and-boost/6094/4 "2012-07-18T16:59:57Z")

</div>

Stefan,

By the characteristics of the TD-IDF formula, shorter sentences should be  
boosted. Lucene (and therefore Elasticsearch) uses norms and they are  
enabled by default.

Here is a good explanation of norms and term frequencies in Lucene:

> **[Home](https://lucidworks.com/)**
>
> Lucidworks' Fusion platform uses industry-leading search technology to power search & discovery for the largest & most successful companies. Request a demo today.

Cheers,

Ivan

On Sat, Jul 14, 2012 at 6:18 AM, Stefan Nguyen [stnguyenvn@gmail.com](mailto:stnguyenvn@gmail.com) wrote:

> Great answer!
> 
> Would you pls help with my situation?  
> I have documents with a 'sentence' field of type String. How can I boost  
> the score of documents with shorter 'sentence' values?
> 
> On Thursday, December 8, 2011 4:01:57 PM UTC+7, Ali Loghmani wrote:
> 
> > I believe this is a lucene thing more than anything. Try this URL:
> > 
> > _[Index of /\_\_root/docs.lucene.apache.org/core/3\_0\_3/api/core/org/apache](http://lucene.apache.org/java/3_0_2/api/core/org/apache/)  
> > lucene/search/Similarity.html_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)
> > 
> > This is the core of it:
> > 
> > ```
> > score(q,d) = coord(q,d)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_coord>
> > 
> > ```
> > 
> > · queryNorm(q)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_queryNorm](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_queryNorm)  
> > \*\* · ∑ ( tf(t in d)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_tf](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_tf)  
> > · idf(t)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_idf](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_idf)  
> > 2 · t.getBoost(\*\*)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_termBoost](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_termBoost)  
> > · norm(t,d)[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#formula\_norm](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_norm)  
> > ) t in q _Lucene Practical Scoring Function_
> > 
> > where
> > 
> > 1. _tf(t in d)_ correlates to the term's _frequency_, defined as the  
> > number of times term _t_ appears in the currently scored document _d_.  
> > Documents that have more occurrences of a given term receive a higher  
> > score. Note that _tf(t in q)_ is assumed to be _1_ and therefore it  
> > does not appear in this equation, However if a query contains twice the  
> > same term, there will be two term-queries with that same term and hence the  
> > computation would still be correct (although not very efficient). The  
> > default computation for _tf(t in d)_ in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float))  
> > is:
> > 
> > ```
> > tf(t in d)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#tf(float)>
> > = frequency½
> > 
> > ```
> > 
> > 1. _idf(t)_ stands for Inverse Document Frequency. This value  
> > correlates to the inverse of _docFreq_ (the number of documents in  
> > which the term _t_ appears). This means rarer terms give higher  
> > contribution to the total score. _idf(t)_ appears for _t_ in both the  
> > query and the document, hence it is squared in the equation. The default  
> > computation for _idf(t)_ in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf(int,+int)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf(int,+int))  
> > is:
> > 
> > ```
> > idf(t)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#idf(int,+int)>
> > 
> > ```
> > 
> > = 1 + log ( numDocs ––––––––– docFreq+1 )
> > 
> > 1. _coord(q,d)_ is a score factor based on how many of the query  
> > terms are found in the specified document. Typically, a document that  
> > contains more of the query's terms will receive a higher score than another  
> > document with fewer query terms. This is a search time factor computed in  
> > _coord(q,d)_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#coord(int,+int)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#coord(int,+int))  
> > by the Similarity in effect at search time.
> > 
> > 2. \*queryNorm(q) \*is a normalizing factor used to make scores between  
> > queries comparable. This factor does not affect document ranking (since all  
> > ranked documents are multiplied by the same factor), but rather just  
> > attempts to make scores from different queries (or even different indexes)  
> > comparable. This is a search time factor computed by the Similarity in  
> > effect at search time. The default computation in _DefaultSimilarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float))  
> > produces a _Euclidean norm_[http://en.wikipedia.org/wiki/Euclidean\_norm#Euclidean\_norm](http://en.wikipedia.org/wiki/Euclidean_norm#Euclidean_norm)  
> > :
> > 
> > ```
> > queryNorm(q) = queryNorm(**sumOfSquaredWeights)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/DefaultSimilarity.html#queryNorm(float)>
> > = 1 –––––––––––––– sumOfSquaredWeights½
> > 
> > ```
> > 
> > The sum of squared weights (of the query terms) is computed by the  
> > query _Weight_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Weight.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Weight.html)  
> > object. For example, a _boolean query_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/BooleanQuery.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/BooleanQuery.html)  
> > computes this value as:
> > 
> > ```
> > sumOfSquaredWeights<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Weight.html#sumOfSquaredWeights()>
> > = q.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#getBoost()>
> > 2 · ∑ ( idf(t)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_idf>
> > · t.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#formula_termBoost>
> > ) 2 t in q
> > 
> > ```
> > 
> > 1. _t.getBoost()_ is a search time boost of term _t_ in the query _q_  
> > as specified in the query text (see _query syntax_[http://lucene.apache.org/java/3\_0\_2/queryparsersyntax.html#Boosting+a+Term](http://lucene.apache.org/java/3_0_2/queryparsersyntax.html#Boosting+a+Term)),  
> > or as set by application calls to _setBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Query.html#setBoost(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#setBoost(float)).  
> > Notice that there is really no direct API for accessing a boost of one term  
> > in a multi term query, but rather multi terms are represented in a query as  
> > multi _TermQuery_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/TermQuery.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/TermQuery.html)  
> > objects, and so the boost of a term in the query is accessible by  
> > calling the sub-query _getBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Query.html#getBoost()](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Query.html#getBoost())  
> > .
> > 
> > 2. _norm(t,d)_ encapsulates a few (indexing time) boost and length  
> > factors:  
> > - _Document boost_ - set by calling _doc.setBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/document/Document.html#setBoost(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Document.html#setBoost(float))  
> > before adding the document to the index.  
> > - _Field boost_ - set by calling _field.setBoost()_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/document/Fieldable.html#setBoost(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Fieldable.html#setBoost(float))  
> > befor\*\*e adding the field to a document.  
> > - _lengthNorm(field)_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String,+int)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String,+int))
> > 
> > When a document is added to the index, all the above factors are  
> > multiplied. If the document has multiple fields with the same name, all  
> > their boosts are multiplied together:
> > 
> > ```
> > norm(t,d) = doc.getBoost()<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Document.html#getBoost()>
> > · lengthNor**m(field)<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#lengthNorm(java.lang.String,+int)>
> > · ∏ f.getBoost<http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/document/Fieldable.html#getBoost()>
> > 
> > ```
> > 
> > () field _f_ in _d_ named as _t_
> > 
> > However the resulted _norm_ value is _encoded_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#encodeNorm(float)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#encodeNorm(float))  
> > as a single byte before being stored. At search time, the norm byte  
> > value is read from the index _directory_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/store/Directory.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/store/Directory.html)  
> > and _decoded_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html#decodeNorm(byte)](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html#decodeNorm(byte))  
> > ba\*\*ck to a float _norm_ value. This encoding/decoding, while  
> > reducing index size, comes with the price of precision loss - it is not  
> > guaranteed that _decode(encode(x)) = x_. For instance, _decode(encode(0.89))  
> > = 0.75_.
> > 
> > Compression of norm values to a single byte saves memory at search  
> > time, because once a field is referenced at search time, its norms - for  
> > all documents - are maintained in memory.
> > 
> > The rationale supporting such lossy compression of norm values is  
> > that given the difficulty (and inaccuracy) of users to express their true  
> > information need by a query, only big differences matter.
> > 
> > Last, note that search time is too late to modify this _norm_ part of  
> > scoring, e.g. by using a different _Similarity_[http://lucene.apache.org/java/3\_0\_2/api/core/org/apache/lucene/search/Similarity.html](http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html)  
> > for search.

---

<div class="post-metadata">

**Author:** ![timg](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@timg](https://discuss.elastic.co/u/timg)\
**Post date:** [February 26, 2016, 9:41am UTC](https://discuss.elastic.co/t/scoring-and-boost/6094/5 "2016-02-26T09:41:52Z")

</div>

can you site some more examples of numbers and their equivalent when encoded/decoded to/from single byte pls. thanks

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [February 26, 2016, 4:00pm UTC](https://discuss.elastic.co/t/scoring-and-boost/6094/6 "2016-02-26T16:00:44Z")

</div>

Wow, this is an old topic. I have never seen examples of the actual numbers  
of the encoded norm, except for examples showing the lossy nature of the  
encode-decode-encode process.

I am assuming you are talking about the length norm and not the overall  
norm value which includes boosting. Do not use index-time boosting, use  
query-time boosting and/or function scores.

You might have better look on the Lucene list since this is more of an  
internal Lucene question. Elasticsearch users tend to be more interested in  
scaling and how to make Kibana look better than in pure search. The code  
for encoding norms is in ClassicSimilarity:

> <https://github.com/apache/lucene-solr/blob/master/lucene/core/src/java/org/apache/lucene/search/similarities/ClassicSimilarity.java>

There is a static lookup table, but I have not executed the code to see  
what those values actually are.

I should point out that Lucene 6, and therefore Elasticsearch 5.0, will  
have the BM25 similarity enabled by default. BM25 provides a more  
tunable/adaptive approach to field length normalization. Before going down  
the route of attempting to alter the length norm yourself, it might be  
worthwhile to check out BM25.

Cheers,

Ivan

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:13pm UTC](https://discuss.elastic.co/t/scoring-and-boost/6094/7 "2017-07-05T23:13:19Z")

</div>


