# Normalizing MLT score

**URL:** <https://discuss.elastic.co/t/normalizing-mlt-score/9899>\
**Category:** Elasticsearch\
**Created:** [November 30, 2012, 7:43am UTC](https://discuss.elastic.co/t/normalizing-mlt-score/9899 "2012-11-30T07:43:23Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Atharva\_Patel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atharva_patel/32/2414_2.png) [@Atharva\_Patel](https://discuss.elastic.co/u/Atharva_Patel)\
**Post date:** [November 30, 2012, 7:43am UTC](https://discuss.elastic.co/t/normalizing-mlt-score/9899/1 "2012-11-30T07:43:23Z")

</div>

I am using MLT queries to find out similar documents. I have a case where I  
would like to set a threshold on the score for deciding which documents  
should be considered as similar to the given document passed in the like  
text.

In the response hits I am observing the scores ranging from 0 to 2.5. The  
2.5 is the upper limit of the few test cases that I have considered while  
in development. In production it may even go higher! Therefore I am  
interested in knowing if there is a way to normalize the score to bring  
them between 0 and 1. Naive strategy of dividing each hit score by max  
score at the client side will be useless as it will produce score 1.0 for  
the first hit(the one with highest score) in the ranked hits, so it will  
always pass the threshold (say 0.3).

It can be also useful if I can some how predict the highest possible score  
on my MLT query based on some internal formula being used by MLT for  
scoring.

Can somebody please help me with these approaches?

Thanks!

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:02am UTC](https://discuss.elastic.co/t/normalizing-mlt-score/9899/2 "2017-07-06T03:02:03Z")

</div>


