# More like this scoring algorithm unclear

**URL:** <https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136>\
**Category:** Elasticsearch\
**Created:** [January 8, 2014, 8:04am UTC](https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136 "2014-01-08T08:04:47Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Maarten\_Roosendaal](https://avatars.discourse-cdn.com/v4/letter/m/49beb7/32.png) [@Maarten\_Roosendaal](https://discuss.elastic.co/u/Maarten_Roosendaal)\
**Post date:** [January 8, 2014, 8:04am UTC](https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136/1 "2014-01-08T08:04:47Z")

</div>

Hi,

I have a question about why the 'more like this' algorithm scores documents  
higher than others, while they are (at first glance) the same.

What i've done is index wishlist-documents which contain 1 property:  
product\_id, this property contains an array of product\_id's (e.g. [1234,  
4444, 5555, 6666]. What i'm trying to do is find similair wishlist for a  
given wishlist with id x. The MLT API seems to work, it returns other  
documents which contain at least 1 of the product\_id's from the original  
list.

But what is see is that, for example. i get 10 hits, the first 6 hits  
contain the same (and only 1) product\_id, this product\_id is present in the  
original wishlist. What i would expect is that the score of the first 6 is  
the same. However what i see is that only the first 2 have the same, the  
next 2 a lower score and the next 2 even lower. Why is this?

Also, i'm trying to write the MLT API as an MLT query, but somehow it  
doesn't work. I would expect that i need to take the entire content of the  
original product\_id property and feed is as input for the 'like\_text'. The  
documentation is not very clear and doesn't provide examples so i'm a  
little lost.

Hope someone can give some pointers.

Thanks,  
Maarten

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/0e2827b2-5a21-4cff-b773-ebdd861c5972%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/0e2827b2-5a21-4cff-b773-ebdd861c5972%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Justin\_Treher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_treher/32/45243_2.png) [@Justin\_Treher](https://discuss.elastic.co/u/Justin_Treher)\
**Post date:** [January 8, 2014, 3:14pm UTC](https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136/2 "2014-01-08T15:14:25Z")

</div>

Hey Maarten,

I would use the "explain":true option to see just why your documents are  
being scored higher than others. MoreLikeThis using the same fulltext  
scoring as far as I know, so term position would affect score.

[http://lucene.apache.org/core/3\_0\_3/api/contrib-queries/org/apache/lucene/search/similar/MoreLikeThis.html](http://lucene.apache.org/core/3_0_3/api/contrib-queries/org/apache/lucene/search/similar/MoreLikeThis.html)

Justin

On Wednesday, January 8, 2014 3:04:47 AM UTC-5, Maarten Roosendaal wrote:

> Hi,
> 
> I have a question about why the 'more like this' algorithm scores  
> documents higher than others, while they are (at first glance) the same.
> 
> What i've done is index wishlist-documents which contain 1 property:  
> product\_id, this property contains an array of product\_id's (e.g. [1234,  
> 4444, 5555, 6666]. What i'm trying to do is find similair wishlist for a  
> given wishlist with id x. The MLT API seems to work, it returns other  
> documents which contain at least 1 of the product\_id's from the original  
> list.
> 
> But what is see is that, for example. i get 10 hits, the first 6 hits  
> contain the same (and only 1) product\_id, this product\_id is present in the  
> original wishlist. What i would expect is that the score of the first 6 is  
> the same. However what i see is that only the first 2 have the same, the  
> next 2 a lower score and the next 2 even lower. Why is this?
> 
> Also, i'm trying to write the MLT API as an MLT query, but somehow it  
> doesn't work. I would expect that i need to take the entire content of the  
> original product\_id property and feed is as input for the 'like\_text'. The  
> documentation is not very clear and doesn't provide examples so i'm a  
> little lost.
> 
> Hope someone can give some pointers.
> 
> Thanks,  
> Maarten

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a0e9a58d-89e7-4084-b7ed-7f34c8514ce5%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a0e9a58d-89e7-4084-b7ed-7f34c8514ce5%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Maarten\_Roosendaal](https://avatars.discourse-cdn.com/v4/letter/m/49beb7/32.png) [@Maarten\_Roosendaal](https://discuss.elastic.co/u/Maarten_Roosendaal)\
**Post date:** [January 8, 2014, 3:50pm UTC](https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136/3 "2014-01-08T15:50:53Z")

</div>

Hi,

Thanks, i'm not quite sure how to do that. I'm using:  
[http://localhost:9200/lists/list/](http://localhost:9200/lists/list/)[id of  
list]/\_mlt?mlt\_field=product\_id&min\_term\_freq=1&min\_doc\_freq=1

the body does not seem to be respected (i'm using the elasticsearch head  
plugin) if i ad:  
{  
"explain": true  
}

i've been trying to rewrite the mlt api as an mlt query but no luck so far.  
Any suggestions?

Thanks,  
Maarten

Op woensdag 8 januari 2014 16:14:25 UTC+1 schreef Justin Treher:

> Hey Maarten,
> 
> I would use the "explain":true option to see just why your documents are  
> being scored higher than others. MoreLikeThis using the same fulltext  
> scoring as far as I know, so term position would affect score.
> 
> [MoreLikeThis (Lucene 3.0.3 API)](http://lucene.apache.org/core/3_0_3/api/contrib-queries/org/apache/lucene/search/similar/MoreLikeThis.html)
> 
> Justin
> 
> On Wednesday, January 8, 2014 3:04:47 AM UTC-5, Maarten Roosendaal wrote:
> 
> > Hi,
> > 
> > I have a question about why the 'more like this' algorithm scores  
> > documents higher than others, while they are (at first glance) the same.
> > 
> > What i've done is index wishlist-documents which contain 1 property:  
> > product\_id, this property contains an array of product\_id's (e.g. [1234,  
> > 4444, 5555, 6666]. What i'm trying to do is find similair wishlist for a  
> > given wishlist with id x. The MLT API seems to work, it returns other  
> > documents which contain at least 1 of the product\_id's from the original  
> > list.
> > 
> > But what is see is that, for example. i get 10 hits, the first 6 hits  
> > contain the same (and only 1) product\_id, this product\_id is present in the  
> > original wishlist. What i would expect is that the score of the first 6 is  
> > the same. However what i see is that only the first 2 have the same, the  
> > next 2 a lower score and the next 2 even lower. Why is this?
> > 
> > Also, i'm trying to write the MLT API as an MLT query, but somehow it  
> > doesn't work. I would expect that i need to take the entire content of the  
> > original product\_id property and feed is as input for the 'like\_text'. The  
> > documentation is not very clear and doesn't provide examples so i'm a  
> > little lost.
> > 
> > Hope someone can give some pointers.
> > 
> > Thanks,  
> > Maarten

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/5f1b4a50-8862-42e8-a3a8-532f88757a48%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/5f1b4a50-8862-42e8-a3a8-532f88757a48%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Maarten\_Roosendaal](https://avatars.discourse-cdn.com/v4/letter/m/49beb7/32.png) [@Maarten\_Roosendaal](https://discuss.elastic.co/u/Maarten_Roosendaal)\
**Post date:** [January 8, 2014, 6:20pm UTC](https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136/4 "2014-01-08T18:20:05Z")

</div>

scoring algorithm is still vague but i got the query to act like the API,  
although the results are different so i'm still doing it wrong, here's an  
example:  
{  
"explain": true,  
"query": {  
"more\_like\_this": {  
"fields": [  
"PRODUCT\_ID"  
],  
"like\_text": "1000004004855475 1001004002067765 1002004000094210  
1002004004499883",  
"min\_term\_freq": 1,  
"min\_doc\_freq": 1,  
"max\_query\_terms": 1,  
"percent\_terms\_to\_match": 0.5  
}  
},  
"from": 0,  
"size": 50,  
"sort": ,  
"facets": {}  
}

the like\_text contains product\_id's from a wishlist for which i want to  
find similair lists

Op woensdag 8 januari 2014 16:50:53 UTC+1 schreef Maarten Roosendaal:

> Hi,
> 
> Thanks, i'm not quite sure how to do that. I'm using:  
> [http://localhost:9200/lists/list/](http://localhost:9200/lists/list/)[id of  
> list]/\_mlt?mlt\_field=product\_id&min\_term\_freq=1&min\_doc\_freq=1
> 
> the body does not seem to be respected (i'm using the elasticsearch head  
> plugin) if i ad:  
> {  
> "explain": true  
> }
> 
> i've been trying to rewrite the mlt api as an mlt query but no luck so  
> far. Any suggestions?
> 
> Thanks,  
> Maarten
> 
> Op woensdag 8 januari 2014 16:14:25 UTC+1 schreef Justin Treher:
> 
> > Hey Maarten,
> > 
> > I would use the "explain":true option to see just why your documents are  
> > being scored higher than others. MoreLikeThis using the same fulltext  
> > scoring as far as I know, so term position would affect score.
> > 
> > [MoreLikeThis (Lucene 3.0.3 API)](http://lucene.apache.org/core/3_0_3/api/contrib-queries/org/apache/lucene/search/similar/MoreLikeThis.html)
> > 
> > Justin
> > 
> > On Wednesday, January 8, 2014 3:04:47 AM UTC-5, Maarten Roosendaal wrote:
> > 
> > > Hi,
> > > 
> > > I have a question about why the 'more like this' algorithm scores  
> > > documents higher than others, while they are (at first glance) the same.
> > > 
> > > What i've done is index wishlist-documents which contain 1 property:  
> > > product\_id, this property contains an array of product\_id's (e.g. [1234,  
> > > 4444, 5555, 6666]. What i'm trying to do is find similair wishlist for a  
> > > given wishlist with id x. The MLT API seems to work, it returns other  
> > > documents which contain at least 1 of the product\_id's from the original  
> > > list.
> > > 
> > > But what is see is that, for example. i get 10 hits, the first 6 hits  
> > > contain the same (and only 1) product\_id, this product\_id is present in the  
> > > original wishlist. What i would expect is that the score of the first 6 is  
> > > the same. However what i see is that only the first 2 have the same, the  
> > > next 2 a lower score and the next 2 even lower. Why is this?
> > > 
> > > Also, i'm trying to write the MLT API as an MLT query, but somehow it  
> > > doesn't work. I would expect that i need to take the entire content of the  
> > > original product\_id property and feed is as input for the 'like\_text'. The  
> > > documentation is not very clear and doesn't provide examples so i'm a  
> > > little lost.
> > > 
> > > Hope someone can give some pointers.
> > > 
> > > Thanks,  
> > > Maarten

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c7032391-2456-47a0-a3b8-1f5fe61127e7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c7032391-2456-47a0-a3b8-1f5fe61127e7%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Alex\_Ksikes1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alex_ksikes1/32/1534_2.png) [@Alex\_Ksikes1](https://discuss.elastic.co/u/Alex_Ksikes1)\
**Post date:** [May 6, 2014, 11:33am UTC](https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136/5 "2014-05-06T11:33:47Z")

</div>

Hi Maarten,

Your 'like\_text' is analyzed, the same way your 'product\_id' field is  
analyzed, unless specified by 'analyzer'. I would recommend setting  
'percent\_terms\_to\_match' to 0. However, if you are only searching over  
product ids then a simple boolean query would do. If not, then I would  
create a boolean query where each clause is a 'more like this field' for  
each field of the queried document. This is actually what the mlt API does.

Cheers,

Alex

On Wednesday, January 8, 2014 7:20:05 PM UTC+1, Maarten Roosendaal wrote:

> scoring algorithm is still vague but i got the query to act like the API,  
> although the results are different so i'm still doing it wrong, here's an  
> example:  
> {  
> "explain": true,  
> "query": {  
> "more\_like\_this": {  
> "fields": [  
> "PRODUCT\_ID"  
> ],  
> "like\_text": "1000004004855475 1001004002067765 1002004000094210  
> 1002004004499883",  
> "min\_term\_freq": 1,  
> "min\_doc\_freq": 1,  
> "max\_query\_terms": 1,  
> "percent\_terms\_to\_match": 0.5  
> }  
> },  
> "from": 0,  
> "size": 50,  
> "sort": ,  
> "facets": {}  
> }
> 
> the like\_text contains product\_id's from a wishlist for which i want to  
> find similair lists
> 
> Op woensdag 8 januari 2014 16:50:53 UTC+1 schreef Maarten Roosendaal:
> 
> > Hi,
> > 
> > Thanks, i'm not quite sure how to do that. I'm using:  
> > [http://localhost:9200/lists/list/](http://localhost:9200/lists/list/)[id of  
> > list]/\_mlt?mlt\_field=product\_id&min\_term\_freq=1&min\_doc\_freq=1
> > 
> > the body does not seem to be respected (i'm using the elasticsearch head  
> > plugin) if i ad:  
> > {  
> > "explain": true  
> > }
> > 
> > i've been trying to rewrite the mlt api as an mlt query but no luck so  
> > far. Any suggestions?
> > 
> > Thanks,  
> > Maarten
> > 
> > Op woensdag 8 januari 2014 16:14:25 UTC+1 schreef Justin Treher:
> > 
> > > Hey Maarten,
> > > 
> > > I would use the "explain":true option to see just why your documents are  
> > > being scored higher than others. MoreLikeThis using the same fulltext  
> > > scoring as far as I know, so term position would affect score.
> > > 
> > > [MoreLikeThis (Lucene 3.0.3 API)](http://lucene.apache.org/core/3_0_3/api/contrib-queries/org/apache/lucene/search/similar/MoreLikeThis.html)
> > > 
> > > Justin
> > > 
> > > On Wednesday, January 8, 2014 3:04:47 AM UTC-5, Maarten Roosendaal wrote:
> > > 
> > > > Hi,
> > > > 
> > > > I have a question about why the 'more like this' algorithm scores  
> > > > documents higher than others, while they are (at first glance) the same.
> > > > 
> > > > What i've done is index wishlist-documents which contain 1 property:  
> > > > product\_id, this property contains an array of product\_id's (e.g. [1234,  
> > > > 4444, 5555, 6666]. What i'm trying to do is find similair wishlist for a  
> > > > given wishlist with id x. The MLT API seems to work, it returns other  
> > > > documents which contain at least 1 of the product\_id's from the original  
> > > > list.
> > > > 
> > > > But what is see is that, for example. i get 10 hits, the first 6 hits  
> > > > contain the same (and only 1) product\_id, this product\_id is present in the  
> > > > original wishlist. What i would expect is that the score of the first 6 is  
> > > > the same. However what i see is that only the first 2 have the same, the  
> > > > next 2 a lower score and the next 2 even lower. Why is this?
> > > > 
> > > > Also, i'm trying to write the MLT API as an MLT query, but somehow it  
> > > > doesn't work. I would expect that i need to take the entire content of the  
> > > > original product\_id property and feed is as input for the 'like\_text'. The  
> > > > documentation is not very clear and doesn't provide examples so i'm a  
> > > > little lost.
> > > > 
> > > > Hope someone can give some pointers.
> > > > 
> > > > Thanks,  
> > > > Maarten

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/91734252-74d0-4001-becc-a184af0f2997%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/91734252-74d0-4001-becc-a184af0f2997%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:31am UTC](https://discuss.elastic.co/t/more-like-this-scoring-algorithm-unclear/15136/6 "2017-07-06T01:31:24Z")

</div>


