# MoreLikeThis percent\_terms\_to\_match

**URL:** <https://discuss.elastic.co/t/morelikethis-percent-terms-to-match/14099>\
**Category:** Elasticsearch\
**Created:** [October 24, 2013, 6:30pm UTC](https://discuss.elastic.co/t/morelikethis-percent-terms-to-match/14099 "2013-10-24T18:30:04Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Justin\_Treher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_treher/32/45243_2.png) [@Justin\_Treher](https://discuss.elastic.co/u/Justin_Treher)\
**Post date:** [October 24, 2013, 6:30pm UTC](https://discuss.elastic.co/t/morelikethis-percent-terms-to-match/14099/1 "2013-10-24T18:30:04Z")

</div>

Hello,

I have a query like this with 0.93. I don't quite understand what  
percent\_terms\_to\_match is doing here. While I have it set to .7, the  
"explain" is clearly telling me that it is only matching one of the three  
terms in the like text. Why would it not filter this match out when 1 of 3  
words match? My interpretation of the documents is that it builds a bool  
query with every term in the like\_text and then would set a minimum should  
match based on the percent. In this case, with three words, if all three  
don't match, it should never give results. However, I suspect my  
interpretation is wrong. Thanks!

{  
"query": {  
"more\_like\_this": {  
"fields": [  
"title\_alias"  
],  
"like\_text": "fish tree lounge",  
"min\_term\_freq": 0,  
"max\_query\_terms": 25,  
"percent\_terms\_to\_match": 0.7,  
"min\_doc\_freq": 1,  
"analyzer": "standard"  
}  
},"explain":true  
}

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood1/32/101255_2.png) [@Mark\_Harwood1](https://discuss.elastic.co/u/Mark_Harwood1)\
**Post date:** [October 25, 2013, 10:27am UTC](https://discuss.elastic.co/t/morelikethis-percent-terms-to-match/14099/2 "2013-10-25T10:27:36Z")

</div>

It looks like the % of terms to match is based on the number of input terms  
that exist in a shard with the required frequency and not just the total  
number of words in the input string.  
Terms that have zero doc frequency on a shard (ie. do not even exist in the  
index) are not added to the final boolean query so on that shard you may  
have only one or two query clauses and not 3. The logic on that shard is  
then that 70% of the 2 clauses relevant to that shard must match giving the  
possibility of a match on a single term.

Cheers  
Mark

On Thursday, October 24, 2013 7:30:04 PM UTC+1, Justin Treher wrote:

> Hello,
> 
> I have a query like this with 0.93. I don't quite understand what  
> percent\_terms\_to\_match is doing here. While I have it set to .7, the  
> "explain" is clearly telling me that it is only matching one of the three  
> terms in the like text. Why would it not filter this match out when 1 of 3  
> words match? My interpretation of the documents is that it builds a bool  
> query with every term in the like\_text and then would set a minimum should  
> match based on the percent. In this case, with three words, if all three  
> don't match, it should never give results. However, I suspect my  
> interpretation is wrong. Thanks!
> 
> {  
> "query": {  
> "more\_like\_this": {  
> "fields": [  
> "title\_alias"  
> ],  
> "like\_text": "fish tree lounge",  
> "min\_term\_freq": 0,  
> "max\_query\_terms": 25,  
> "percent\_terms\_to\_match": 0.7,  
> "min\_doc\_freq": 1,  
> "analyzer": "standard"  
> }  
> },"explain":true  
> }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Justin\_Treher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_treher/32/45243_2.png) [@Justin\_Treher](https://discuss.elastic.co/u/Justin_Treher)\
**Post date:** [October 25, 2013, 12:55pm UTC](https://discuss.elastic.co/t/morelikethis-percent-terms-to-match/14099/3 "2013-10-25T12:55:56Z")

</div>

Thanks. I thought something like this must have been happening, but I was  
not quite sure of the intent of the parameter and function altogether. I  
don't think it does what I thought, so I will stick with a match with min  
should match parameter.  
On Oct 25, 2013 6:27 AM, "Mark Harwood" [markharwood@gmail.com](mailto:markharwood@gmail.com) wrote:

> It looks like the % of terms to match is based on the number of input  
> terms that exist in a shard with the required frequency and not just the  
> total number of words in the input string.  
> Terms that have zero doc frequency on a shard (ie. do not even exist in  
> the index) are not added to the final boolean query so on that shard you  
> may have only one or two query clauses and not 3. The logic on that shard  
> is then that 70% of the 2 clauses relevant to that shard must match giving  
> the possibility of a match on a single term.
> 
> Cheers  
> Mark
> 
> On Thursday, October 24, 2013 7:30:04 PM UTC+1, Justin Treher wrote:
> 
> > Hello,
> > 
> > I have a query like this with 0.93. I don't quite understand what  
> > percent\_terms\_to\_match is doing here. While I have it set to .7, the  
> > "explain" is clearly telling me that it is only matching one of the three  
> > terms in the like text. Why would it not filter this match out when 1 of 3  
> > words match? My interpretation of the documents is that it builds a bool  
> > query with every term in the like\_text and then would set a minimum should  
> > match based on the percent. In this case, with three words, if all three  
> > don't match, it should never give results. However, I suspect my  
> > interpretation is wrong. Thanks!
> > 
> > {  
> > "query": {  
> > "more\_like\_this": {  
> > "fields": [  
> > "title\_alias"  
> > ],  
> > "like\_text": "fish tree lounge",  
> > "min\_term\_freq": 0,  
> > "max\_query\_terms": 25,  
> > "percent\_terms\_to\_match": 0.7,  
> > "min\_doc\_freq": 1,  
> > "analyzer": "standard"  
> > }  
> > },"explain":true  
> > }
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/rSre9kSXAqQ/unsubscribe](https://groups.google.com/d/topic/elasticsearch/rSre9kSXAqQ/unsubscribe).  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:10am UTC](https://discuss.elastic.co/t/morelikethis-percent-terms-to-match/14099/4 "2017-07-06T02:10:35Z")

</div>


