# MoreLikeThis query, what does percent\_terms\_to\_match do?

**URL:** <https://discuss.elastic.co/t/morelikethis-query-what-does-percent-terms-to-match-do/7789>\
**Category:** Elasticsearch\
**Created:** [May 21, 2012, 4:30pm UTC](https://discuss.elastic.co/t/morelikethis-query-what-does-percent-terms-to-match-do/7789 "2012-05-21T16:30:46Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nick\_Dunn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_dunn/32/2315_2.png) [@Nick\_Dunn](https://discuss.elastic.co/u/Nick_Dunn)\
**Post date:** [May 21, 2012, 4:30pm UTC](https://discuss.elastic.co/t/morelikethis-query-what-does-percent-terms-to-match-do/7789/1 "2012-05-21T16:30:46Z")

</div>

I'm trying out morelikethis  
([http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html)) and  
it's working well. So easy 🙂

I'm finding entries by related tags so I dropped `min_doc_freq` to 1 (one  
or more tag required for a match) and `max_query_terms` to 100 (an entry  
could be tagged with up to 100 tags) however from the docs it's not clear  
to me what `percent_terms_to_match` does:

The percentage of terms to match on (float value). Defaults to 0.3 (30  
percent).

Could someone explain this in other words please, perhaps an example of  
what might happen if I increase or decrease from the default? When I try it  
on my sample data it increases/reduces the number of hits and doesn't seem  
to affect the score of each hit, so I'm just not sure what it's doing.

Thanks.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 23, 2012, 10:30pm UTC](https://discuss.elastic.co/t/morelikethis-query-what-does-percent-terms-to-match-do/7789/2 "2012-05-23T22:30:51Z")

</div>

Effectively, what happens in the more like this query is that it builds a  
big boolean query with should clauses for each term. The percent terms to  
match means that out of all the terms built, at least X percent should  
match (effectively, setting the minimum\_should\_match parameter on it).

On Mon, May 21, 2012 at 6:30 PM, Nick Dunn [nick@nick-dunn.co.uk](mailto:nick@nick-dunn.co.uk) wrote:

> I'm trying out morelikethis (  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html))  
> and it's working well. So easy 🙂
> 
> I'm finding entries by related tags so I dropped `min_doc_freq` to 1 (one  
> or more tag required for a match) and `max_query_terms` to 100 (an entry  
> could be tagged with up to 100 tags) however from the docs it's not clear  
> to me what `percent_terms_to_match` does:
> 
> The percentage of terms to match on (float value). Defaults to 0.3 (30  
> percent).
> 
> Could someone explain this in other words please, perhaps an example of  
> what might happen if I increase or decrease from the default? When I try it  
> on my sample data it increases/reduces the number of hits and doesn't seem  
> to affect the score of each hit, so I'm just not sure what it's doing.
> 
> Thanks.

---

<div class="post-metadata">

**Author:** ![Nick\_Dunn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_dunn/32/2315_2.png) [@Nick\_Dunn](https://discuss.elastic.co/u/Nick_Dunn)\
**Post date:** [May 24, 2012, 8:31am UTC](https://discuss.elastic.co/t/morelikethis-query-what-does-percent-terms-to-match-do/7789/3 "2012-05-24T08:31:55Z")

</div>

Ah that makes sense, thanks Shay.

So increasing `percent_terms_to_match` to 0.5 means that if I have 20 tags, 50% (10) of these should match in the other document for it to be returned. Increasing the value increases precision, while deceasing it decreases precision but increases recall.

Cheers.

On 23 May 2012, at 23:30, Shay Banon wrote:

> Effectively, what happens in the more like this query is that it builds a big boolean query with should clauses for each term. The percent terms to match means that out of all the terms built, at least X percent should match (effectively, setting the minimum\_should\_match parameter on it).
> 
> On Mon, May 21, 2012 at 6:30 PM, Nick Dunn [nick@nick-dunn.co.uk](mailto:nick@nick-dunn.co.uk) wrote:  
> I'm trying out morelikethis ([Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html)) and it's working well. So easy 🙂
> 
> I'm finding entries by related tags so I dropped `min_doc_freq` to 1 (one or more tag required for a match) and `max_query_terms` to 100 (an entry could be tagged with up to 100 tags) however from the docs it's not clear to me what `percent_terms_to_match` does:
> 
> The percentage of terms to match on (float value). Defaults to 0.3 (30 percent).
> 
> Could someone explain this in other words please, perhaps an example of what might happen if I increase or decrease from the default? When I try it on my sample data it increases/reduces the number of hits and doesn't seem to affect the score of each hit, so I'm just not sure what it's doing.
> 
> Thanks.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 25, 2012, 10:44pm UTC](https://discuss.elastic.co/t/morelikethis-query-what-does-percent-terms-to-match-do/7789/4 "2012-05-25T22:44:34Z")

</div>

Yep.

On Thu, May 24, 2012 at 10:31 AM, Nick Dunn [nick@nick-dunn.co.uk](mailto:nick@nick-dunn.co.uk) wrote:

> Ah that makes sense, thanks Shay.
> 
> So increasing `percent_terms_to_match` to 0.5 means that if I have 20  
> tags, 50% (10) of these should match in the other document for it to be  
> returned. Increasing the value increases precision, while deceasing it  
> decreases precision but increases recall.
> 
> Cheers.
> 
> On 23 May 2012, at 23:30, Shay Banon wrote:
> 
> Effectively, what happens in the more like this query is that it builds a  
> big boolean query with should clauses for each term. The percent terms to  
> match means that out of all the terms built, at least X percent should  
> match (effectively, setting the minimum\_should\_match parameter on it).
> 
> On Mon, May 21, 2012 at 6:30 PM, Nick Dunn [nick@nick-dunn.co.uk](mailto:nick@nick-dunn.co.uk) wrote:
> 
> > I'm trying out morelikethis (  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/mlt-query.html))  
> > and it's working well. So easy 🙂
> > 
> > I'm finding entries by related tags so I dropped `min_doc_freq` to 1 (one  
> > or more tag required for a match) and `max_query_terms` to 100 (an entry  
> > could be tagged with up to 100 tags) however from the docs it's not clear  
> > to me what `percent_terms_to_match` does:
> > 
> > The percentage of terms to match on (float value). Defaults to 0.3 (30  
> > percent).
> > 
> > Could someone explain this in other words please, perhaps an example of  
> > what might happen if I increase or decrease from the default? When I try it  
> > on my sample data it increases/reduces the number of hits and doesn't seem  
> > to affect the score of each hit, so I'm just not sure what it's doing.
> > 
> > Thanks.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:26am UTC](https://discuss.elastic.co/t/morelikethis-query-what-does-percent-terms-to-match-do/7789/5 "2017-07-06T03:26:53Z")

</div>


