# MoreLikeThis query performance with some extremely common words

**URL:** <https://discuss.elastic.co/t/morelikethis-query-performance-with-some-extremely-common-words/54775>\
**Category:** Elasticsearch\
**Created:** [July 5, 2016, 6:57pm UTC](https://discuss.elastic.co/t/morelikethis-query-performance-with-some-extremely-common-words/54775 "2016-07-05T18:57:40Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![J\_Campbell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/j_campbell/32/10732_2.png) [@J\_Campbell](https://discuss.elastic.co/u/J_Campbell)\
**Post date:** [July 5, 2016, 6:57pm UTC](https://discuss.elastic.co/t/morelikethis-query-performance-with-some-extremely-common-words/54775/1 "2016-07-05T18:57:40Z")

</div>

Hi, I'm using the latest stable ElasticSearch (2.3.3) and I am using the More Like This query to find similar documents. My documents are highly structured with lots of fields. The problem I am running into occurs when large groups, or even most of my documents contain fields with the same terms (values): the More Like This query becomes slower and slower as more documents are added. Initially, I tried using the max\_doc\_freq to solve this problem. While max\_doc\_freq did solve the performance problem, it creates a new problem in that sometimes the More Like This query returns 0 results even when there are many similar documents in the index once the max\_doc\_freq threshold is exceeded for too many fields. I'm also aware of the stop words option, but I don't have advance knowledge of what words will become common.

It is my understanding that the performance impact comes from retrieving all the documents with _any_ words in common (when max\_doc\_freq is not used) from the index and then computing scores for each one.

What I want is some way to run the More Like This query in such a way so that it maintains the performance while also returning similar documents. For example, if there were some way to only retrieve _some_ documents (like 10 or 100) with common words that exceed the max\_doc\_freq, rather than _every_ document, from the index, but still apply the full scoring function to every field regardless of document frequency. This way, if I inserted 1 million identical documents and set the max\_doc\_freq to 1000, I would still get _some_ results rather than 0 when running More Like This for one of the 1 million identical documents, and for documents with less common terms, More Like This would work as expected.

Thanks!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:37pm UTC](https://discuss.elastic.co/t/morelikethis-query-performance-with-some-extremely-common-words/54775/2 "2017-07-05T22:37:51Z")

</div>


