# Near duplicate document detection

**URL:** <https://discuss.elastic.co/t/near-duplicate-document-detection/241334>\
**Category:** Elasticsearch\
**Created:** [July 15, 2020, 4:36pm UTC](https://discuss.elastic.co/t/near-duplicate-document-detection/241334 "2020-07-15T16:36:30Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![tallison](https://avatars.discourse-cdn.com/v4/letter/t/258eb7/32.png) [@tallison](https://discuss.elastic.co/u/tallison)\
**Post date:** [July 15, 2020, 4:36pm UTC](https://discuss.elastic.co/t/near-duplicate-document-detection/241334/1 "2020-07-15T16:36:30Z")

</div>

I have an index w ~12 million files. We have a serious duplication issue and near duplication issue.

The exact duplicate problem is trivial because we store a digest. The near duplicate issue is more profound, and I've read the several posts on [discuss.elastic.co](http://discuss.elastic.co) that deal with near duplicates.

I reindexed with a min\_hash filter with default 512 buckets and other defaults, and I stored term\_vectors for that field on the theory that I'd improve search speed. Reindexing time and extra required space were both impressive...no problems there.

When I run a MoreLikeThisQuery, the performance is not great (200ms up to 45 seconds per query), even if I limit the terms to the top 5. I saw similar performance on MLT with the straight content field (with no termvectors).

I got much better performance when I programmatically created my own Boolean AND with a subset of the terms stored in the termvector in the min hash field, but then I had to run jaccard on the matches, which can be expensive given the number of near duplicates we have.

Are there other, more efficient strategies you'd recommend for finding near duplicates?

My next thought is to randomly select strings of five words and submit a couple of phrase queries...

Thank you!

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [July 15, 2020, 5:35pm UTC](https://discuss.elastic.co/t/near-duplicate-document-detection/241334/2 "2020-07-15T17:35:03Z")

</div>

Hi Tim,  
I was quite pleased with the encoding I used in the `significant_text` aggregation's near-duplicate filter.  
It's used in-memory on result streams but I guess you could apply the same approach to indexing content signatures. I discuss it [here](https://youtu.be/zH7bizwjj20?t=509).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 12, 2020, 5:35pm UTC](https://discuss.elastic.co/t/near-duplicate-document-detection/241334/3 "2020-08-12T17:35:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
