# Finding similarity between docs to avoid data duplication by abusers

**URL:** <https://discuss.elastic.co/t/finding-similarity-between-docs-to-avoid-data-duplication-by-abusers/181434>\
**Category:** Elasticsearch\
**Created:** [May 16, 2019, 3:00pm UTC](https://discuss.elastic.co/t/finding-similarity-between-docs-to-avoid-data-duplication-by-abusers/181434 "2019-05-16T15:00:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![es\_noobie](https://avatars.discourse-cdn.com/v4/letter/e/edb3f5/32.png) [@es\_noobie](https://discuss.elastic.co/u/es_noobie)\
**Post date:** [May 16, 2019, 3:00pm UTC](https://discuss.elastic.co/t/finding-similarity-between-docs-to-avoid-data-duplication-by-abusers/181434/1 "2019-05-16T15:00:18Z")

</div>

Hi all, I have a simple (to explain) task and I'm researching into Elasticsearch to see if is the right tool.

**The task**

In a marketplace platform some (ab)users insert the (text) content they are advertising multiples times a day to get more visibility. They do modifications, like inserting small random strings in the middle of the text to avoid exact matching.

What I want to do is given a document find the max "similarity score" (as a percentage) of that document vs all other documents stored in the datastore, so that I can place for human review those documents with a score above a threshold.

**The question**

Is there a way I can obtain that max score (as a percentage or a coefficient between 0 and 1) using a Elasticsearch query?

I have already look into using a _More Like This_ query, but it seems like it's not precise enough for this task. Also read about Fuzzy matching, but with a edit distance limited to 2 it can't do the job for me.

Thanks in advance

---

<div class="post-metadata">

**Author:** ![balazs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/balazs/32/45225_2.png) [@balazs](https://discuss.elastic.co/u/balazs)\
**Post date:** [May 24, 2019, 9:47am UTC](https://discuss.elastic.co/t/finding-similarity-between-docs-to-avoid-data-duplication-by-abusers/181434/2 "2019-05-24T09:47:33Z")

</div>

Hi Joao,

In addition to the possibilities you've already mentioned, another option to consider for finding similar text is using [text embeddings](https://github.com/jtibshirani/text-embeddings) and calculating similarity with [vector fields](https://www.elastic.co/guide/en/elasticsearch/reference/7.1/dense-vector.html) using [cosine similarity](https://www.elastic.co/guide/en/elasticsearch/reference/7.2/query-dsl-script-score-query.html#vector-functions) script scoring (not yet released) to find the similarity between the vector generated for the new content string queried against the corpus. Given "they are advertising multiples times a day", I imagine the query could also be heavily restricted to just the current day's documents.

You may also want to consider analyzing the text in a way that discards emojis, special characters, or any other such elements that submitters are using to avoid exact matching. That way matching `Awesome Ugly ❤️Sweater❤️ $$$` to `Awesome Ugly Sweater` becomes much easier. In fact, if there is a strong match when using this method _and_ the raw text _also_ considers such characters, it may be an even stronger indication of abusive behavior. This is a challenging problem though, because if multiple people are legitimately trying to advertise something like concert tickets, they may all use almost the same terms to describe the tickets (e.g. `Tickets- <musician name> <venue> <date>`) and may deliberately add emojis or other special characters to callout attention to their advertisement, and this wouldn't be abusive behavior.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 21, 2019, 9:47am UTC](https://discuss.elastic.co/t/finding-similarity-between-docs-to-avoid-data-duplication-by-abusers/181434/3 "2019-06-21T09:47:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
