# Percolator in ES 5 - minimum similarity

**URL:** <https://discuss.elastic.co/t/percolator-in-es-5-minimum-similarity/53754>\
**Category:** Elasticsearch\
**Created:** [June 23, 2016, 9:07am UTC](https://discuss.elastic.co/t/percolator-in-es-5-minimum-similarity/53754 "2016-06-23T09:07:40Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![heino1986](https://avatars.discourse-cdn.com/v4/letter/h/c4cdca/32.png) [@heino1986](https://discuss.elastic.co/u/heino1986)\
**Post date:** [June 23, 2016, 9:07am UTC](https://discuss.elastic.co/t/percolator-in-es-5-minimum-similarity/53754/1 "2016-06-23T09:07:40Z")

</div>

Similar to this post here

> <https://stackoverflow.com/questions/24335715/scoring-in-elasticsearch-percolate-response>

I am wondering if the percolate query in ES 5 (or in 2.3) allows full-text style similarity queries. If possible I would like to percolate existing and new documents and only 'tag/classify' them if the similarity score is above a certain threshold value. Is this possible out of the box, can it be done with a script, or not at all.

I understand that for this to work I would need to have access to the tf-idf of my index of documents to be percolated, which will become computationally intensive as the size of said index grows. But if I'm not mistaken percolation works differently in ES5 compared to ES2.3?

Any thoughts are very welcome. Thanks,  
Chris

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [June 23, 2016, 10:18am UTC](https://discuss.elastic.co/t/percolator-in-es-5-minimum-similarity/53754/2 "2016-06-23T10:18:42Z")

</div>

In ES 2.x and before scoring based on the document being percolated isn't possible.

In ES 5 many things have changed in the percolator and the `percolate` query does produce a score based on the document being percolated. The document being percolated is indexed into an temporary in-memory index as part of query parsing. This index only contains a single document and at search time all percolator queries are executed on that index. So when it comes to scoring for each percolator query, only the tf statistics has an impact, the df statistic is always 1. There is no way you can get access to the df statistic of another ES index.

The thresholding you mention can be achieved using [min\_score](https://www.elastic.co/guide/en/elasticsearch/reference/master/search-request-min-score.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:41pm UTC](https://discuss.elastic.co/t/percolator-in-es-5-minimum-similarity/53754/3 "2017-07-05T22:41:11Z")

</div>


