# Score based on Term Frequency alone

**URL:** <https://discuss.elastic.co/t/score-based-on-term-frequency-alone/83323>\
**Category:** Elasticsearch\
**Created:** [April 23, 2017, 12:51pm UTC](https://discuss.elastic.co/t/score-based-on-term-frequency-alone/83323 "2017-04-23T12:51:14Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![hey-arnold](https://avatars.discourse-cdn.com/v4/letter/h/a4c791/32.png) [@hey-arnold](https://discuss.elastic.co/u/hey-arnold)\
**Post date:** [April 23, 2017, 12:51pm UTC](https://discuss.elastic.co/t/score-based-on-term-frequency-alone/83323/1 "2017-04-23T12:51:14Z")

</div>

i want to disable IDF, and maybe TF so that I just have a score based on how many terms are present from the given query. I've looked into solutions writing custom scripts, and playing around with the query, but these all involve splitting up the query itself into individual terms. The problem with this approach is that if you have some wrapper service which takes a query with multiple tokens in a string, you need to find a way to split the query into tokens before feeding them into the ES search request. The only way to do this reliably is to first make a request to the analyze endpoint, but this just slows things down.  
(Examples: [How to complete disable TF-IDF?](https://discuss.elastic.co/t/how-to-complete-disable-tf-idf/70570/4), [https://www.elastic.co/guide/en/elasticsearch/guide/current/ignoring-tfidf.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/ignoring-tfidf.html))

I think ES is awesome, but I think it would be cool if there were more similarity modules which cover simple use cases, like when you want a simple count on term presence in a document. Why not just have a bunch load of similarity modules: summing one hot vectors, cosine, etc...

Any advice on how I can achieve my goal? I am using elastic-search 5.2.2.

Thanks!

EDIT: I got this working by writing a plugin. There are examples online but they are outdated, I will formalise my solution, post to GIT, and update this answer in due time.

---

<div class="post-metadata">

**Author:** ![hey-arnold](https://avatars.discourse-cdn.com/v4/letter/h/a4c791/32.png) [@hey-arnold](https://discuss.elastic.co/u/hey-arnold)\
**Post date:** [April 25, 2017, 11:04pm UTC](https://discuss.elastic.co/t/score-based-on-term-frequency-alone/83323/2 "2017-04-25T23:04:41Z")

</div>

Ok, If anyone wants to disable IDF, disable TF, as to just score based on the presence of a term and boost value on the field, in elastic search v5+, then see the following plugin:

> **[farisk/elasticsearch-boolean-similarity-plugin](https://github.com/farisk/elasticsearch-boolean-similarity-plugin)**
>
> elasticsearch-boolean-similarity-plugin - Boolean similarity from Lucene made available in elasticsearch v5.x via a plugin

If you don't want to disable TF but don't know how to make a plugin, the code in the repo above should help, adding TF should be simple.

Also note some guy has implemented this into the latest elasticsearch code, see:

> <https://github.com/elastic/elasticsearch/commit/8359dd05c9aa75588c679a8ae151951fa00f001b>

  
You will just need to set "similarity": "boolean" on properties. This is available in elasticsearch 5.4.0 +

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 23, 2017, 11:05pm UTC](https://discuss.elastic.co/t/score-based-on-term-frequency-alone/83323/3 "2017-05-23T23:05:27Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
