# How to change similarity without actual code

**URL:** <https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824>\
**Category:** Elasticsearch\
**Created:** [February 20, 2013, 4:31pm UTC](https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824 "2013-02-20T16:31:54Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [February 20, 2013, 4:31pm UTC](https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824/1 "2013-02-20T16:31:54Z")

</div>

Hey,

I am fresh to ES, and i have a task that i dont know what approach is best  
to take.

our data is a simple line of text and some number fields, and our queries  
are only on the line of text.  
when I query a few terms, (as far as i understand) the score gets  
calculated in such way, that prefers multiple occurrences of terms in the  
text, and also prefers longer matches.

if i would want to change that (say, dont mind how many times a term  
appeared, and dont mind the length), i would write this in lucene:

public class MySimilarity extends DefaultSimilarity {

```
@Override

//We don't care about how many times a term appears in the text

public float tf(float freq) {

    return freq == 0 ? 0 : 1;

}   

@Override

 public float computeNorm(String field, FieldInvertState state) {

    return state.getBoost(); //ignore length factor

}  

```

}

now my question is - is there a way to do this kind of things in ES, so i  
dont have to actually write code, ie use the dsl?

Thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [February 20, 2013, 10:41pm UTC](https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824/2 "2013-02-20T22:41:35Z")

</div>

Hey,

for the Term Frequency part I would just recommend to omit the TF and index  
document ids only. this will effectively what you showed in your example  
just without the "branch" in the similarity. (see 'index\_options' here:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html))  
For the length normalization I'd omitNorms ('omit\_norms'  
here: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html))  
in the mapping and use a custom score like shown  
here: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/custom-score-query.html)

this should be equivalent to what you want and you can influence how much  
weight the boost gets at runtime.

simon

On Wednesday, February 20, 2013 5:31:54 PM UTC+1, Shlomi wrote:

> Hey,
> 
> I am fresh to ES, and i have a task that i dont know what approach is best  
> to take.
> 
> our data is a simple line of text and some number fields, and our queries  
> are only on the line of text.  
> when I query a few terms, (as far as i understand) the score gets  
> calculated in such way, that prefers multiple occurrences of terms in the  
> text, and also prefers longer matches.
> 
> if i would want to change that (say, dont mind how many times a term  
> appeared, and dont mind the length), i would write this in lucene:
> 
> public class MySimilarity extends DefaultSimilarity {
> 
> ```
> @Override
> 
> //We don't care about how many times a term appears in the text
> 
> public float tf(float freq) {
> 
> return freq == 0 ? 0 : 1;
> 
> }   
> 
> @Override
> 
> public float computeNorm(String field, FieldInvertState state) {
> 
> return state.getBoost(); //ignore length factor
> 
> }  
> 
> ```
> 
> }
> 
> now my question is - is there a way to do this kind of things in ES, so i  
> dont have to actually write code, ie use the dsl?
> 
> Thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [February 21, 2013, 10:02am UTC](https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824/3 "2013-02-21T10:02:15Z")

</div>

Hey

Thank you for your response,

"omit\_norms" seemed to do the job right, but "index\_options" set to "docs"  
made searches that are not direct term unavailable, meaning i couldnt do  
query\_string like: +bre\* -break\* AND "Tons of"

I tried my luck with unique token filters[http://www.elasticsearch.org/guide/reference/index-modules/analysis/unique-tokenfilter.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/unique-tokenfilter.html),  
from what i understood (from the really short description), it should give  
me similar results, am I correct?

On Thursday, February 21, 2013 12:41:35 AM UTC+2, simonw wrote:

> Hey,
> 
> for the Term Frequency part I would just recommend to omit the TF and  
> index document ids only. this will effectively what you showed in your  
> example just without the "branch" in the similarity. (see 'index\_options'  
> here: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)  
> )  
> For the length normalization I'd omitNorms ('omit\_norms' here:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)) in  
> the mapping and use a custom score like shown here:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/custom-score-query.html)
> 
> this should be equivalent to what you want and you can influence how much  
> weight the boost gets at runtime.
> 
> simon
> 
> On Wednesday, February 20, 2013 5:31:54 PM UTC+1, Shlomi wrote:
> 
> > Hey,
> > 
> > I am fresh to ES, and i have a task that i dont know what approach is  
> > best to take.
> > 
> > our data is a simple line of text and some number fields, and our queries  
> > are only on the line of text.  
> > when I query a few terms, (as far as i understand) the score gets  
> > calculated in such way, that prefers multiple occurrences of terms in the  
> > text, and also prefers longer matches.
> > 
> > if i would want to change that (say, dont mind how many times a term  
> > appeared, and dont mind the length), i would write this in lucene:
> > 
> > public class MySimilarity extends DefaultSimilarity {
> > 
> > ```
> > @Override
> > 
> > //We don't care about how many times a term appears in the text
> > 
> > public float tf(float freq) {
> > 
> > return freq == 0 ? 0 : 1;
> > 
> > }   
> > 
> > @Override
> > 
> > public float computeNorm(String field, FieldInvertState state) {
> > 
> > return state.getBoost(); //ignore length factor
> > 
> > }  
> > 
> > ```
> > 
> > }
> > 
> > now my question is - is there a way to do this kind of things in ES, so i  
> > dont have to actually write code, ie use the dsl?
> > 
> > Thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [February 22, 2013, 10:39pm UTC](https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824/4 "2013-02-22T22:39:47Z")

</div>

On Thursday, February 21, 2013 11:02:15 AM UTC+1, Shlomi wrote:

> Hey
> 
> Thank you for your response,
> 
> "omit\_norms" seemed to do the job right, but "index\_options" set to "docs"  
> made searches that are not direct term unavailable, meaning i couldnt do  
> query\_string like: +bre\* -break\* AND "Tons of"

ah I see yeah setting this to "docs" will drop positions and queries like  
"Tons of" won't work anymore. UniqueTokenFitler should do the job here!

simon

> I tried my luck with unique token filters[http://www.elasticsearch.org/guide/reference/index-modules/analysis/unique-tokenfilter.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/unique-tokenfilter.html),  
> from what i understood (from the really short description), it should give  
> me similar results, am I correct?
> 
> On Thursday, February 21, 2013 12:41:35 AM UTC+2, simonw wrote:
> 
> > Hey,
> > 
> > for the Term Frequency part I would just recommend to omit the TF and  
> > index document ids only. this will effectively what you showed in your  
> > example just without the "branch" in the similarity. (see 'index\_options'  
> > here:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html))  
> > For the length normalization I'd omitNorms ('omit\_norms' here:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)) in  
> > the mapping and use a custom score like shown here:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/custom-score-query.html)
> > 
> > this should be equivalent to what you want and you can influence how much  
> > weight the boost gets at runtime.
> > 
> > simon
> > 
> > On Wednesday, February 20, 2013 5:31:54 PM UTC+1, Shlomi wrote:
> > 
> > > Hey,
> > > 
> > > I am fresh to ES, and i have a task that i dont know what approach is  
> > > best to take.
> > > 
> > > our data is a simple line of text and some number fields, and our  
> > > queries are only on the line of text.  
> > > when I query a few terms, (as far as i understand) the score gets  
> > > calculated in such way, that prefers multiple occurrences of terms in the  
> > > text, and also prefers longer matches.
> > > 
> > > if i would want to change that (say, dont mind how many times a term  
> > > appeared, and dont mind the length), i would write this in lucene:
> > > 
> > > public class MySimilarity extends DefaultSimilarity {
> > > 
> > > ```
> > > @Override
> > > 
> > > //We don't care about how many times a term appears in the text
> > > 
> > > public float tf(float freq) {
> > > 
> > > return freq == 0 ? 0 : 1;
> > > 
> > > }   
> > > 
> > > @Override
> > > 
> > > public float computeNorm(String field, FieldInvertState state) {
> > > 
> > > return state.getBoost(); //ignore length factor
> > > 
> > > }  
> > > 
> > > ```
> > > 
> > > }
> > > 
> > > now my question is - is there a way to do this kind of things in ES, so  
> > > i dont have to actually write code, ie use the dsl?
> > > 
> > > Thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [February 24, 2013, 3:34pm UTC](https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824/5 "2013-02-24T15:34:34Z")

</div>

thanks!

So I tired that, and it worked fine until i tried to query something like  
"bye bye", which was not distinguishable from "bye" (as opposed with "bye  
now" for instance)..

of course i could do shingle token filter, but that would needlessly  
enlarge my index size..

any other suggestions?

On Saturday, February 23, 2013 12:39:47 AM UTC+2, simonw wrote:

> On Thursday, February 21, 2013 11:02:15 AM UTC+1, Shlomi wrote:
> 
> > Hey
> > 
> > Thank you for your response,
> > 
> > "omit\_norms" seemed to do the job right, but "index\_options" set to  
> > "docs" made searches that are not direct term unavailable, meaning i  
> > couldnt do query\_string like: +bre\* -break\* AND "Tons of"
> 
> ah I see yeah setting this to "docs" will drop positions and queries like  
> "Tons of" won't work anymore. UniqueTokenFitler should do the job here!
> 
> simon
> 
> > I tried my luck with unique token filters[http://www.elasticsearch.org/guide/reference/index-modules/analysis/unique-tokenfilter.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/unique-tokenfilter.html),  
> > from what i understood (from the really short description), it should give  
> > me similar results, am I correct?
> > 
> > On Thursday, February 21, 2013 12:41:35 AM UTC+2, simonw wrote:
> > 
> > > Hey,
> > > 
> > > for the Term Frequency part I would just recommend to omit the TF and  
> > > index document ids only. this will effectively what you showed in your  
> > > example just without the "branch" in the similarity. (see 'index\_options'  
> > > here:  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html))  
> > > For the length normalization I'd omitNorms ('omit\_norms' here:  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html))  
> > > in the mapping and use a custom score like shown here:  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/custom-score-query.html)
> > > 
> > > this should be equivalent to what you want and you can influence how  
> > > much weight the boost gets at runtime.
> > > 
> > > simon
> > > 
> > > On Wednesday, February 20, 2013 5:31:54 PM UTC+1, Shlomi wrote:
> > > 
> > > > Hey,
> > > > 
> > > > I am fresh to ES, and i have a task that i dont know what approach is  
> > > > best to take.
> > > > 
> > > > our data is a simple line of text and some number fields, and our  
> > > > queries are only on the line of text.  
> > > > when I query a few terms, (as far as i understand) the score gets  
> > > > calculated in such way, that prefers multiple occurrences of terms in the  
> > > > text, and also prefers longer matches.
> > > > 
> > > > if i would want to change that (say, dont mind how many times a term  
> > > > appeared, and dont mind the length), i would write this in lucene:
> > > > 
> > > > public class MySimilarity extends DefaultSimilarity {
> > > > 
> > > > ```
> > > > @Override
> > > > 
> > > > //We don't care about how many times a term appears in the text
> > > > 
> > > > public float tf(float freq) {
> > > > 
> > > > return freq == 0 ? 0 : 1;
> > > > 
> > > > }   
> > > > 
> > > > @Override
> > > > 
> > > > public float computeNorm(String field, FieldInvertState state) {
> > > > 
> > > > return state.getBoost(); //ignore length factor
> > > > 
> > > > }  
> > > > 
> > > > ```
> > > > 
> > > > }
> > > > 
> > > > now my question is - is there a way to do this kind of things in ES, so  
> > > > i dont have to actually write code, ie use the dsl?
> > > > 
> > > > Thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:49am UTC](https://discuss.elastic.co/t/how-to-change-similarity-without-actual-code/10824/6 "2017-07-06T02:49:50Z")

</div>


