# Ignore term frequency (not releveant for the type of document I'm using)

**URL:** <https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946>\
**Category:** Elasticsearch\
**Created:** [July 25, 2013, 7:34pm UTC](https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946 "2013-07-25T19:34:54Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 25, 2013, 7:34pm UTC](https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946/1 "2013-07-25T19:34:54Z")

</div>

Hi,

How do I get ES to ignore the term frequency since it is not releveant for  
the type of document I'm using?

I'm using two ES types to handle two kind of data that need different  
analyzers.  
I'm trying to query the 2 types using a multi\_match but I would like to  
ignore the term frequency.  
I tried using "index\_options" : "docs" on my fields but I'm still getting  
different scores depending on the term frequency.

Mapping:  
curl -XPOST "localhost:9200/myindex" -d '  
{  
"settings":{  
"index":{  
"analysis":{  
"filter" : {  
"name\_nGram" : {  
"max\_gram" : 100,  
"min\_gram" : 2,  
"type" : "edge\_ngram"  
},  
"strip\_hydrid\_sign\_filter":{  
"pattern":"\u00D7",  
"replacement":"",  
"type": "pattern\_replace"  
}  
},  
"analyzer":{  
"name\_index" : {  
"filter" : [  
"lowercase","asciifolding","name\_nGram"  
],  
"tokenizer" : "keyword"  
},  
"full\_name\_index" : {  
"filter" : [  
"lowercase","asciifolding"  
],  
"tokenizer" : "keyword"  
},  
"scientificname\_index" : {  
"filter" : [  
"lowercase","asciifolding","strip\_hydrid\_sign\_filter","name\_nGram"  
],  
"tokenizer" : "keyword"  
},  
"name\_search" : {  
"filter" : [  
"lowercase","asciifolding"  
],  
"tokenizer" : "keyword"  
},  
"scientificname\_search" : {  
"filter" : [  
"lowercase","asciifolding","strip\_hydrid\_sign\_filter"  
],  
"tokenizer" : "keyword"  
}  
}  
}  
}  
},  
"mappings" : {  
"taxon" : {  
"properties" : {  
"taxonname" : {  
"type" : "multi\_field",  
"fields":{  
"taxonname":{  
"type" : "string",  
"index\_analyzer" : "full\_name\_index",  
"search\_analyzer" : "name\_search",  
"omit\_norms" : true,  
"index\_options" : "docs"  
},  
"ngrams":{  
"type" : "string",  
"index\_analyzer" : "scientificname\_index",  
"search\_analyzer" : "scientificname\_search",  
"omit\_norms" : true,  
"index\_options" : "docs"  
}  
}  
}  
}  
},  
"vernacular" : {  
"properties" : {  
"vernacularname" : {  
"type" : "multi\_field",  
"fields":{  
"vernacularname":{  
"type" : "string",  
"index\_analyzer" : "full\_name\_index",  
"search\_analyzer" : "name\_search",  
"omit\_norms" : true,  
"index\_options" : "docs"  
},  
"ngrams":{  
"type" : "string",  
"index\_analyzer" : "name\_index",  
"search\_analyzer" : "name\_search",  
"omit\_norms" : true,  
"index\_options" : "docs"  
}  
}  
}  
}  
}  
}  
}'

Data:  
curl -XPUT "localhost:9200/myindex/taxon/1" -d '{  
"taxonname":"Carex capitata"  
}'  
curl -XPUT "localhost:9200/myindex/taxon/2" -d '{  
"taxonname":"Carex heleonastes"  
}'  
curl -XPUT "localhost:9200/myindex/taxon/3" -d '{  
"taxonname":"Carex buckleyi"  
}'

curl -XPUT "localhost:9200/myindex/vernacular/1" -d '{  
"vernacularname":"carex de Richardson"  
}'  
curl -XPUT "localhost:9200/myindex/vernacular/2" -d '{  
"vernacularname":"carex du lac Tahoe"  
}'

Query:  
curl  
"localhost:9200/myindex/\_search?search\_type=dfs\_query\_then\_fetch&pretty=1" -d  
'{  
"query":{  
"bool":{  
"should":[  
{  
"multi\_match" : {  
"query" : "carex",  
"fields" : ["taxonname", "taxonname.ngrams"]  
}  
},  
{  
"multi\_match" : {  
"query" : "carex",  
"fields" : ["vernacularname", "vernacularname.ngrams"]  
}  
}  
]  
}  
}  
}'

This would give a better score for vernacularname than taxonname since they  
have different term frequency.

So, how can I ignore the term frequency so vernacularname and taxonname  
would have the same score or, there is a better way to achieve that?

Thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [July 26, 2013, 9:50am UTC](https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946/2 "2013-07-26T09:50:41Z")

</div>

Hi,

Inverse document frequencies also play a role when scoring with the default  
similarity. You can ignore the scoring of a query by wrapping it inside a  
constant score query[1]. Does it help? Another option would be to write a  
custom similarity extending the default one that would always return 1 for  
the idf.

[1]  
[http://www.elasticsearch.org/guide/reference/query-dsl/constant-score-query/](http://www.elasticsearch.org/guide/reference/query-dsl/constant-score-query/)

--  
Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 26, 2013, 12:08pm UTC](https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946/3 "2013-07-26T12:08:54Z")

</div>

Hi Adrien,  
Indeed, the "explain" returns idf(docFreq=1334, maxDocs=57595) for the type  
taxon and idf(docFreq=366, maxDocs=57595)for the type vernacular so I guess  
the Inverse document frequency is the main reason.  
I'm not sure I understand the purpose of the constant score query.  
Actually, I want to have the score to sort them by relevance (ngrams  
fields) but I don't need the idf since the document frequency is not  
relevant in this specific context.

I guess the custom similarity should be something like that  
: [GitHub - tlrx/elasticsearch-custom-similarity-provider: A custom SimilarityProvider example for Elasticsearch](https://github.com/tlrx/elasticsearch-custom-similarity-provider)

Thanks,

Christian

On Friday, July 26, 2013 5:50:41 AM UTC-4, Adrien Grand wrote:

> Hi,
> 
> Inverse document frequencies also play a role when scoring with the  
> default similarity. You can ignore the scoring of a query by wrapping it  
> inside a constant score query[1]. Does it help? Another option would be to  
> write a custom similarity extending the default one that would always  
> return 1 for the idf.
> 
> [1]  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/constant-score-query/)
> 
> --  
> Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [July 26, 2013, 4:46pm UTC](https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946/4 "2013-07-26T16:46:21Z")

</div>

Hi,

On Fri, Jul 26, 2013 at 2:08 PM, Christian Gendreau \<  
[christiangendreau@gmail.com](mailto:christiangendreau@gmail.com)\> wrote:

> I'm not sure I understand the purpose of the constant score query.  
> Actually, I want to have the score to sort them by relevance (ngrams  
> fields) but I don't need the idf since the document frequency is not  
> relevant in this specific context.

I wanted to mention that if you run a boolean query with two clauses which  
are term queries wrapped into constant score queries, the TF-IDF won't be  
involved in the scoring, the best documents will be those which have the  
higher number of matching clauses.

> I guess the custom similarity should be something like that :  
> [GitHub - tlrx/elasticsearch-custom-similarity-provider: A custom SimilarityProvider example for Elasticsearch](https://github.com/tlrx/elasticsearch-custom-similarity-provider)

Exactly, you can even override tf(float freq) to something like "return  
freq \> 0 ? 1 : 0;" if you don't want to take into account the term  
frequency either.

--  
Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 26, 2013, 6:54pm UTC](https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946/5 "2013-07-26T18:54:49Z")

</div>

Hi,

I tried that:  
curl  
"localhost:9200/myindex/\_search?search\_type=dfs\_query\_then\_fetch&pretty=1"  
-d '{  
"query":{  
"bool":{  
"should":[  
{  
"constant\_score" : {  
"query" : {  
"match":{  
"taxonname":{  
"query":"carex"  
}  
}  
},  
"boost" : 1  
}  
},  
{  
"constant\_score" : {  
"query" : {  
"match":{  
"taxonname.ngrams":{  
"query":"carex"  
}  
}  
},  
"boost" : 1  
}  
}  
,{  
"constant\_score" : {  
"query" : {  
"match":{  
"vernacularname":{  
"query":"carex"  
}  
}  
},  
"boost" : 1  
}  
},  
{  
"constant\_score" : {  
"query" : {  
"match":{  
"vernacularname.ngrams":{  
"query":"carex"  
}  
}  
},  
"boost" : 1  
}  
}  
]  
}  
},  
"size" : 100,  
"sort" : [  
"\_score",  
{ "sortname" : {"order" : "asc"} }  
]  
}'

I'm not sure if this is exactly what you meant but it seems to work!  
I'm also not sure if this is the most efficient way to do this or the  
custom similarity would perform better.

Thanks for your help,

Christian

On Friday, July 26, 2013 12:46:21 PM UTC-4, Adrien Grand wrote:

> Hi,
> 
> On Fri, Jul 26, 2013 at 2:08 PM, Christian Gendreau \<[christia...@gmail.com](mailto:christia...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > I'm not sure I understand the purpose of the constant score query.  
> > Actually, I want to have the score to sort them by relevance (ngrams  
> > fields) but I don't need the idf since the document frequency is not  
> > relevant in this specific context.
> 
> I wanted to mention that if you run a boolean query with two clauses which  
> are term queries wrapped into constant score queries, the TF-IDF won't be  
> involved in the scoring, the best documents will be those which have the  
> higher number of matching clauses.
> 
> > I guess the custom similarity should be something like that :  
> > [GitHub - tlrx/elasticsearch-custom-similarity-provider: A custom SimilarityProvider example for Elasticsearch](https://github.com/tlrx/elasticsearch-custom-similarity-provider)
> 
> Exactly, you can even override tf(float freq) to something like "return  
> freq \> 0 ? 1 : 0;" if you don't want to take into account the term  
> frequency either.
> 
> --  
> Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:24am UTC](https://discuss.elastic.co/t/ignore-term-frequency-not-releveant-for-the-type-of-document-im-using/12946/6 "2017-07-06T02:24:14Z")

</div>


