# Modifying scoring algorithm during search operations

**URL:** <https://discuss.elastic.co/t/modifying-scoring-algorithm-during-search-operations/15428>\
**Category:** Elasticsearch\
**Created:** [January 27, 2014, 7:12am UTC](https://discuss.elastic.co/t/modifying-scoring-algorithm-during-search-operations/15428 "2014-01-27T07:12:57Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Hiro\_Gangwani](https://avatars.discourse-cdn.com/v4/letter/h/9fc29f/32.png) [@Hiro\_Gangwani](https://discuss.elastic.co/u/Hiro_Gangwani)\
**Post date:** [January 27, 2014, 7:12am UTC](https://discuss.elastic.co/t/modifying-scoring-algorithm-during-search-operations/15428/1 "2014-01-27T07:12:57Z")

</div>

Dear Team,

I have been looking at search algorithm being used in elastic search and  
found following set of rules which are applied while calculating the score  
(Boolean Model)

- more occurrences in the document are preferred
- terms rarer in the corpus are preferred
- shorter documents are more heavily weighted
- other functions used to adjust score, boosts, etc.

In my application we are doing text based search across set of word  
documents. We would like to assign the higher scroe to documents having  
more occurances and show at the top irrespective of size of document.  
Primarily our application is recruitment system where is search is based  
upon skill sets. So our business team wants to show the resumes having more  
occurrences of search key words at top irrespective of size and rare terms.  
Is there any mechanism to ignore second and third rules as listed below and  
calculate the score based upon More occurrences condition only. We are  
executing search operations using Java API. Please let me know is it  
possible to achieve the same and if yes how?

Thanks in advance for suggesting solution.

Hiro

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f6936b6f-ef7c-4497-b186-bdba28176d89%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f6936b6f-ef7c-4497-b186-bdba28176d89%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 27, 2014, 6:20pm UTC](https://discuss.elastic.co/t/modifying-scoring-algorithm-during-search-operations/15428/2 "2014-01-27T18:20:41Z")

</div>

For the third rule, you can omit index norms for a field which will prevent  
length normalization. See [1]. The option is either called omit\_norms  
or norms.enabled depending on your version.

For the second rule, it is slightly more complicated. You can define your  
own custom similarity [2] that dictates how the TF, IDF and norms are used.  
You simply extends Lucene's DefaultSimilarity (of TDIDFSimilarity) and at  
it to elasticsearch's classpath.

[1]

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

[2]

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

--  
Ivan

On Sun, Jan 26, 2014 at 11:12 PM, Hiro Gangwani [hiro.gangwani@gmail.com](mailto:hiro.gangwani@gmail.com)wrote:

> Dear Team,
> 
> I have been looking at search algorithm being used in Elasticsearch and  
> found following set of rules which are applied while calculating the score  
> (Boolean Model)
> 
> - more occurrences in the document are preferred
> - terms rarer in the corpus are preferred
> - shorter documents are more heavily weighted
> - other functions used to adjust score, boosts, etc.
> 
> In my application we are doing text based search across set of word  
> documents. We would like to assign the higher scroe to documents having  
> more occurances and show at the top irrespective of size of document.  
> Primarily our application is recruitment system where is search is based  
> upon skill sets. So our business team wants to show the resumes having more  
> occurrences of search key words at top irrespective of size and rare terms.  
> Is there any mechanism to ignore second and third rules as listed below  
> and calculate the score based upon More occurrences condition only. We are  
> executing search operations using Java API. Please let me know is it  
> possible to achieve the same and if yes how?
> 
> Thanks in advance for suggesting solution.
> 
> Hiro
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/f6936b6f-ef7c-4497-b186-bdba28176d89%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f6936b6f-ef7c-4497-b186-bdba28176d89%40googlegroups.com)  
> .  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQA1d7L6ixwNPMtVZ%2BcdsYv8HfAc4CC4gQY%3D%2BavfT-rxEA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQA1d7L6ixwNPMtVZ%2BcdsYv8HfAc4CC4gQY%3D%2BavfT-rxEA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Hiro\_Gangwani](https://avatars.discourse-cdn.com/v4/letter/h/9fc29f/32.png) [@Hiro\_Gangwani](https://discuss.elastic.co/u/Hiro_Gangwani)\
**Post date:** [January 28, 2014, 6:32am UTC](https://discuss.elastic.co/t/modifying-scoring-algorithm-during-search-operations/15428/3 "2014-01-28T06:32:05Z")

</div>

Hi Ivan,  
Thanks for the reply. We tried using norms.enabled property and it is  
working fine. But what we have observed is this attribute works only on  
string types. In our application we are indexing the word (.doc,.docx) and  
pdf documents and performing test based search from document content. When  
we define the norm.enabled for attachments types, normalization is not  
working and size of document is being considered while calculating the  
score.

Please suggest how do resolve this issue for attachment types.

## Code to create the index for attachment types

## XContentBuilder map = XContentFactory.jsonBuilder().startObject() .startObject(idxType) .startObject("properties") .startObject("file") .field("type", "attachement") .field("norms.enabled", false) .startObject("fields") .startObject("refid") .field("store", "yes") .endObject() .startObject("name") .field("store", "yes") .endObject() .startObject("itexp") .field("store", "yes") .endObject() .startObject("totalexp") .field("store", "yes") .endObject() .endObject() .endObject() .endObject() .endObject();

Hiro

On Monday, 27 January 2014 23:50:41 UTC+5:30, Ivan Brusic wrote:

> For the third rule, you can omit index norms for a field which will  
> prevent length normalization. See [1]. The option is either  
> called omit\_norms or norms.enabled depending on your version.
> 
> For the second rule, it is slightly more complicated. You can define your  
> own custom similarity [2] that dictates how the TF, IDF and norms are used.  
> You simply extends Lucene's DefaultSimilarity (of TDIDFSimilarity) and at  
> it to elasticsearch's classpath.
> 
> [1]  
> [Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/mapping-core-types.html#string)  
> [2]  
> [Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/index-modules-similarity.html)
> 
> --  
> Ivan
> 
> On Sun, Jan 26, 2014 at 11:12 PM, Hiro Gangwani \<[hiro.g...@gmail.com](mailto:hiro.g...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > Dear Team,
> > 
> > I have been looking at search algorithm being used in Elasticsearch and  
> > found following set of rules which are applied while calculating the score  
> > (Boolean Model)
> > 
> > - more occurrences in the document are preferred
> > - terms rarer in the corpus are preferred
> > - shorter documents are more heavily weighted
> > - other functions used to adjust score, boosts, etc.
> > 
> > In my application we are doing text based search across set of word  
> > documents. We would like to assign the higher scroe to documents having  
> > more occurances and show at the top irrespective of size of document.  
> > Primarily our application is recruitment system where is search is based  
> > upon skill sets. So our business team wants to show the resumes having more  
> > occurrences of search key words at top irrespective of size and rare terms.  
> > Is there any mechanism to ignore second and third rules as listed below  
> > and calculate the score based upon More occurrences condition only. We are  
> > executing search operations using Java API. Please let me know is it  
> > possible to achieve the same and if yes how?
> > 
> > Thanks in advance for suggesting solution.
> > 
> > Hiro
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/f6936b6f-ef7c-4497-b186-bdba28176d89%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f6936b6f-ef7c-4497-b186-bdba28176d89%40googlegroups.com)  
> > .  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f80933eb-1b68-4c6f-b073-39b78e3f45e9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f80933eb-1b68-4c6f-b073-39b78e3f45e9%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 28, 2014, 4:03pm UTC](https://discuss.elastic.co/t/modifying-scoring-algorithm-during-search-operations/15428/4 "2014-01-28T16:03:58Z")

</div>

Norms are applied at the field level, not at the index level. You would  
need to omit norms for every field it is meant to apply to. Another  
alternative would be to use index templates:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

--  
Ivan

On Mon, Jan 27, 2014 at 10:32 PM, Hiro Gangwani [hiro.gangwani@gmail.com](mailto:hiro.gangwani@gmail.com)wrote:

> Hi Ivan,  
> Thanks for the reply. We tried using norms.enabled property and it is  
> working fine. But what we have observed is this attribute works only on  
> string types. In our application we are indexing the word (.doc,.docx) and  
> pdf documents and performing test based search from document content. When  
> we define the norm.enabled for attachments types, normalization is not  
> working and size of document is being considered while calculating the  
> score.
> 
> Please suggest how do resolve this issue for attachment types.
> 
> ## Code to create the index for attachment types
> 
> ## XContentBuilder map = XContentFactory.jsonBuilder().startObject() .startObject(idxType) .startObject("properties") .startObject("file") .field("type", "attachement") .field("norms.enabled", false) .startObject("fields") .startObject("refid") .field("store", "yes") .endObject() .startObject("name") .field("store", "yes") .endObject() .startObject("itexp") .field("store", "yes") .endObject() .startObject("totalexp") .field("store", "yes") .endObject() .endObject() .endObject() .endObject() .endObject();
> 
> Hiro
> 
> On Monday, 27 January 2014 23:50:41 UTC+5:30, Ivan Brusic wrote:
> 
> > For the third rule, you can omit index norms for a field which will  
> > prevent length normalization. See [1]. The option is either  
> > called omit\_norms or norms.enabled depending on your version.
> > 
> > For the second rule, it is slightly more complicated. You can define your  
> > own custom similarity [2] that dictates how the TF, IDF and norms are used.  
> > You simply extends Lucene's DefaultSimilarity (of TDIDFSimilarity) and at  
> > it to elasticsearch's classpath.
> > 
> > [1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/)  
> > reference/current/mapping-core-types.html#string  
> > [2] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/)  
> > reference/current/index-modules-similarity.html
> > 
> > --  
> > Ivan
> > 
> > On Sun, Jan 26, 2014 at 11:12 PM, Hiro Gangwani [hiro.g...@gmail.com](mailto:hiro.g...@gmail.com)wrote:
> > 
> > > Dear Team,
> > > 
> > > I have been looking at search algorithm being used in Elasticsearch and  
> > > found following set of rules which are applied while calculating the score  
> > > (Boolean Model)
> > > 
> > > - more occurrences in the document are preferred
> > > - terms rarer in the corpus are preferred
> > > - shorter documents are more heavily weighted
> > > - other functions used to adjust score, boosts, etc.
> > > 
> > > In my application we are doing text based search across set of word  
> > > documents. We would like to assign the higher scroe to documents having  
> > > more occurances and show at the top irrespective of size of document.  
> > > Primarily our application is recruitment system where is search is based  
> > > upon skill sets. So our business team wants to show the resumes having more  
> > > occurrences of search key words at top irrespective of size and rare terms.  
> > > Is there any mechanism to ignore second and third rules as listed below  
> > > and calculate the score based upon More occurrences condition only. We are  
> > > executing search operations using Java API. Please let me know is it  
> > > possible to achieve the same and if yes how?
> > > 
> > > Thanks in advance for suggesting solution.
> > > 
> > > Hiro
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).
> > > 
> > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > msgid/elasticsearch/f6936b6f-ef7c-4497-b186-bdba28176d89%  
> > > [40googlegroups.com](http://40googlegroups.com).  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/f80933eb-1b68-4c6f-b073-39b78e3f45e9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f80933eb-1b68-4c6f-b073-39b78e3f45e9%40googlegroups.com)  
> > .
> 
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCje\_CLTBjf%2B6E9x0w\_tn61GP9TvaacgC4d%2B38EPodAtA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCje_CLTBjf%2B6E9x0w_tn61GP9TvaacgC4d%2B38EPodAtA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:54am UTC](https://discuss.elastic.co/t/modifying-scoring-algorithm-during-search-operations/15428/5 "2017-07-06T01:54:21Z")

</div>


