# Scoring is lower after adding new field to a document

**URL:** <https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283>\
**Category:** Elasticsearch\
**Created:** [July 2, 2012, 12:56pm UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283 "2012-07-02T12:56:09Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![philipDS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/philipds/32/2817_2.png) [@philipDS](https://discuss.elastic.co/u/philipDS)\
**Post date:** [July 2, 2012, 12:56pm UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283/1 "2012-07-02T12:56:09Z")

</div>

Hey,

We're building an application that allows people to save the Twitter handle  
of a person and an associated tweet. The user has to select the handle and  
then it fetches the tweet using the Twitter API. This information is stored  
in an ElasticSearch document afterwards. The Twitter field is initialized  
at {} (empty JSON object, we're using a Javascript ES wrapper). The  
following is (part of) the mapping:

user: { type: 'String', index: 'not\_analyzed' },  
location : { type: 'String', analyzer: 'lowercase\_keyword' },  
tags : { type: 'String', index: 'not\_analyzed' },  
twitter : {  
properties : {  
handle : { type: 'String', index: 'not\_analyzed' },  
tweet : { type: 'String', index: 'not\_analyzed' },  
age : { type: 'Date' }  
},  
type : 'nested'  
}

The location and tags fields are already filled in when we find the handle  
and tweet of the user. When we find a handle and a tweet, we set all the  
twitter fields (including the age, which is set at the current date), so it  
becomes something like:  
{  
handle: 'twitter\_handle',  
tweet: 'my last tweet',  
age: '....'  
}

Now, after this twitter info is added, the score of the document changes,  
when using a query that does not search these specific twitter fields. The  
following query is used to search for documents we need:

query {  
bool: {  
must: {  
[{  
text: {  
'tags': {  
query: 'Technology',  
operator: 'and'  
}  
}  
},  
{  
text: {  
'location': {  
query: 'Madrid',  
operator: 'and'  
}  
}  
},  
{  
text: {  
'user' : {  
query: queryArgs.email,  
operator: 'and'  
}  
}  
}]  
}  
}  
}

I have set the explanation to true and see that the score is lower after  
adding the twitter information (to the same document as the tags, location  
and user), but I don't understand why. Does this have to do with the term  
frequency factor _tf_? Or what's the exact explanation for this behaviour?  
Should I keep this twitter specific information in another type in the  
index, with a link (by ID) to it to not affect the score? How can I avoid  
the change in score in general while still adding fields to the document?

Thanks!

---

<div class="post-metadata">

**Author:** ![philipDS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/philipds/32/2817_2.png) [@philipDS](https://discuss.elastic.co/u/philipDS)\
**Post date:** [July 4, 2012, 10:01am UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283/2 "2012-07-04T10:01:05Z")

</div>

Anyone that could help me?

On Monday, 2 July 2012 14:56:09 UTC+2, philipDS wrote:

> Hey,
> 
> We're building an application that allows people to save the Twitter  
> handle of a person and an associated tweet. The user has to select the  
> handle and then it fetches the tweet using the Twitter API. This  
> information is stored in an Elasticsearch document afterwards. The Twitter  
> field is initialized at {} (empty JSON object, we're using a Javascript ES  
> wrapper). The following is (part of) the mapping:
> 
> user: { type: 'String', index: 'not\_analyzed' },  
> location : { type: 'String', analyzer: 'lowercase\_keyword' },  
> tags : { type: 'String', index: 'not\_analyzed' },  
> twitter : {  
> properties : {  
> handle : { type: 'String', index: 'not\_analyzed' },  
> tweet : { type: 'String', index: 'not\_analyzed' },  
> age : { type: 'Date' }  
> },  
> type : 'nested'  
> }
> 
> The location and tags fields are already filled in when we find the handle  
> and tweet of the user. When we find a handle and a tweet, we set all the  
> twitter fields (including the age, which is set at the current date), so it  
> becomes something like:  
> {  
> handle: 'twitter\_handle',  
> tweet: 'my last tweet',  
> age: '....'  
> }
> 
> Now, after this twitter info is added, the score of the document changes,  
> when using a query that does not search these specific twitter fields. The  
> following query is used to search for documents we need:
> 
> query {  
> bool: {  
> must: {  
> [{  
> text: {  
> 'tags': {  
> query: 'Technology',  
> operator: 'and'  
> }  
> }  
> },  
> {  
> text: {  
> 'location': {  
> query: 'Madrid',  
> operator: 'and'  
> }  
> }  
> },  
> {  
> text: {  
> 'user' : {  
> query: queryArgs.email,  
> operator: 'and'  
> }  
> }  
> }]  
> }  
> }  
> }
> 
> I have set the explanation to true and see that the score is lower after  
> adding the twitter information (to the same document as the tags, location  
> and user), but I don't understand why. Does this have to do with the term  
> frequency factor _tf_? Or what's the exact explanation for this  
> behaviour? Should I keep this twitter specific information in another type  
> in the index, with a link (by ID) to it to not affect the score? How can I  
> avoid the change in score in general while still adding fields to the  
> document?
> 
> Thanks!

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [July 4, 2012, 10:19am UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283/3 "2012-07-04T10:19:57Z")

</div>

Have you disabled the \_all field? It seems you are not interested in it, so  
you can just disable it. With the mapping shown, at least the "age" field  
will also get indexed into the \_all field and this might change the score.

Best regards,

Jörg

On Wednesday, July 4, 2012 12:01:05 PM UTC+2, philipDS wrote:

> Anyone that could help me?
> 
> On Monday, 2 July 2012 14:56:09 UTC+2, philipDS wrote:
> 
> > Hey,
> > 
> > We're building an application that allows people to save the Twitter  
> > handle of a person and an associated tweet. The user has to select the  
> > handle and then it fetches the tweet using the Twitter API. This  
> > information is stored in an Elasticsearch document afterwards. The Twitter  
> > field is initialized at {} (empty JSON object, we're using a Javascript ES  
> > wrapper). The following is (part of) the mapping:
> > 
> > user: { type: 'String', index: 'not\_analyzed' },  
> > location : { type: 'String', analyzer: 'lowercase\_keyword' },  
> > tags : { type: 'String', index: 'not\_analyzed' },  
> > twitter : {  
> > properties : {  
> > handle : { type: 'String', index: 'not\_analyzed' },  
> > tweet : { type: 'String', index: 'not\_analyzed' },  
> > age : { type: 'Date' }  
> > },  
> > type : 'nested'  
> > }
> > 
> > The location and tags fields are already filled in when we find the  
> > handle and tweet of the user. When we find a handle and a tweet, we set all  
> > the twitter fields (including the age, which is set at the current date),  
> > so it becomes something like:  
> > {  
> > handle: 'twitter\_handle',  
> > tweet: 'my last tweet',  
> > age: '....'  
> > }
> > 
> > Now, after this twitter info is added, the score of the document changes,  
> > when using a query that does not search these specific twitter fields. The  
> > following query is used to search for documents we need:
> > 
> > query {  
> > bool: {  
> > must: {  
> > [{  
> > text: {  
> > 'tags': {  
> > query: 'Technology',  
> > operator: 'and'  
> > }  
> > }  
> > },  
> > {  
> > text: {  
> > 'location': {  
> > query: 'Madrid',  
> > operator: 'and'  
> > }  
> > }  
> > },  
> > {  
> > text: {  
> > 'user' : {  
> > query: queryArgs.email,  
> > operator: 'and'  
> > }  
> > }  
> > }]  
> > }  
> > }  
> > }
> > 
> > I have set the explanation to true and see that the score is lower after  
> > adding the twitter information (to the same document as the tags, location  
> > and user), but I don't understand why. Does this have to do with the term  
> > frequency factor _tf_? Or what's the exact explanation for this  
> > behaviour? Should I keep this twitter specific information in another type  
> > in the index, with a link (by ID) to it to not affect the score? How can I  
> > avoid the change in score in general while still adding fields to the  
> > document?
> > 
> > Thanks!

---

<div class="post-metadata">

**Author:** ![philipDS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/philipds/32/2817_2.png) [@philipDS](https://discuss.elastic.co/u/philipDS)\
**Post date:** [July 4, 2012, 11:54am UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283/4 "2012-07-04T11:54:44Z")

</div>

I just removed all my indexes and disabled the all field:

this.liMapping = {  
contacts: {  
"\_all" : { "enabled" : false },  
properties: {  
// mapping here  
}  
}  
};

Still the same result though ☹ My query also doesn't search on the \_all  
field, but on specific fields. So, I don't know if this \_all field has  
anything to do with this issue?

On Wednesday, 4 July 2012 12:19:57 UTC+2, Jörg Prante wrote:

> Have you disabled the \_all field? It seems you are not interested in it,  
> so you can just disable it. With the mapping shown, at least the "age"  
> field will also get indexed into the \_all field and this might change the  
> score.
> 
> Best regards,
> 
> Jörg
> 
> On Wednesday, July 4, 2012 12:01:05 PM UTC+2, philipDS wrote:
> 
> > Anyone that could help me?
> > 
> > On Monday, 2 July 2012 14:56:09 UTC+2, philipDS wrote:
> > 
> > > Hey,
> > > 
> > > We're building an application that allows people to save the Twitter  
> > > handle of a person and an associated tweet. The user has to select the  
> > > handle and then it fetches the tweet using the Twitter API. This  
> > > information is stored in an Elasticsearch document afterwards. The Twitter  
> > > field is initialized at {} (empty JSON object, we're using a Javascript ES  
> > > wrapper). The following is (part of) the mapping:
> > > 
> > > user: { type: 'String', index: 'not\_analyzed' },  
> > > location : { type: 'String', analyzer: 'lowercase\_keyword' },  
> > > tags : { type: 'String', index: 'not\_analyzed' },  
> > > twitter : {  
> > > properties : {  
> > > handle : { type: 'String', index: 'not\_analyzed' },  
> > > tweet : { type: 'String', index: 'not\_analyzed' },  
> > > age : { type: 'Date' }  
> > > },  
> > > type : 'nested'  
> > > }
> > > 
> > > The location and tags fields are already filled in when we find the  
> > > handle and tweet of the user. When we find a handle and a tweet, we set all  
> > > the twitter fields (including the age, which is set at the current date),  
> > > so it becomes something like:  
> > > {  
> > > handle: 'twitter\_handle',  
> > > tweet: 'my last tweet',  
> > > age: '....'  
> > > }
> > > 
> > > Now, after this twitter info is added, the score of the document  
> > > changes, when using a query that does not search these specific twitter  
> > > fields. The following query is used to search for documents we need:
> > > 
> > > query {  
> > > bool: {  
> > > must: {  
> > > [{  
> > > text: {  
> > > 'tags': {  
> > > query: 'Technology',  
> > > operator: 'and'  
> > > }  
> > > }  
> > > },  
> > > {  
> > > text: {  
> > > 'location': {  
> > > query: 'Madrid',  
> > > operator: 'and'  
> > > }  
> > > }  
> > > },  
> > > {  
> > > text: {  
> > > 'user' : {  
> > > query: queryArgs.email,  
> > > operator: 'and'  
> > > }  
> > > }  
> > > }]  
> > > }  
> > > }  
> > > }
> > > 
> > > I have set the explanation to true and see that the score is lower after  
> > > adding the twitter information (to the same document as the tags, location  
> > > and user), but I don't understand why. Does this have to do with the term  
> > > frequency factor _tf_? Or what's the exact explanation for this  
> > > behaviour? Should I keep this twitter specific information in another type  
> > > in the index, with a link (by ID) to it to not affect the score? How can I  
> > > avoid the change in score in general while still adding fields to the  
> > > document?
> > > 
> > > Thanks!

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 4, 2012, 12:15pm UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283/5 "2012-07-04T12:15:21Z")

</div>

> ```
> I have set the explanation to true and see that the score is
> lower after adding the twitter information (to the same
> document as the tags, location and user), but I don't
> understand why. Does this have to do with the term frequency
> factor tf? Or what's the exact explanation for this behaviour?
> Should I keep this twitter specific information in another
> type in the index, with a link (by ID) to it to not affect the
> score?
> 
> ```

Term frequency, field length, document length - it's a complex business.

Have a look at this page for a short summary:  
[http://www.lucenetutorial.com/advanced-topics/scoring.html](http://www.lucenetutorial.com/advanced-topics/scoring.html)

And at this page for a longer version:  
[http://lucene.apache.org/core/3\_6\_0/scoring.html](http://lucene.apache.org/core/3_6_0/scoring.html)

> How can I avoid the change in score in general while still adding  
> fields to the document?

My question is: why does this matter? Scoring is relative. There are  
some situations where these differences may impact your search, but  
generally they'll even out.

You may want to look at the omit\_norms and omit\_term\_freq\_and\_positions  
options on this page:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

clint

---

<div class="post-metadata">

**Author:** ![philipDS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/philipds/32/2817_2.png) [@philipDS](https://discuss.elastic.co/u/philipDS)\
**Post date:** [July 5, 2012, 10:04pm UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283/6 "2012-07-05T22:04:51Z")

</div>

Thanks clint. The omit\_norms or omit\_term\_freq\_and\_positions fields didn't  
help (which is weird, because I really thought that was the problem).

I will read through the Lucene scoring documentation later.

More suggestions always welcome! Thanks.

2012/7/4 Clinton Gormley [clint@traveljury.com](mailto:clint@traveljury.com)

> > ```
> > I have set the explanation to true and see that the score is
> > lower after adding the twitter information (to the same
> > document as the tags, location and user), but I don't
> > understand why. Does this have to do with the term frequency
> > factor tf? Or what's the exact explanation for this behaviour?
> > Should I keep this twitter specific information in another
> > type in the index, with a link (by ID) to it to not affect the
> > score?
> > 
> > ```
> 
> Term frequency, field length, document length - it's a complex business.
> 
> Have a look at this page for a short summary:  
> [Lucene Scoring - Lucene Tutorial.com](http://www.lucenetutorial.com/advanced-topics/scoring.html)
> 
> And at this page for a longer version:  
> [Apache Lucene - Scoring](http://lucene.apache.org/core/3_6_0/scoring.html)
> 
> > How can I avoid the change in score in general while still adding  
> > fields to the document?
> 
> My question is: why does this matter? Scoring is relative. There are  
> some situations where these differences may impact your search, but  
> generally they'll even out.
> 
> You may want to look at the omit\_norms and omit\_term\_freq\_and\_positions  
> options on this page:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> 
> clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:21am UTC](https://discuss.elastic.co/t/scoring-is-lower-after-adding-new-field-to-a-document/8283/7 "2017-07-06T03:21:22Z")

</div>


