# How to calculate Inverse Document Frequency for particular term?

**URL:** <https://discuss.elastic.co/t/how-to-calculate-inverse-document-frequency-for-particular-term/36830>\
**Category:** Elasticsearch\
**Created:** [December 10, 2015, 5:38am UTC](https://discuss.elastic.co/t/how-to-calculate-inverse-document-frequency-for-particular-term/36830 "2015-12-10T05:38:04Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![kartheek91](https://avatars.discourse-cdn.com/v4/letter/k/cc9497/32.png) [@kartheek91](https://discuss.elastic.co/u/kartheek91)\
**Post date:** [December 10, 2015, 5:38am UTC](https://discuss.elastic.co/t/how-to-calculate-inverse-document-frequency-for-particular-term/36830/1 "2015-12-10T05:38:04Z")

</div>

PUT /my\_index/doc/1  
{ "text" : "quick brown fox" }

GET /my\_index/doc/\_search?explain  
{  
"query": {  
"term": {  
"text": "fox"  
}  
}  
}

weight(text:fox in 0) [PerFieldSimilarity]: 0.15342641  
result of:  
fieldWeight in 0 0.15342641  
product of:  
tf(freq=1.0), with freq of 1: 1.0  
idf(docFreq=1, maxDocs=1): 0.30685282  
fieldNorm(doc=0): 0.5

Can anyone can explain how these values are came by manual calcualtion.

---

<div class="post-metadata">

**Author:** ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)\
**Post date:** [December 10, 2015, 1:27pm UTC](https://discuss.elastic.co/t/how-to-calculate-inverse-document-frequency-for-particular-term/36830/2 "2015-12-10T13:27:14Z")

</div>

Hi,

```auto
tf(freq=1.0), with freq of 1: 1.0

```

This is the frequency of your search term in the matched doc.

```auto
idf(docFreq=1, maxDocs=1): 0.30685282

```

The ClassicSimilarity in Lucene calclulates this as: log(numDocs/(docFreq+1)) + 1, so if you fill in the values you get log(1/(1+1)) + 1

```auto
fieldNorm(doc=0): 0.5 

```

Here it gets more complicated and hard to track by hand. Essentially the normalization factor for a document field should lower the score for documents with long fields. Lucene caclulates this already at index time, the formular is roughly `1.0 / Math.sqrt(numTerms)` according to ClassicSimilarity#lengthNorm, so for three terms like in the example you would get `~0.577`. This however is stored as a single byte and later converted back to float, so there are rounding issues as e.g. described [here](http://stackoverflow.com/questions/15135872/lucene-fieldnorm-discrepancy-between-similarity-calculation-and-query-time-value).

If you care to dive deeper into Scoring there's lots of general resources describing TF/IDF (with sometimes slightly different implementation details). The description of Lucenes [TFIDF similarity](http://lucene.apache.org/core/5_0_0/core/org/apache/lucene/search/similarities/TFIDFSimilarity.html) looks complicated but worth taking a look at. For other than simple examples like the one you gave, calculating scores by hand is a complex task.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:32pm UTC](https://discuss.elastic.co/t/how-to-calculate-inverse-document-frequency-for-particular-term/36830/3 "2017-07-05T23:32:02Z")

</div>


