# What is the scope of TF & IDF calculation?

**URL:** <https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283>\
**Category:** Elasticsearch\
**Created:** [May 18, 2016, 2:41am UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283 "2016-05-18T02:41:02Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [May 18, 2016, 2:41am UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/1 "2016-05-18T02:41:02Z")

</div>

I am not quire clear how ES calculate TF/IDF in some situations, like cross index/type search, search with filters etc.

Assume I have two indices, index1 and index2, each of which has two types, type1, and type2. All types of all indies have a filed: language which could be used as filter.

1. Cross index search  
GET /index1,index2/type1,type2/\_search

2. Search with filter  
GET /index1/type1/\_search  
{  
"filter": {  
"term": {  
"language": "english"  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 18, 2016, 5:29am UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/2 "2016-05-18T05:29:26Z")

</div>

1. The score is done per shard, then results are compared across all indices and reduced.
2. Filters do not score, they are a simple match or no-match.

See [https://www.elastic.co/guide/en/elasticsearch/reference/2.3/query-filter-context.html](https://www.elastic.co/guide/en/elasticsearch/reference/2.3/query-filter-context.html) for more

---

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [May 18, 2016, 9:56am UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/3 "2016-05-18T09:56:31Z")

</div>

thanks walkolm  
for #1, I mean, the IDF is calculated before filtering or after filtering?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 18, 2016, 10:17am UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/4 "2016-05-18T10:17:55Z")

</div>

You mean with a search in your second point, but over multiple indices?

---

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [May 19, 2016, 4:32pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/5 "2016-05-19T16:32:39Z")

</div>

Sorry for poor expression.  
Let me explain my question with example

Assume I create an index: myindex with one type: mytype

And put 3 docs to /myindex/mytype

{  
"title": "your search you data",  
"language": "english"  
}

{  
"title": "hello there",  
"language": "french"  
}

{  
"title": "I love elasticsearch",  
"language": "english"  
}

And search with following DSL (please ignore the incorrect synctax)  
GET /myindex/mytype/\_search  
{  
"query": {  
"title": "hello"  
}

"filter": {  
"term": {  
"language": "english"  
}  
}  
}

With above filtered query, when ES calculate the IDF value, it counts the frequency of query term in all 3 documents or only in 2 documents whose language is "english"?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 23, 2016, 3:51am UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/6 "2016-05-23T03:51:16Z")

</div>

It will filter only the `"language": "english"`, then score any that pass that filter.

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 23, 2016, 4:21pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/7 "2016-05-23T16:21:14Z")

</div>

The IDF value is per shard, irregardless of the type. The "type" is an  
Elasticsearch construct, and the Lucene shard knows nothing about them.

And if you noticed, I said it was per shard, not even per index. Since each  
index has its own shards, the IDF values are never shared between indices.  
But by default, they are not even shared between the same index. You need  
to enable distributed queries for that to occur. Small performance hit, but  
it is worth it in finely tuned search environments IMHO.

[https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-search-type.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-search-type.html)

Cheers,

Ivan

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 23, 2016, 4:24pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/8 "2016-05-23T16:24:05Z")

</div>

And to answer your last question, the IDF of a term is pre-calculated and  
not dependent on the documents returned. In other words, it is calculated  
pre-filtering.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [May 23, 2016, 5:42pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/9 "2016-05-23T17:42:40Z")

</div>

> [@Ivan](#):
>
> And to answer your last question, the IDF of a term is pre-calculated andnot dependent on the documents returned. In other words, it is calculatedpre-filtering.

It still includes deleted/old copies of updated documents as well.

> [@Ivan](#):
>
> You need to enable distributed queries for that to occur. Small performance hit, but it is worth it in finely tuned search environments IMHO.

I think that depends on uniformity of the index. [This](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-search-type.html#dfs-query-then-fetch) is the setting. In many cases you are better off just having a single shard for small index so they are automatically "uniform". If the index gets large enough to run multiple shards (a couple of GB) then it is worth playing with the search\_type if score is important to you. The default is the default because for lots of people the indexes are pretty uniform and/or deviations in score aren't a huge problem. But the score can deviate quite a bit if you have terms that are genuinely rare.

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 23, 2016, 6:56pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/10 "2016-05-23T18:56:34Z")

</div>

Uniformity is indeed important. IDF values tend to normalize over large  
data sets. The bigger the shard, the better. Rare terms, which is what IDF  
was meant to improve, suffer on multi-sharded indices. I emphasize using  
single shards for test shards when dealing with relevancy. Many issues with  
test cases is simply because there is not enough data for relevant TF/IDF  
values.

And I think the default is because no one uses Elasticsearch for search  
anymore, so why go through the extra search tuning step. 🙂

Ivan

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [May 23, 2016, 7:40pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/11 "2016-05-23T19:40:44Z")

</div>

> [@Ivan](#):
>
> And I think the default is because no one uses Elasticsearch for searchanymore, so why go through the extra search tuning step.

I used it for search, but yeah, lots of use cases aren't search and they'd just be paying the extra query phase price for nothing.

---

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [May 24, 2016, 2:01pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/12 "2016-05-24T14:01:42Z")

</div>

> [@Ivan](#):
>
> And I think the default is because no one uses Elasticsearch for search  
> anymore, so why go through the extra search tuning step. 🙂

Why do you say "no one uses elasticsearch for search anymore..."? Does this conclusion come from statistics of user scenario? And if it is true, does this mean Search functionality (including relevance tuning) will have low pri in ES's road map?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:49pm UTC](https://discuss.elastic.co/t/what-is-the-scope-of-tf-idf-calculation/50283/13 "2017-07-05T22:49:28Z")

</div>


