# Detail questions about significant terms aggregation

**URL:** <https://discuss.elastic.co/t/detail-questions-about-significant-terms-aggregation/16963>\
**Category:** Elasticsearch\
**Created:** [April 11, 2014, 3:58pm UTC](https://discuss.elastic.co/t/detail-questions-about-significant-terms-aggregation/16963 "2014-04-11T15:58:36Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Valentin\_Pletzer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/valentin_pletzer/32/105548_2.png) [@Valentin\_Pletzer](https://discuss.elastic.co/u/Valentin_Pletzer)\
**Post date:** [April 11, 2014, 3:58pm UTC](https://discuss.elastic.co/t/detail-questions-about-significant-terms-aggregation/16963/1 "2014-04-11T15:58:36Z")

</div>

Hi,

first of all: I really love the new significant terms aggregation as well  
as the cardinal count aggregation. Thanks a lot!

I have some detail questions:

- What is bg\_count (I assume background count) but what is the meaning of  
it?
- At first I thought the score values are between 0 and 1 but there are  
much bigger values. Can anyone give me a rough explanation?

Cheers  
Valentin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/0aa544b5-a2a4-40ae-986d-03955a27ea60%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/0aa544b5-a2a4-40ae-986d-03955a27ea60%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Hannes\_Korte](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hannes_korte/32/909_2.png) [@Hannes\_Korte](https://discuss.elastic.co/u/Hannes_Korte)\
**Post date:** [April 11, 2014, 5:40pm UTC](https://discuss.elastic.co/t/detail-questions-about-significant-terms-aggregation/16963/2 "2014-04-11T17:40:14Z")

</div>

Hi Valentin,

> - What is bg\_count (I assume background count) but what is the meaning of  
> it?

The bg\_count is the number of documents, which contain the term in the  
whole index (not just in the search result).

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

> - At first I thought the score values are between 0 and 1 but there are  
> much bigger values. Can anyone give me a rough explanation?

You can see the code of the computation here:

[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/aggregations/bucket/significant/InternalSignificantTerms.java#L94](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/aggregations/bucket/significant/InternalSignificantTerms.java#L94)

This is a summarized version of the formula:

double subsetProb = #relative frequency in the search result#;  
double supersetProb = #relative frequency in the whole index#;  
double absoluteProbChange = subsetProb - supersetProb;  
if (absoluteProbChange \<= 0) {  
return 0;  
}  
double relativeProbChange = (subsetProb / supersetProb);  
return absoluteProbChange \* relativeProbChange;

I guess in the future there will be support for other scorings like  
mutual information, chi squared or information gain.

Best regards,  
Hannes

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/534828FE.20806%40hkorte.com](https://groups.google.com/d/msgid/elasticsearch/534828FE.20806%40hkorte.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Valentin\_Pletzer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/valentin_pletzer/32/105548_2.png) [@Valentin\_Pletzer](https://discuss.elastic.co/u/Valentin_Pletzer)\
**Post date:** [April 12, 2014, 9:29am UTC](https://discuss.elastic.co/t/detail-questions-about-significant-terms-aggregation/16963/3 "2014-04-12T09:29:08Z")

</div>

Hi Hannes,

thanks for the info. Scoring like mutual information sound fun.

Best regards,  
Valentin

On Friday, April 11, 2014 7:40:14 PM UTC+2, Hannes Korte wrote:

> Hi Valentin,
> 
> > - What is bg\_count (I assume background count) but what is the meaning  
> > of  
> > it?
> 
> The bg\_count is the number of documents, which contain the term in the  
> whole index (not just in the search result).
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/search-aggregations-bucket-significantterms-aggregation.html)
> 
> > - At first I thought the score values are between 0 and 1 but there are  
> > much bigger values. Can anyone give me a rough explanation?
> 
> You can see the code of the computation here:
> 
> [https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/aggregations/bucket/significant/InternalSignificantTerms.java#L94](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/aggregations/bucket/significant/InternalSignificantTerms.java#L94)
> 
> This is a summarized version of the formula:
> 
> double subsetProb = #relative frequency in the search result#;  
> double supersetProb = #relative frequency in the whole index#;  
> double absoluteProbChange = subsetProb - supersetProb;  
> if (absoluteProbChange \<= 0) {  
> return 0;  
> }  
> double relativeProbChange = (subsetProb / supersetProb);  
> return absoluteProbChange \* relativeProbChange;
> 
> I guess in the future there will be support for other scorings like  
> mutual information, chi squared or information gain.
> 
> Best regards,  
> Hannes

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/2490b9bd-4531-4964-9f21-6e18d2a92c7e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/2490b9bd-4531-4964-9f21-6e18d2a92c7e%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:36am UTC](https://discuss.elastic.co/t/detail-questions-about-significant-terms-aggregation/16963/4 "2017-07-06T01:36:17Z")

</div>


