# How can I aggregate terms by their tf-idf score in elasticsearch?

**URL:** <https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548>\
**Category:** Elasticsearch\
**Created:** [June 21, 2016, 5:58pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548 "2016-06-21T17:58:02Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![apanimesh061](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/apanimesh061/32/927_2.png) [@apanimesh061](https://discuss.elastic.co/u/apanimesh061)\
**Post date:** [June 21, 2016, 5:58pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/1 "2016-06-21T17:58:02Z")

</div>

Suppose I run a query which returns a total of 1000 documents and want to aggregate the top 500 documents with terms sorted in order of their `tf-idf` scores.

Is it possible to do that in Elasticsearch?

I am using `v2.3.3`.

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [June 21, 2016, 8:06pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/2 "2016-06-21T20:06:38Z")

</div>

I'm not sure I understand what you mean by "terms sorted in order of their `tf-idf` score"? Hits are already returned in relevance ordering.

Are you wanting a list of the top 500 terms from the matching 1000 documents?

Could you perhaps give an example of what you're looking for?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [June 21, 2016, 8:07pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/3 "2016-06-21T20:07:49Z")

</div>

TF is a per-document score so it doesn't make sense to have a unique list of terms each with a single score that includes any notion of TF.  
See the "explain" api instead [https://www.elastic.co/guide/en/elasticsearch/reference/2.3/search-request-explain.html](https://www.elastic.co/guide/en/elasticsearch/reference/2.3/search-request-explain.html)

---

<div class="post-metadata">

**Author:** ![apanimesh061](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/apanimesh061/32/927_2.png) [@apanimesh061](https://discuss.elastic.co/u/apanimesh061)\
**Post date:** [June 21, 2016, 8:20pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/4 "2016-06-21T20:20:51Z")

</div>

Hi,  
Thanks for the reply.

Yes you are right I want the top 50 terms from the matching documents?

I added this to the query I was running:

```auto
"aggregations": {
    "importantTerms": {
      "terms": {
        "size": 25,
        "field" : "title"
      }
    }
  }

```

I got the terms aggregated and sorted by `doc_count`.

I want the sorting to be done by the `tf * idf` value instead. Is it even possible to get `tf` and `idf` of a particular term this way?

I also tried `significant_terms`but it is just too slow.

Just to add to this, the terms that I get are `unigrams`, is there a way to get `bigrams`?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [June 22, 2016, 6:56am UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/5 "2016-06-22T06:56:12Z")

</div>

> [@apanimesh061](#):
>
> tried significant\_termsbut it is just too slow.

Try wrapping it in the sampler aggregation to focus the inspection on only the top N docs rather than all.

---

<div class="post-metadata">

**Author:** ![apanimesh061](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/apanimesh061/32/927_2.png) [@apanimesh061](https://discuss.elastic.co/u/apanimesh061)\
**Post date:** [June 22, 2016, 2:12pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/6 "2016-06-22T14:12:23Z")

</div>

@Mark_Harwood  
I am not sure if this aggregation is correct:

```
"aggs": {
        "sample": {
            "sampler": {
                "shard_size": 10,
                "field" : "title"
            },
            "aggs": {
                "keywords": {
                    "significant_terms": {
                        "field": "title"
                    }
                }
            }
        }
    }

```

When I run this query I get 0 keywords and a `CircuitBreakingException`.

Is there something I am not doing correctly?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [June 22, 2016, 2:16pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/7 "2016-06-22T14:16:50Z")

</div>

Try this:

```
"aggs": {
		"sample": {
			"sampler": {
				"shard_size": 1000
			},
			"aggs": {
				"keywords": {
					"significant_terms": {
						"field": "title"
					}
				}
			}
		}
	}

```

You want a reasonable sample size to get any sensible stats (like a survey - you wouldn't sample just 10 people). You don't need the "field" property - that only gets used to control diversity in the sample.

---

<div class="post-metadata">

**Author:** ![apanimesh061](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/apanimesh061/32/927_2.png) [@apanimesh061](https://discuss.elastic.co/u/apanimesh061)\
**Post date:** [June 22, 2016, 2:24pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/8 "2016-06-22T14:24:59Z")

</div>

@Mark_Harwood

Thanks for the suggestion.

I got a result like:

```
{
  "took": 927,
  "timed_out": false,
  "_shards": {
    "total": 24,
    "successful": 24,
    "failed": 0
  },
  "hits": {
    "total": 79,
    "max_score": 0,
    "hits": []
  },
  "aggregations": {
    "sample": {
      "doc_count": 79,
      "keywords": {
        "doc_count": 79,
        "buckets": [
          {
            "key": "data",
            "doc_count": 3,
            "score": 234.19592738702784,
            "bg_count": 6916
          }
        ]
      }
    }
  }
} 

```

I only get one term. I tried increasing the `shard_size` from `1000` to `10,000` and `20,000` but I got the same result. I also tried adding `"size": 10,` to the aggregations but it is not helping.

Is it right to be getting only one `significant_term`?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [June 22, 2016, 2:29pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/9 "2016-06-22T14:29:48Z")

</div>

> [@apanimesh061](#):
>
> Is it right to be getting only one significant\_term?

I don't know what your search is but it only matches 79 documents in total across 24 shards.  
That could mean each shard is looking for statistically significant changes in a sample that may be as small as only 2 or 3 docs. That doesn't provide enough of a signal. You could lower the min\_doc\_count and shard\_min\_doc\_count from the default settings but that likely won't help - we need a reasonable number of docs before the recommendations will be any good.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:41pm UTC](https://discuss.elastic.co/t/how-can-i-aggregate-terms-by-their-tf-idf-score-in-elasticsearch/53548/10 "2017-07-05T22:41:24Z")

</div>


