# Output n-gram frequency distributions in Elasticsearch?

**URL:** <https://discuss.elastic.co/t/output-n-gram-frequency-distributions-in-elasticsearch/154302>\
**Category:** Elasticsearch\
**Created:** [October 27, 2018, 7:17pm UTC](https://discuss.elastic.co/t/output-n-gram-frequency-distributions-in-elasticsearch/154302 "2018-10-27T19:17:10Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![hacker\_21](https://avatars.discourse-cdn.com/v4/letter/h/a6a055/32.png) [@hacker\_21](https://discuss.elastic.co/u/hacker_21)\
**Post date:** [October 27, 2018, 7:17pm UTC](https://discuss.elastic.co/t/output-n-gram-frequency-distributions-in-elasticsearch/154302/1 "2018-10-27T19:17:10Z")

</div>

Say I have a field called "webpages" with an array of webpage plain-text content.

Through Elasticsearch, is there a way I can have it output the top 1 million 6 word long n-gram phrases based on term frequency? Maybe the top 3 million 2 word long n gram phrases?

By frequency, I'm referring to the number of times the n-gram appears in the "webpages" array across the whole index.

I'm thinking maybe ES already computes this information in advance for tf-idf computations, would be useful if I could output it and save it to a text file in a reasonable amount of time.

This would really save me time because all relevant data is already stored in one centralized place - Elasticsearch. Thanks in advance!

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [October 29, 2018, 11:33am UTC](https://discuss.elastic.co/t/output-n-gram-frequency-distributions-in-elasticsearch/154302/2 "2018-10-29T11:33:58Z")

</div>

So this is actually fairly tricky. Elasticsearch has the information as you said, but it's not available for easy retrieval. Term and Doc frequencies are stored on a per-segment basis to facilitate distributed search. There's no global registry of all the term/doc frequencies... to get that information we'd have to compile all the dictionaries from all the segments across the cluster, dedupe, aggregate and order. Which is reasonably expensive 🙂

There's a TermVector API ([https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-termvectors.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-termvectors.html)) which returns some of the information you're interested in, but the API interface is on a per-document basis. It's done this way because the information is expensive to retrieve so we only allow it being pulled out doc-by-doc.

I think the best option for you is to run a composite aggregation ([https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html)) to collect all the ngrams, then order by count client-side. Composite agg is designed to be a memory friendly streaming aggregation, but lacks some niceties like sorting by count for that reason.

More traditional aggregations like the `terms` agg won't work because trying to sort a few million ngrams will quickly break your nodes with out-of-memory operations.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 26, 2018, 11:33am UTC](https://discuss.elastic.co/t/output-n-gram-frequency-distributions-in-elasticsearch/154302/3 "2018-11-26T11:33:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
