# Get unique words (tokens ?) from all text fields from all documents in an Index

**URL:** <https://discuss.elastic.co/t/get-unique-words-tokens-from-all-text-fields-from-all-documents-in-an-index/385424>\
**Category:** Elasticsearch\
**Created:** [March 12, 2026, 1:20pm UTC](https://discuss.elastic.co/t/get-unique-words-tokens-from-all-text-fields-from-all-documents-in-an-index/385424 "2026-03-12T13:20:47Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![sivakumarp11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sivakumarp11/32/99470_2.png) [@sivakumarp11](https://discuss.elastic.co/u/sivakumarp11)\
**Post date:** [March 12, 2026, 1:20pm UTC](https://discuss.elastic.co/t/get-unique-words-tokens-from-all-text-fields-from-all-documents-in-an-index/385424/1 "2026-03-12T13:20:47Z")

</div>

Is there a way to get list of ‘unique’ words Or terms used in a large text field in ALL the documents ? for ex. if the text field has content like _ **‘It is the most confidential because confidential or transcendental knowledge involves understanding the difference between the soul and the body’** _ then the output should be something like -

_ **confidential, transcendental, knowledge, understanding, soul, body** _

---

<div class="post-metadata">

**Author:** ![Rafa\_Silva](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rafa_silva/32/147814_2.png) [@Rafa\_Silva](https://discuss.elastic.co/u/Rafa_Silva)\
**Post date:** [March 14, 2026, 12:54am UTC](https://discuss.elastic.co/t/get-unique-words-tokens-from-all-text-fields-from-all-documents-in-an-index/385424/2 "2026-03-14T00:54:14Z")

</div>

Yes, but it’s not straightforward with a large text field.  
Elasticsearch stores analyzed tokens in the inverted index, but it does not provide a simple way to retrieve a full global list of unique terms from all documents.  
If you need something like this, a few common approaches are:  
Use term vectors to inspect the tokens generated for a field.  
Use the \_analyze API to see how the text is tokenized.  
If aggregation of unique values is required, store the terms in a separate field (keyword or normalized) during ingestion and run a terms aggregation on that field.  
In practice, if this is a recurring requirement, the best approach is usually to extract the relevant terms during ingestion and index them in a dedicated field so they can be aggregated efficiently.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood1/32/101255_2.png) [@Mark\_Harwood1](https://discuss.elastic.co/u/Mark_Harwood1)\
**Post date:** [March 14, 2026, 8:16am UTC](https://discuss.elastic.co/t/get-unique-words-tokens-from-all-text-fields-from-all-documents-in-an-index/385424/3 "2026-03-14T08:16:29Z")

</div>

Adding to what Rafa already said - you can get the most popular words from a random sample of content using something like this aggregation: [Background stats query (extracts popular words and counts from a random sample for use as background in significant text analysis) · GitHub](https://gist.github.com/markharwood/74bb6b8523f6a5746b1b758da2a5372e)

It’s too expensive to do this on too many documents and would cause memory issues so keep the numbers low.

---

<div class="post-metadata">

**Author:** ![sivakumarp11](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sivakumarp11/32/99470_2.png) [@sivakumarp11](https://discuss.elastic.co/u/sivakumarp11)\
**Post date:** [March 15, 2026, 5:26am UTC](https://discuss.elastic.co/t/get-unique-words-tokens-from-all-text-fields-from-all-documents-in-an-index/385424/4 "2026-03-15T05:26:40Z")

</div>

Thank you so much Rafa, Mark for taking time to share such a detailed response, will review the same. I’m afraid, that I have to go for some custom logic, since the requirements are getting complex, as I get into more details.
