# Elasticsearch word frequency and relations

**URL:** <https://discuss.elastic.co/t/elasticsearch-word-frequency-and-relations/23519>\
**Category:** Elasticsearch\
**Created:** [May 3, 2015, 7:49am UTC](https://discuss.elastic.co/t/elasticsearch-word-frequency-and-relations/23519 "2015-05-03T07:49:07Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![redserpent7](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/redserpent7/32/467_2.png) [@redserpent7](https://discuss.elastic.co/u/redserpent7)\
**Post date:** [May 3, 2015, 7:49am UTC](https://discuss.elastic.co/t/elasticsearch-word-frequency-and-relations/23519/1 "2015-05-03T07:49:07Z")

</div>

Hi,

I am wondering if it is possible at all to get the top ten most frequent  
words in an Elasticsearch field across an entire index or alias.

Here is what I'm trying to do:

I am indexing text documents extracted from various document types (Word,  
Powerpoint, PDF, etc) these are analyzed and stored in a field called  
doc\_content. I would like to know if there is a way to find the most  
frequent word(s) in a particular index that are stored in the doc\_content  
field.

To make it clearer, lets assume I am indexing invoices from Amazon and eBay  
for example. Now lets assume I have 100 invoices from amazon and 20  
invoices from ebay. Lets also assume that the word "amazon" occurs twice in  
each amazon invoice and the word "ebay" occurs 3 times in each ebay  
invoice.

Now, is there a way to get an aggregate of sort that tells me that the word  
"amazon" appears in my index 200 times (100 invoices x 2  
occurrences/invoice) and the word "ebay" occurs 60 times (20 invoices x 3  
occurrences/invoice).

My other question is if the former is possible, then is there a way to  
determine what is the most frequent word that comes after a certain word?

For example: lets assume I have 100 documents. 60 of these documents  
contains the term "Old Cat" and 40 contains the term "Old Dog" and for the  
sake of argument lets assume that these words only appear once in each  
document.

Now, if we can get the frequency of the word "old" which in our case should  
be 100. Can we then determine a relation to the word that comes right after  
it to have something like this:

```
          __________ Cat (60)
          |

```

Old (100) |  
|\_\_\_\_\_\_\_\_\_\_ Dog (40)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b8056758-902f-4361-bb60-a8930aaa9725%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b8056758-902f-4361-bb60-a8930aaa9725%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![e\_moeller](https://avatars.discourse-cdn.com/v4/letter/e/ce7236/32.png) [@e\_moeller](https://discuss.elastic.co/u/e_moeller)\
**Post date:** [May 4, 2015, 3:32am UTC](https://discuss.elastic.co/t/elasticsearch-word-frequency-and-relations/23519/2 "2015-05-04T03:32:47Z")

</div>

Similar question as Zaid's first question - keyword extraction along TF-IDF  
logic. Specifically, I have a corpus of ~10K articles and am looking to  
get a ranking of all the tokenized terms in each article based on their  
frequency in the article and the terms relative frequency across the  
corpus. Thanks!

On Sunday, May 3, 2015 at 3:49:07 AM UTC-4, Zaid Amir wrote:

> Hi,
> 
> I am wondering if it is possible at all to get the top ten most frequent  
> words in an Elasticsearch field across an entire index or alias.
> 
> Here is what I'm trying to do:
> 
> I am indexing text documents extracted from various document types (Word,  
> Powerpoint, PDF, etc) these are analyzed and stored in a field called  
> doc\_content. I would like to know if there is a way to find the most  
> frequent word(s) in a particular index that are stored in the doc\_content  
> field.
> 
> To make it clearer, lets assume I am indexing invoices from Amazon and  
> eBay for example. Now lets assume I have 100 invoices from amazon and 20  
> invoices from ebay. Lets also assume that the word "amazon" occurs twice in  
> each amazon invoice and the word "ebay" occurs 3 times in each ebay  
> invoice.
> 
> Now, is there a way to get an aggregate of sort that tells me that the  
> word "amazon" appears in my index 200 times (100 invoices x 2  
> occurrences/invoice) and the word "ebay" occurs 60 times (20 invoices x 3  
> occurrences/invoice).
> 
> My other question is if the former is possible, then is there a way to  
> determine what is the most frequent word that comes after a certain word?
> 
> For example: lets assume I have 100 documents. 60 of these documents  
> contains the term "Old Cat" and 40 contains the term "Old Dog" and for the  
> sake of argument lets assume that these words only appear once in each  
> document.
> 
> Now, if we can get the frequency of the word "old" which in our case  
> should be 100. Can we then determine a relation to the word that comes  
> right after it to have something like this:
> 
> ```
> __________ Cat (60)
> |
> 
> ```
> 
> Old (100) |  
> |\_\_\_\_\_\_\_\_\_\_ Dog (40)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/aebf97f7-e20f-4d8e-a513-f79df4256b71%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/aebf97f7-e20f-4d8e-a513-f79df4256b71%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:16am UTC](https://discuss.elastic.co/t/elasticsearch-word-frequency-and-relations/23519/3 "2017-07-06T00:16:04Z")

</div>


