# Terms / Documents Matrix

**URL:** <https://discuss.elastic.co/t/terms-documents-matrix/6953>\
**Category:** Elasticsearch\
**Created:** [March 8, 2012, 8:07pm UTC](https://discuss.elastic.co/t/terms-documents-matrix/6953 "2012-03-08T20:07:34Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jimmy\_Krehl](https://avatars.discourse-cdn.com/v4/letter/j/34f0e0/32.png) [@Jimmy\_Krehl](https://discuss.elastic.co/u/Jimmy_Krehl)\
**Post date:** [March 8, 2012, 8:07pm UTC](https://discuss.elastic.co/t/terms-documents-matrix/6953/1 "2012-03-08T20:07:34Z")

</div>

I've been looking for a way to extract n-gram frequencies from  
ElasticSearch as though it were a large table of n-grams by documents. I  
found this thread from about a year ago:

[http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html](http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html)

"3. The above, 1 and 2, talk about having map reduce implemented on the  
"search" aspect. One thing that I would love to also tackle is the "terms"  
aspect of a search engine. Being able to run (streaming) map reduce jobs on  
terms, especially ones with term vector information, can provide a strong  
infrastructure for implementing algos like clustering and the like.

So, yes, it has crossed my mind :), and it is on the roadmap."

I'm wondering what the status of this is today. Is something similar  
supported in a different way? I could begin work on a plugin or I could  
help with a module in development.

Thanks,  
Jim Krehl

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 9, 2012, 6:58pm UTC](https://discuss.elastic.co/t/terms-documents-matrix/6953/2 "2012-03-09T18:58:53Z")

</div>

Nothing has happened on that front, though I still toy with the idea 🙂

On Thursday, March 8, 2012 at 10:07 PM, Jimmy Krehl wrote:

> I've been looking for a way to extract n-gram frequencies from Elasticsearch as though it were a large table of n-grams by documents. I found this thread from about a year ago:
> 
> [http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html](http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html)
> 
> "3. The above, 1 and 2, talk about having map reduce implemented on the "search" aspect. One thing that I would love to also tackle is the "terms" aspect of a search engine. Being able to run (streaming) map reduce jobs on terms, especially ones with term vector information, can provide a strong infrastructure for implementing algos like clustering and the like.
> 
> So, yes, it has crossed my mind :), and it is on the roadmap."
> 
> I'm wondering what the status of this is today. Is something similar supported in a different way? I could begin work on a plugin or I could help with a module in development.
> 
> Thanks,  
> Jim Krehl

---

<div class="post-metadata">

**Author:** ![Jimmy\_Krehl](https://avatars.discourse-cdn.com/v4/letter/j/34f0e0/32.png) [@Jimmy\_Krehl](https://discuss.elastic.co/u/Jimmy_Krehl)\
**Post date:** [March 9, 2012, 8:18pm UTC](https://discuss.elastic.co/t/terms-documents-matrix/6953/3 "2012-03-09T20:18:12Z")

</div>

Is a search plugin the route to go? I'm pretty new to ES and I'm not if  
there's a framework for those. I'm hoping to be able to leverage the  
search infrastructure in ES to distribute the collation of n-grams.  
Googling has lead to me to believe that people link ES's indices to HDFS  
and use Mahout to extract TF/IDF data. I'd prefer using ES entirely,  
however.

Thanks!  
jimmyk

On Friday, March 9, 2012 10:58:53 AM UTC-8, kimchy wrote:

> Nothing has happened on that front, though I still toy with the idea 🙂
> 
> On Thursday, March 8, 2012 at 10:07 PM, Jimmy Krehl wrote:
> 
> I've been looking for a way to extract n-gram frequencies from  
> Elasticsearch as though it were a large table of n-grams by documents. I  
> found this thread from about a year ago:
> 
> [http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html](http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html)
> 
> "3. The above, 1 and 2, talk about having map reduce implemented on the  
> "search" aspect. One thing that I would love to also tackle is the "terms"  
> aspect of a search engine. Being able to run (streaming) map reduce jobs on  
> terms, especially ones with term vector information, can provide a strong  
> infrastructure for implementing algos like clustering and the like.
> 
> So, yes, it has crossed my mind :), and it is on the roadmap."
> 
> I'm wondering what the status of this is today. Is something similar  
> supported in a different way? I could begin work on a plugin or I could  
> help with a module in development.
> 
> Thanks,  
> Jim Krehl

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 10, 2012, 6:45pm UTC](https://discuss.elastic.co/t/terms-documents-matrix/6953/4 "2012-03-10T18:45:39Z")

</div>

It should be possible with a plugin, and it might not be that difficult if you have a very specific use case.

On Friday, March 9, 2012 at 10:18 PM, Jimmy Krehl wrote:

> Is a search plugin the route to go? I'm pretty new to ES and I'm not if there's a framework for those. I'm hoping to be able to leverage the search infrastructure in ES to distribute the collation of n-grams. Googling has lead to me to believe that people link ES's indices to HDFS and use Mahout to extract TF/IDF data. I'd prefer using ES entirely, however.
> 
> Thanks!  
> jimmyk
> 
> On Friday, March 9, 2012 10:58:53 AM UTC-8, kimchy wrote:
> 
> > Nothing has happened on that front, though I still toy with the idea 🙂
> > 
> > On Thursday, March 8, 2012 at 10:07 PM, Jimmy Krehl wrote:
> > 
> > > I've been looking for a way to extract n-gram frequencies from Elasticsearch as though it were a large table of n-grams by documents. I found this thread from about a year ago:
> > > 
> > > [http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html](http://elasticsearch-users.115913.n3.nabble.com/Pseudo-map-reduce-for-searchresults-td2683300.html)
> > > 
> > > "3. The above, 1 and 2, talk about having map reduce implemented on the "search" aspect. One thing that I would love to also tackle is the "terms" aspect of a search engine. Being able to run (streaming) map reduce jobs on terms, especially ones with term vector information, can provide a strong infrastructure for implementing algos like clustering and the like.
> > > 
> > > So, yes, it has crossed my mind :), and it is on the roadmap."
> > > 
> > > I'm wondering what the status of this is today. Is something similar supported in a different way? I could begin work on a plugin or I could help with a module in development.
> > > 
> > > Thanks,  
> > > Jim Krehl

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:36am UTC](https://discuss.elastic.co/t/terms-documents-matrix/6953/5 "2017-07-06T03:36:25Z")

</div>


