# Recreating Google's Ngram Viewer with elasticsearch

**URL:** <https://discuss.elastic.co/t/recreating-googles-ngram-viewer-with-elasticsearch/20655>\
**Category:** Elasticsearch\
**Created:** [November 9, 2014, 7:16pm UTC](https://discuss.elastic.co/t/recreating-googles-ngram-viewer-with-elasticsearch/20655 "2014-11-09T19:16:37Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![jarib](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jarib/32/882_2.png) [@jarib](https://discuss.elastic.co/u/jarib)\
**Post date:** [November 9, 2014, 7:16pm UTC](https://discuss.elastic.co/t/recreating-googles-ngram-viewer-with-elasticsearch/20655/1 "2014-11-09T19:16:37Z")

</div>

Hello,

I'm looking for tips on how to recreate something like Google's Ngram viewer  
[https://books.google.com/ngrams](https://books.google.com/ngrams) with elasticsearch. I have a text corpus  
of \< 500 MB for which this kind of tool would be very valuable.

I've had some success with the shingle token filter  
[http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/analysis-shingle-tokenfilter.html](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/analysis-shingle-tokenfilter.html) and  
the date histogram aggregation  
[http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/search-aggregations-bucket-datehistogram-aggregation.html](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/search-aggregations-bucket-datehistogram-aggregation.html),  
but the results are not ideal: I'd like to get a histogram of word/phrase  
frequencies, not a histogram of how many documents the word/phrase occurs  
in.

It looks like what I need is some kind of combination of shingles, term  
vectors  
[http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/docs-termvectors.html](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/docs-termvectors.html) and the  
date histogram aggregation, but I'm not sure how to proceed. I can improve  
my current approach by breaking the corpus into smaller pieces, i.e. make  
my documents be paragraphs instead of chapters. But what I really want is a  
"shingle frequency date histogram".

Is this something that can be accomplished with elasticsearch?

Jari

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4b37f0a1-4611-4260-85fb-36b4d67c6076%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4b37f0a1-4611-4260-85fb-36b4d67c6076%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:51am UTC](https://discuss.elastic.co/t/recreating-googles-ngram-viewer-with-elasticsearch/20655/2 "2017-07-06T00:51:16Z")

</div>


