# What are the most popular contextual terms (after/before) of an expression?

**URL:** <https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654>\
**Category:** Elasticsearch\
**Created:** [September 20, 2015, 1:39pm UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654 "2015-09-20T13:39:54Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![softwaredoug](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/softwaredoug/32/22681_2.png) [@softwaredoug](https://discuss.elastic.co/u/softwaredoug)\
**Post date:** [September 21, 2015, 2:05am UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654/2 "2015-09-21T02:05:41Z")

</div>

Well one simple way, depending on the size of your data, is to create an index of bigrams by using a custom analyzer.

So for the input to analysis, you'd have

> the great house at

and instead of breaking it up into words modify analysis to break it up into bigrams (two word tokens) using the [shingle filter](https://www.elastic.co/guide/en/elasticsearch/reference/1.4/analysis-shingle-tokenfilter.html), like

> [the great] [great house] [house at]

A prefix query on `house\ *` here yields all the occurrences of house SPACE some word, then simply do a [terms aggregration](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html), and you'll see an ordering of all the bigrams as a facet, ordered by how frequently the terms occur in the search results. You may need to further [filter this](https://www.elastic.co/guide/en/elasticsearch/reference/2.0/search-aggregations-bucket-filter-aggregation.html) so you don't see every bigram in these documents.

```json
"buckets" : [ 
                {
                    "key" : "house rules",
                    "doc_count" : 52
                },
                {
                    "key" : "house sucks",
                    "doc_count" : 42
                },
               ...
            ]
        }

```

The OTHER direction though is a bit trickier. You may need to duplicate your data to another field to get a different view. You can to wildcard `* house` queries, but they don't perform that well. Instead, you need to reverse the tokens BEFORE you do the prefix query. So in a completely separate field, you want to add a [reverse filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-reverse-tokenfilter.html) to reverse the text AFTER shingling.

So:

> [good house]

becomes for examining the other direction:

> [esuoh doog]

Then repeat the process for the other direction with a `esuoh\ *` query 😄 getting terms aggregations that you'll have to reverse yourself 🙂

Fun problem, Hope that helps

---

_[View the full topic](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654)._
