# What are the most popular contextual terms (after/before) of an expression?

**URL:** <https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654>\
**Category:** Elasticsearch\
**Created:** [September 20, 2015, 1:39pm UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654 "2015-09-20T13:39:54Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![nicom](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicom/32/488_2.png) [@nicom](https://discuss.elastic.co/u/nicom)\
**Post date:** [September 20, 2015, 1:39pm UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654/1 "2015-09-20T13:39:54Z")

</div>

Hi,

I have a text field.

I would like to get a list of all the most popular contextual terms related to an expression e.g. "great house". By context I mean most popular terms next or before the expression found in the corpus. e.g. xx great house xx.

if lots of documents have in the text "nice great house" -\> "nice" should be in such a list.

How to do such this in ES? / is ES the right tools for that?

---

<div class="post-metadata">

**Author:** ![softwaredoug](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/softwaredoug/32/22681_2.png) [@softwaredoug](https://discuss.elastic.co/u/softwaredoug)\
**Post date:** [September 21, 2015, 2:05am UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654/2 "2015-09-21T02:05:41Z")

</div>

Well one simple way, depending on the size of your data, is to create an index of bigrams by using a custom analyzer.

So for the input to analysis, you'd have

> the great house at

and instead of breaking it up into words modify analysis to break it up into bigrams (two word tokens) using the [shingle filter](https://www.elastic.co/guide/en/elasticsearch/reference/1.4/analysis-shingle-tokenfilter.html), like

> [the great] [great house] [house at]

A prefix query on `house\ *` here yields all the occurrences of house SPACE some word, then simply do a [terms aggregration](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html), and you'll see an ordering of all the bigrams as a facet, ordered by how frequently the terms occur in the search results. You may need to further [filter this](https://www.elastic.co/guide/en/elasticsearch/reference/2.0/search-aggregations-bucket-filter-aggregation.html) so you don't see every bigram in these documents.

```json
"buckets" : [ 
                {
                    "key" : "house rules",
                    "doc_count" : 52
                },
                {
                    "key" : "house sucks",
                    "doc_count" : 42
                },
               ...
            ]
        }

```

The OTHER direction though is a bit trickier. You may need to duplicate your data to another field to get a different view. You can to wildcard `* house` queries, but they don't perform that well. Instead, you need to reverse the tokens BEFORE you do the prefix query. So in a completely separate field, you want to add a [reverse filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-reverse-tokenfilter.html) to reverse the text AFTER shingling.

So:

> [good house]

becomes for examining the other direction:

> [esuoh doog]

Then repeat the process for the other direction with a `esuoh\ *` query 😄 getting terms aggregations that you'll have to reverse yourself 🙂

Fun problem, Hope that helps

---

<div class="post-metadata">

**Author:** ![nicom](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicom/32/488_2.png) [@nicom](https://discuss.elastic.co/u/nicom)\
**Post date:** [September 21, 2015, 5:18am UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654/3 "2015-09-21T05:18:59Z")

</div>

Great doug,

1. thanks for the the "before term" trick !

2. about shingle filter, is there a way in ES to do a skip-gram modeling?

---

<div class="post-metadata">

**Author:** ![softwaredoug](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/softwaredoug/32/22681_2.png) [@softwaredoug](https://discuss.elastic.co/u/softwaredoug)\
**Post date:** [September 21, 2015, 4:03pm UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654/4 "2015-09-21T16:03:38Z")

</div>

Not that I know of. Probably not directly in ES, but that's not quite my baliwick. My coauthor [John Berryman](https://twitter.com/JnBrymn) who's much more of a data scientist would probably know better than I, you might try pinging him?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:49pm UTC](https://discuss.elastic.co/t/what-are-the-most-popular-contextual-terms-after-before-of-an-expression/29654/5 "2017-07-05T23:49:01Z")

</div>


