# Using a dictionary in es tokenization for filtering?

**URL:** <https://discuss.elastic.co/t/using-a-dictionary-in-es-tokenization-for-filtering/48074>\
**Category:** Elasticsearch\
**Created:** [April 21, 2016, 4:24pm UTC](https://discuss.elastic.co/t/using-a-dictionary-in-es-tokenization-for-filtering/48074 "2016-04-21T16:24:01Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ApproximateIdentity](https://avatars.discourse-cdn.com/v4/letter/a/7993a0/32.png) [@ApproximateIdentity](https://discuss.elastic.co/u/ApproximateIdentity)\
**Post date:** [April 21, 2016, 4:24pm UTC](https://discuss.elastic.co/t/using-a-dictionary-in-es-tokenization-for-filtering/48074/1 "2016-04-21T16:24:01Z")

</div>

I'm talking about functionality similar to the documentation here:

[https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-hunspell-tokenfilter.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-hunspell-tokenfilter.html)

My question is, is it possible to use a dictionary, such as hunspell or custom, to filter out tokens; for example, invalid English words (similar to the python nltk library `nltk.is_english_word(word)` method)? Even though the link I posted refers to a "filter" it doesn't seem to be filtering in the way I understand the term and instead does stemming, but leaves in words that aren't in the dictionary.

Thanks for any help.

---

<div class="post-metadata">

**Author:** ![ApproximateIdentity](https://avatars.discourse-cdn.com/v4/letter/a/7993a0/32.png) [@ApproximateIdentity](https://discuss.elastic.co/u/ApproximateIdentity)\
**Post date:** [April 21, 2016, 7:33pm UTC](https://discuss.elastic.co/t/using-a-dictionary-in-es-tokenization-for-filtering/48074/2 "2016-04-21T19:33:11Z")

</div>

I'll update for anyone who happens upon this question. There is an option in es for this:

[https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-keep-words-tokenfilter.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-keep-words-tokenfilter.html)

You can just use any set of words (say from open source dictionaries online) and put them in a file. Then you use the `keep_words_path` and you're cooking with gas.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:57pm UTC](https://discuss.elastic.co/t/using-a-dictionary-in-es-tokenization-for-filtering/48074/3 "2017-07-05T22:57:11Z")

</div>


