# Get stem for word in elastic

**URL:** <https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599>\
**Category:** Elasticsearch\
**Created:** [September 3, 2015, 10:58am UTC](https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599 "2015-09-03T10:58:18Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![edeak](https://avatars.discourse-cdn.com/v4/letter/e/a698b9/32.png) [@edeak](https://discuss.elastic.co/u/edeak)\
**Post date:** [September 3, 2015, 10:58am UTC](https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599/1 "2015-09-03T10:58:18Z")

</div>

Hi!

I'd like to pass a sentence to Elastic and want the word's root/etymon/stem (don't know how to call).  
I'm using the latest elastic and if it's possible I'd do it through the REST api.

Thanks!

---

<div class="post-metadata">

**Author:** ![mikemccand](https://avatars.discourse-cdn.com/v4/letter/m/f04885/32.png) [@mikemccand](https://discuss.elastic.co/u/mikemccand)\
**Post date:** [September 3, 2015, 12:59pm UTC](https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599/2 "2015-09-03T12:59:02Z")

</div>

You could use e.g. Porter Stem filter (if the text is english), or Snowball if it's english or many other languages, and then use the analyze API ([https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-analyze.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-analyze.html) ) to see how a given chunk of text is translated to tokens...

---

<div class="post-metadata">

**Author:** ![edeak](https://avatars.discourse-cdn.com/v4/letter/e/a698b9/32.png) [@edeak](https://discuss.elastic.co/u/edeak)\
**Post date:** [September 3, 2015, 1:04pm UTC](https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599/3 "2015-09-03T13:04:22Z")

</div>

> [@mikemccand](#):
>
> Stem

Problem is that the text is hungarian. But I read here: [Getting Started with Languages | Elasticsearch: The Definitive Guide [2.x] | Elastic](https://www.elastic.co/guide/en/elasticsearch/guide/current/language-intro.html)  
that there is a 'built-in' language analyzer, and I want to test that if it's fit for my requirements or not.

EDIT: and the other thing is that I'm not interested in the tokens, but in what is the result of stemming for the words...

---

<div class="post-metadata">

**Author:** ![mikemccand](https://avatars.discourse-cdn.com/v4/letter/m/f04885/32.png) [@mikemccand](https://discuss.elastic.co/u/mikemccand)\
**Post date:** [September 3, 2015, 5:08pm UTC](https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599/4 "2015-09-03T17:08:49Z")

</div>

The stemmer runs after tokenization for these language specific analyzers, I think (not sure if it does for Hungarian though). Just try using the Hungarian analyzer and see how it tokenizes?

---

<div class="post-metadata">

**Author:** ![edeak](https://avatars.discourse-cdn.com/v4/letter/e/a698b9/32.png) [@edeak](https://discuss.elastic.co/u/edeak)\
**Post date:** [September 7, 2015, 1:01pm UTC](https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599/5 "2015-09-07T13:01:00Z")

</div>

Okay, finally I managed to configure a hunspell analyzer with a hungarian dictionary and it works.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:51pm UTC](https://discuss.elastic.co/t/get-stem-for-word-in-elastic/28599/6 "2017-07-05T23:51:47Z")

</div>


