# Ways to handle umlauts

**URL:** <https://discuss.elastic.co/t/ways-to-handle-umlauts/91435>\
**Category:** Elasticsearch\
**Created:** [June 30, 2017, 1:18pm UTC](https://discuss.elastic.co/t/ways-to-handle-umlauts/91435 "2017-06-30T13:18:56Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![skauk](https://avatars.discourse-cdn.com/v4/letter/s/b9bd4f/32.png) [@skauk](https://discuss.elastic.co/u/skauk)\
**Post date:** [June 30, 2017, 1:18pm UTC](https://discuss.elastic.co/t/ways-to-handle-umlauts/91435/1 "2017-06-30T13:18:56Z")

</div>

I will very much appreciate an advice from the community on the best practices of handling umlauts for search.  
What I have right now in my setup which is a mix of German and English is a `asciifolding` token filter with preserving the original which covers 90% of use cases.  
In effect what it does is it emits an additional token for each token containing an umlaut with the umlaut replaced with a single character. However, to cover the rest 10% of cases I would like to consider words which are written with expanded umlauts. So for "Köln" I would like all of the following to be able to yield a match:

- köln
- koln
- koeln

I've tried to add the missing third variant by using `german_normalization` filter. It works as intended but because it uses just simple substitution it also mangles words like "Raphael" to "raphal" which is something I don't want.  
It appears that a good solution would be normalization filter which also preserves original token. However I can't find a way to create such a filter chain.

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [June 30, 2017, 3:19pm UTC](https://discuss.elastic.co/t/ways-to-handle-umlauts/91435/2 "2017-06-30T15:19:20Z")

</div>

Unfortunately, there is no "on size fits all" solution.

Mixing german and english words in the index is generally not a good idea, but it should work for umlauts, because they do not often appear in english words.

I use two alternatives, one with stemming using "German2" snowball stemmer, the other without stemming. See

> <https://github.com/jprante/elasticsearch-plugin-bundle/blob/master/src/test/resources/org/xbib/elasticsearch/index/analysis/german/unstemmed.json>

and

> <https://github.com/jprante/elasticsearch-plugin-bundle/blob/master/src/test/java/org/xbib/elasticsearch/index/analysis/german/UnstemmedTests.java>

The stemming variant may not be acceptable, because many german words collapse into the same word form in the index (known as "overstemming"). I try to soften this effect by indexing the original word form, too, with the help of the keyword repeat filter.

Because I keep the original form, I do not protect words like "Raphael" from stemming but it should be possible with the keyword marker token filter, see

[https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-keyword-marker-tokenfilter.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-keyword-marker-tokenfilter.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 28, 2017, 3:19pm UTC](https://discuss.elastic.co/t/ways-to-handle-umlauts/91435/3 "2017-07-28T15:19:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
