# Analyze German words with umlauts

**URL:** <https://discuss.elastic.co/t/analyze-german-words-with-umlauts/27743>\
**Category:** Elasticsearch\
**Created:** [August 20, 2015, 9:14am UTC](https://discuss.elastic.co/t/analyze-german-words-with-umlauts/27743 "2015-08-20T09:14:25Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![usarskyy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/usarskyy/32/4017_2.png) [@usarskyy](https://discuss.elastic.co/u/usarskyy)\
**Post date:** [August 20, 2015, 9:14am UTC](https://discuss.elastic.co/t/analyze-german-words-with-umlauts/27743/1 "2015-08-20T09:14:25Z")

</div>

Hello everyone!

I have a German word with umlaut, lets say it is "läuft". My target is to create an analyzer that produces three tokens at the end: "läuft", "laeuft" and "lauft".

I have tried different combinations with _icu\_normalizer_, _asciifolding_ and _snowball for German2_ filters but no results. The best result I've got from _asciifolding_ token filter that emits two out of three required tokens: "läuft" and "lauft".

So, basically, I need to create some kind of custom _asciifolding_ filter for German language that will allow to emit additional variations for words with umlauts.

My configuration for _asciifolding_ and _snowball_ filters are the following:

```
"ascii2": {
              "type": "asciifolding",
              "preserve_original": "true"
            },

"snow-german2": {
              "type": "snowball",
              "language": "German2"
            },

```

I would be really appreciated for your help!

---

<div class="post-metadata">

**Author:** ![francoisguerin](https://avatars.discourse-cdn.com/v4/letter/f/c6cbf5/32.png) [@francoisguerin](https://discuss.elastic.co/u/francoisguerin)\
**Post date:** [August 20, 2015, 4:14pm UTC](https://discuss.elastic.co/t/analyze-german-words-with-umlauts/27743/2 "2015-08-20T16:14:33Z")

</div>

Hi,  
You should try the Combo analyzer plugin : [https://github.com/yakaz/elasticsearch-analysis-combo/](https://github.com/yakaz/elasticsearch-analysis-combo/)  
it can combine multiple analyzers. For example, the one you mentioned (läuft =\> läuft, lauft) and another one (läuft =\> laeuft), with a regexp (char mapping or pattern replace).

---

<div class="post-metadata">

**Author:** ![usarskyy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/usarskyy/32/4017_2.png) [@usarskyy](https://discuss.elastic.co/u/usarskyy)\
**Post date:** [August 20, 2015, 11:30pm UTC](https://discuss.elastic.co/t/analyze-german-words-with-umlauts/27743/3 "2015-08-20T23:30:14Z")

</div>

Yes, we came up to the same conclusion on SO ([http://stackoverflow.com/questions/32114129/elasticsearch-analyzer-for-german-language](http://stackoverflow.com/questions/32114129/elasticsearch-analyzer-for-german-language)). It seems to be the only possible solution for now.

Thank you for your advise!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:54pm UTC](https://discuss.elastic.co/t/analyze-german-words-with-umlauts/27743/4 "2017-07-05T23:54:42Z")

</div>


