# U-umlaut search --\> indexing user name müller , search fails for müller but success for muller

**URL:** <https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317>\
**Category:** Elasticsearch\
**Created:** [September 12, 2016, 4:41pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317 "2016-09-12T16:41:32Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![rmsnandha](https://avatars.discourse-cdn.com/v4/letter/r/3be4f8/32.png) [@rmsnandha](https://discuss.elastic.co/u/rmsnandha)\
**Post date:** [September 12, 2016, 4:41pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317/1 "2016-09-12T16:41:32Z")

</div>

indexing user name müller , i am able to search by query muller , but when i query for müller .. it's not returning anything .. could you please help me on this

below is my analyzer configuration.

{  
"settings":{  
"analysis":{  
"analyzer":{  
"myOwn\_index\_analyzer":{  
"tokenizer":[  
"standard"  
],  
"filter":[  
"standard",  
"my\_delimiter",  
"lowercase",  
"icu\_folding"  
]  
},  
"myOwn\_search\_analyzer":{  
"tokenizer":[  
"standard"  
],  
"filter":[  
"standard",  
"my\_delimiter",  
"lowercase",  
"icu\_folding"  
]  
}  
},  
"filter":{  
"my\_delimiter":{  
"type":"word\_delimiter",  
"generate\_word\_parts":true,  
"catenate\_words":true,  
"catenate\_numbers":true,  
"catenate\_all":true,  
"split\_on\_case\_change":true,  
"preserve\_original":true,  
"split\_on\_numerics":true,  
"stem\_english\_possessive":true  
}  
}  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![simi](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@simi](https://discuss.elastic.co/u/simi)\
**Post date:** [September 15, 2016, 12:05pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317/2 "2016-09-15T12:05:07Z")

</div>

I too have a similar issue. My mapping looks pretty much the same. It looks like the analyzer is replacing the umlaut with a standard ascii character but also not keeping the umlaut. My use case is to allow for searches with or without the special characters. So "ju" and "jü" should both work. I'm trying to use the icu\_folder as well - does anyone have any insight or examples of this?

Thanks!

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 15, 2016, 12:27pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317/3 "2016-09-15T12:27:44Z")

</div>

Use german normalizer

[https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-normalization-tokenfilter.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-normalization-tokenfilter.html)

---

<div class="post-metadata">

**Author:** ![simi](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@simi](https://discuss.elastic.co/u/simi)\
**Post date:** [September 15, 2016, 12:47pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317/4 "2016-09-15T12:47:11Z")

</div>

Hi Jorg,

Thanks for the response. I'm currently using the icu\_folding filter which seems to take care of the special characters like the german normalizer you mentioned. But it seems like both of these filters "replace" the special characters with an non-extended ascii equivelant. So "ß: get converted to "ss" and "ö" gets converted to "o". That great, but I want to search to succeed for both. So a prefix query of "das" or "daß" should BOTH return the correct document. I was thinking the preserver original set to true would keep both tokens, but it doesn't seem to do so. I'm still pretty new to ES, so forgive me if I'm asking simple questions.

Thanks!

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 15, 2016, 1:59pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317/5 "2016-09-15T13:59:03Z")

</div>

A token filter works best by reducing words to a base form which can be indexed, and apply the token filter also at search time.

Keeping the original token in the token stream can be achieved by the token filter `keyword_repeat`. It distorts the frequency of words in the index so you must live with it when you wonder about different scoring values. You should add the `unique` filter to avoid double tokens. Highlighting is supposed not to work any more.

Also, when analyzing german, your method is not complete. Folding is just one part. German umlauts are also valid in expanded form: ä-\>ae, ö-\>oe, ü-\>ue and also ae-\>ä, oe-\>ö, and ue-\>ü. This umlaut conversion has to be performed in a grammar context to avoid errors. The Snowball analyzer is able to do this conversion (see below `snowball_german_umlaut`)

Also, there is the ICU normalizer. Normalization is an important step before folding if you don't know how the input text is encoded. It converts characters which might be decomposed into a Unicode normalized form. Unicode does not have a distinction between umlaut and diaresis (trema).

With the correct analyzer, you can index

Köln -\> koln  
Koeln -\> koln  
Koln -\> koln

I have added an `unstemmed` variant, it omits the german word stemming which Snowball performs.

Here is my solution for german:

```auto
{
  "index" : {
    "analysis" : {
      "filter" : {
        "snowball_german_umlaut" : {
          "type" : "snowball",
          "name" : "German2"
        }
      },
      "analyzer" : {
        "stemmed" : {
          "type" : "custom",
          "tokenizer" : "hyphen",
          "filter" : [
            "lowercase",
            "keyword_repeat",
            "icu_normalizer",
            "icu_folding",
            "snowball_german_umlaut",
            "unique"
          ]
        },
        "unstemmed" : {
          "type" : "custom",
          "tokenizer" : "hyphen",
          "filter" : [
            "lowercase",
            "keyword_repeat",
            "icu_normalizer",
            "icu_folding",
            "german_normalize",
            "unique"
          ]
        }
      }
    }
  }
}

```

The tokenizer `hyphen` is one of my custom tokenizers which can preserve composite words (Bindestrichwörter) which are important in german language. You can also use default or whitespace tokenizer instead.

---

<div class="post-metadata">

**Author:** ![simi](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@simi](https://discuss.elastic.co/u/simi)\
**Post date:** [September 16, 2016, 2:29pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317/6 "2016-09-16T14:29:47Z")

</div>

Hey, thanks for the info! I think I've got things working well now.

Cheers!

Simi

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:19pm UTC](https://discuss.elastic.co/t/u-umlaut-search-indexing-user-name-muller-search-fails-for-muller-but-success-for-muller/60317/7 "2017-07-05T22:19:36Z")

</div>


