# Documents with german umlauts

**URL:** <https://discuss.elastic.co/t/documents-with-german-umlauts/95513>\
**Category:** Elasticsearch\
**Created:** [August 2, 2017, 11:51am UTC](https://discuss.elastic.co/t/documents-with-german-umlauts/95513 "2017-08-02T11:51:36Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![astropanic](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/astropanic/32/20726_2.png) [@astropanic](https://discuss.elastic.co/u/astropanic)\
**Post date:** [August 2, 2017, 11:51am UTC](https://discuss.elastic.co/t/documents-with-german-umlauts/95513/1 "2017-08-02T11:51:36Z")

</div>

I have two documents:

1. {"name": "Drucker"}
2. {"name": "Drücker"}

How I should index it and how the query should be build so I can:

a) find both documents querying for "drucker"  
b) sort the documents according to the search query (the searched document should appear before the others)

Regards,  
Wojciech

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 2, 2017, 12:16pm UTC](https://discuss.elastic.co/t/documents-with-german-umlauts/95513/2 "2017-08-02T12:16:27Z")

</div>

Using an asciifolding token filter would probably help here.

See [https://www.elastic.co/guide/en/elasticsearch/reference/5.5/analysis-asciifolding-tokenfilter.html](https://www.elastic.co/guide/en/elasticsearch/reference/5.5/analysis-asciifolding-tokenfilter.html)

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 2, 2017, 5:52pm UTC](https://discuss.elastic.co/t/documents-with-german-umlauts/95513/3 "2017-08-02T17:52:20Z")

</div>

The asciifolding filter will normalize the extended characters so that  
those words are equivalent. It was solve your first case, but not the  
second. ICU collation might help with the latter, but sorting would be  
language specific and not based on the query:

[https://www.elastic.co/guide/en/elasticsearch/plugins/current/analysis-icu-collation-keyword-field.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/analysis-icu-collation-keyword-field.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 30, 2017, 5:52pm UTC](https://discuss.elastic.co/t/documents-with-german-umlauts/95513/4 "2017-08-30T17:52:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
