# ICU Analysers for Elastic search

**URL:** <https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542>\
**Category:** Elasticsearch\
**Created:** [January 19, 2016, 1:07pm UTC](https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542 "2016-01-19T13:07:48Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Akhil\_Suresh](https://avatars.discourse-cdn.com/v4/letter/a/91b2a8/32.png) [@Akhil\_Suresh](https://discuss.elastic.co/u/Akhil_Suresh)\
**Post date:** [January 19, 2016, 1:07pm UTC](https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542/1 "2016-01-19T13:07:48Z")

</div>

Do Elastic search support a korean language Analyser? Need help on that

---

<div class="post-metadata">

**Author:** ![igor\_k](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_k/32/1157_2.png) [@igor\_k](https://discuss.elastic.co/u/igor_k)\
**Post date:** [January 19, 2016, 1:19pm UTC](https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542/2 "2016-01-19T13:19:57Z")

</div>

Hi @Akhil_Suresh,

When I worked at Egnyte we where able to tokenize Korean using ICU Tokenizer. Please take a look at this blog post [https://www.egnyte.com/blog/2015/07/indexing-multilingual-documents-with-elasticsearch/](https://www.egnyte.com/blog/2015/07/indexing-multilingual-documents-with-elasticsearch/)

In general ICU will let you tokenize langauges where words are not space delimited (like Korean) and will fold national character to their ascii versions (like in French or Polish, `é --> e`).

Hope this helps.

Thanks,  
Igor

---

<div class="post-metadata">

**Author:** ![Akhil\_Suresh](https://avatars.discourse-cdn.com/v4/letter/a/91b2a8/32.png) [@Akhil\_Suresh](https://discuss.elastic.co/u/Akhil_Suresh)\
**Post date:** [January 19, 2016, 1:24pm UTC](https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542/3 "2016-01-19T13:24:55Z")

</div>

Thanks @igor_k for the response.

This is how i used the language analyzer. I am not able to query out all korean words. Some of them are ok. Please help if any modifications required.

```
  analysis: {
    char_filter: {
      hyphen_mapping: {
        type: "mapping",
        mappings: [
          "-=>"
        ]
      }
    },
    filter: {
      korean_collation: {
        type: "icu_collation",
        language: "ko",
        country: "KR",
        decomposition: "canonical"
      }
    },
    analyzer: {
      custom_with_char_filter: {
        tokenizer: "standard",
        char_filter: [
          "hyphen_mapping"
        ],
        filter: ["standard", "lowercase", "stop", "porter_stem"]
      },
      korean: {
        tokenizer: "icu_tokenizer",
        char_filter: [
          "hyphen_mapping"
        ],
         filter: ["icu_normalizer", "lowercase", "stop", "porter_stem", "korean_collation"]
      }

    }
  }
},
mappings: {
  document: {
    properties: {
```

---

<div class="post-metadata">

**Author:** ![igor\_k](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_k/32/1157_2.png) [@igor\_k](https://discuss.elastic.co/u/igor_k)\
**Post date:** [January 21, 2016, 8:31am UTC](https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542/4 "2016-01-21T08:31:56Z")

</div>

Hi, I never tried to stem Korean words. I think the issue is in your pipeline of filter. You have `porter_stem`, but its [web page](http://tartarus.org/martin/PorterStemmer/) suggests it is english-only stemmer.

> The Porter stemming algorithm (or ‘Porter stemmer’) is a process for removing the commoner morphological and inflexional endings from words in English.

Try removing it. Also, you can start simple, with `icu_tokenizer` and `icu_folding` and see where that will lead you. For example if you use folding you do not need to use `lowercase` filter.

You can start with this example [https://www.found.no/play/gist/81780a22b33efa60f439](https://www.found.no/play/gist/81780a22b33efa60f439) and try your Korean searches there (I do not know Korean, so it is hard for me to give more than a general tips). And then you can build it up if you need more fancy features.

Hope this helps,  
Igor

---

<div class="post-metadata">

**Author:** ![Akhil\_Suresh](https://avatars.discourse-cdn.com/v4/letter/a/91b2a8/32.png) [@Akhil\_Suresh](https://discuss.elastic.co/u/Akhil_Suresh)\
**Post date:** [February 22, 2016, 6:36am UTC](https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542/5 "2016-02-22T06:36:15Z")

</div>

Thanks @igor_k Partial text search for Korean text is not working . For eg: if we search "에프알엘코리아" we will get 100 results but if we search "에프알" i am not getting any results. This text belong to a field name "sections". Do i need to add any particular analyzer for this particular field to enable partial text search? Please help

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:14pm UTC](https://discuss.elastic.co/t/icu-analysers-for-elastic-search/39542/6 "2017-07-05T23:14:32Z")

</div>


