# ICU Analysis Plugin doesn't normalize some characters like other languages(PHP, Python)

**URL:** <https://discuss.elastic.co/t/icu-analysis-plugin-doesnt-normalize-some-characters-like-other-languages-php-python/314155>\
**Category:** Elasticsearch\
**Created:** [September 12, 2022, 2:15am UTC](https://discuss.elastic.co/t/icu-analysis-plugin-doesnt-normalize-some-characters-like-other-languages-php-python/314155 "2022-09-12T02:15:43Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![mictsai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mictsai/32/110728_2.png) [@mictsai](https://discuss.elastic.co/u/mictsai)\
**Post date:** [September 12, 2022, 2:15am UTC](https://discuss.elastic.co/t/icu-analysis-plugin-doesnt-normalize-some-characters-like-other-languages-php-python/314155/1 "2022-09-12T02:15:43Z")

</div>

Elasticsearch version: 7.10.1  
Installed plugins: [analysis-icu, analysis-kuromoji, analysis-nori]

그래비티 and 그래비티 looks the same. When encoding to URL, 그래비티 is %EA%B7%B8%EB%9E%98%EB%B9%84%ED%8B%B0%20 and 그래비티 is %E1%84%80%E1%85%B3%E1%84%85%E1%85%A2%E1%84%87%E1%85%B5%E1%84%90%E1%85%B5. The difference can be seen by the length of the encoding result.

To solve this issue, I start a test on Python.

PYTHON 3.8.13

```auto
import unicodedata

def normalize(word):
    print(len(word), word)
    normalized_word = unicodedata.normalize('NFKC', word)
    print(len(normalized_word), normalized_word)
    
for word in ['그래비티', '그래비티']:
    normalize(word)

```

The output do solve the problem

```auto
8 그래비티
4 그래비티
4 그래비티
4 그래비티

```

When I test on elasticsearch, using the icu analyzer.  
The result don't change.

GET /\_analyze

```auto
{
    "char_filter": [
        {
            "type": "icu_normalizer",
            "name": "nfkc"
        }
    ],
    "text": "그래비티"
}

```

The result is the same

```auto
그래비티
{
    "tokens": [
        {
            "token": "그래비티",
            "start_offset": 0,
            "end_offset": 4,
            "type": "word",
            "position": 0
        }
    ]
}
그래비티
{
    "tokens": [
        {
            "token": "그래비티",
            "start_offset": 0,
            "end_offset": 8,
            "type": "word",
            "position": 0
        }
    ]
}

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 10, 2022, 2:15am UTC](https://discuss.elastic.co/t/icu-analysis-plugin-doesnt-normalize-some-characters-like-other-languages-php-python/314155/2 "2022-10-10T02:15:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
