# Getting Accented Text Indexed Properly

**URL:** <https://discuss.elastic.co/t/getting-accented-text-indexed-properly/38194>\
**Category:** Elasticsearch\
**Created:** [December 30, 2015, 7:08pm UTC](https://discuss.elastic.co/t/getting-accented-text-indexed-properly/38194 "2015-12-30T19:08:23Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![dadepo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadepo/32/19944_2.png) [@dadepo](https://discuss.elastic.co/u/dadepo)\
**Post date:** [December 30, 2015, 7:08pm UTC](https://discuss.elastic.co/t/getting-accented-text-indexed-properly/38194/1 "2015-12-30T19:08:23Z")

</div>

Following the article here You have an accent [https://www.elastic.co/guide/en/elasticsearch/guide/current/asciifolding-token-filter.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/asciifolding-token-filter.html) I added the following analysis to my index:

```
PUT /blog
{
  "settings": {
    "analysis": {
      "analyzer": {
        "folding": {
          "tokenizer": "standard",
          "filter": ["lowercase", "asciifolding"]
        }
      }
    }
  }
}

```

and according to the article when I test the analysis out like this:

```
GET /my_index?analyzer=folding
My œsophagus caused a débâcle

```

should yield this:

`my, oesophagus, caused, a, debacle`

But this is not what I get. Instead I get the following output:

```
{
   "tokens": [
      {
         "token": "my",
         "start_offset": 0,
         "end_offset": 2,
         "type": "<ALPHANUM>",
         "position": 1
      },
      {
         "token": "sophagus",
         "start_offset": 4,
         "end_offset": 12,
         "type": "<ALPHANUM>",
         "position": 2
      },
      {
         "token": "caused",
         "start_offset": 13,
         "end_offset": 19,
         "type": "<ALPHANUM>",
         "position": 3
      },
      {
         "token": "a",
         "start_offset": 20,
         "end_offset": 21,
         "type": "<ALPHANUM>",
         "position": 4
      },
      {
         "token": "d",
         "start_offset": 22,
         "end_offset": 23,
         "type": "<ALPHANUM>",
         "position": 5
      },
      {
         "token": "b",
         "start_offset": 24,
         "end_offset": 25,
         "type": "<ALPHANUM>",
         "position": 6
      },
      {
         "token": "cle",
         "start_offset": 26,
         "end_offset": 29,
         "type": "<ALPHANUM>",
         "position": 7
      }
   ]
}

```

Any idea why the **débâcle** get's broken down the way it does on my machine?

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 30, 2015, 8:48pm UTC](https://discuss.elastic.co/t/getting-accented-text-indexed-properly/38194/2 "2015-12-30T20:48:43Z")

</div>

`asciifolding` is for ASCII only.

You should use the ICU tokenizer/token filter

[https://www.elastic.co/guide/en/elasticsearch/plugins/current/analysis-icu.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/analysis-icu.html)

---

<div class="post-metadata">

**Author:** ![dadepo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadepo/32/19944_2.png) [@dadepo](https://discuss.elastic.co/u/dadepo)\
**Post date:** [December 31, 2015, 5:41am UTC](https://discuss.elastic.co/t/getting-accented-text-indexed-properly/38194/3 "2015-12-31T05:41:26Z")

</div>

Hi @jprante, Thanks for the suggestion, would check it out, but still that does not explain why I am getting a different result from the article I linked to.

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 31, 2015, 9:54am UTC](https://discuss.elastic.co/t/getting-accented-text-indexed-properly/38194/4 "2015-12-31T09:54:24Z")

</div>

> [@dadepo](#):
>
> [You Have an Accent | Elasticsearch: The Definitive Guide [2.x] | Elastic](https://www.elastic.co/guide/en/elasticsearch/guide/current/asciifolding-token-filter.html)

It must be

```auto
PUT /my_index
{
  "settings": {
    "analysis": {
      "analyzer": {
        "folding": {
          "tokenizer": "standard",
          "filter": ["lowercase", "asciifolding"]
        }
      }
    }
  }
}

POST /my_index/_analyze?analyzer=folding
My œsophagus caused a débâcle

```

Note that you must pass UTF-8 in the POST body.

But, for processing all kind of Unicode characters and for using correct folding, you should use ICU folding in tokenizer/tokenizer filter.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:27pm UTC](https://discuss.elastic.co/t/getting-accented-text-indexed-properly/38194/5 "2017-07-05T23:27:39Z")

</div>


