# Issue with asciiFolding filter and accents

**URL:** <https://discuss.elastic.co/t/issue-with-asciifolding-filter-and-accents/31422>\
**Category:** Elasticsearch\
**Created:** [September 30, 2015, 3:19pm UTC](https://discuss.elastic.co/t/issue-with-asciifolding-filter-and-accents/31422 "2015-09-30T15:19:52Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![jb95](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jb95/32/5079_2.png) [@jb95](https://discuss.elastic.co/u/jb95)\
**Post date:** [September 30, 2015, 3:19pm UTC](https://discuss.elastic.co/t/issue-with-asciifolding-filter-and-accents/31422/1 "2015-09-30T15:19:52Z")

</div>

Hi,

i have an issue with the how asciiFolding filter works ...

I explain :

I have an analyzer

```
"folding": {
       "tokenizer": "standard",
      "filter": ["asciifolding"]
 }

```

I thought (in french) the tokens for pate and pâte will be the same --\> pate without accent

But no

```
GET /cac/_analyze?analyzer=folding&text=pate
{
   "tokens": [
      {
         "token": "pate",
         "start_offset": 0,
         "end_offset": 4,
         "type": "<ALPHANUM>",
         "position": 1
      }
   ]
}

```

AND

```
GET /cac/_analyze?analyzer=folding&text=pâte
{
   "tokens": [
      {
         "token": "p",
         "start_offset": 0,
         "end_offset": 1,
         "type": "<ALPHANUM>",
         "position": 1
      },
      {
         "token": "te",
         "start_offset": 2,
         "end_offset": 4,
         "type": "<ALPHANUM>",
         "position": 2
      }
   ]
}

```

Why i hav two tokens with the second word with accent ? ☹ I have searched a lot but nothing ☹ all my tests are bad !

Thank you for your help !

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 30, 2015, 3:49pm UTC](https://discuss.elastic.co/t/issue-with-asciifolding-filter-and-accents/31422/2 "2015-09-30T15:49:59Z")

</div>

Well. Be careful with the tool you are using to send those tests.

It must be sent in UTF-8 otherwise the standard analyzer might produce bad results.

For example, on my french laptop with curl, I get:

```auto
curl -XGET "http://localhost:9200/_analyze?tokenizer=standard&text=pâte&pretty"

```

```auto
{
  "tokens" : [ {
    "token" : "pￃﾢte",
    "start_offset" : 0,
    "end_offset" : 5,
    "type" : "<ALPHANUM>",
    "position" : 1
  } ]
}

```

---

<div class="post-metadata">

**Author:** ![jb95](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jb95/32/5079_2.png) [@jb95](https://discuss.elastic.co/u/jb95)\
**Post date:** [September 30, 2015, 3:52pm UTC](https://discuss.elastic.co/t/issue-with-asciifolding-filter-and-accents/31422/3 "2015-09-30T15:52:08Z")

</div>

yes no soucy i used tools with utf-8 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:47pm UTC](https://discuss.elastic.co/t/issue-with-asciifolding-filter-and-accents/31422/4 "2017-07-05T23:47:20Z")

</div>


