# Accent with edge ngram token filter

**URL:** <https://discuss.elastic.co/t/accent-with-edge-ngram-token-filter/324863>\
**Category:** Elasticsearch\
**Created:** [February 7, 2023, 7:10am UTC](https://discuss.elastic.co/t/accent-with-edge-ngram-token-filter/324863 "2023-02-07T07:10:04Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bob\_Guo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bob_guo/32/116843_2.png) [@Bob\_Guo](https://discuss.elastic.co/u/Bob_Guo)\
**Post date:** [February 7, 2023, 7:10am UTC](https://discuss.elastic.co/t/accent-with-edge-ngram-token-filter/324863/1 "2023-02-07T07:10:04Z")

</div>

Hi,

I have a custom analyzer which uses the edge\_ngram token filer. Below is the setup:

```auto
"analysis": {
        "filter": {
          "my_filter": {
            "type": "edge_ngram",
            "min_gram": "1",
            "max_gram": "10"
          }
        },
        "analyzer": {
          "my_analyzer": {
            "filter": ["lowercase", "my_filter"],
            "type": "custom",
            "tokenizer": "standard"
          }
        }
      }

```

When I use the above analyzer to index some Thai contents, it seems that the edge\_ngram filter also takes accent into account when it produces the tokens. For example:

`โจ้ นากา` would give `โ`, `โจ`, `โจ้`, `น`, `นา`, `นาก`, `นากา`. So when someone searches either `โจ้` (with diacritic) or `โจ` (without diacritic), that document will be returned, which is ok.

The problem is that when the user searches the term `โจ` (without diacritic), some documents that contain `โจ้` (with diacritic) have a higher rank than those that contain the exact term `โจ`. I understand why this is happening but not sure how I can solve it. So in this case, I want to give extra scores to those that contain the exact term. Is this possible?

Also, is it possible to disable the accent folding on the edge\_ngram filter so that it wouldn't produce the token `โจ` (without diacritic) in this case?

Any help would be appreciated.

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [February 7, 2023, 11:45am UTC](https://discuss.elastic.co/t/accent-with-edge-ngram-token-filter/324863/2 "2023-02-07T11:45:02Z")

</div>

Hi @Bob_Guo

I think you are working with the [Thai tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-thai-tokenizer.html).  
In that case I would try to use the Thai tokenizer to preserve accents and better punctuate the documents.

---

<div class="post-metadata">

**Author:** ![Bob\_Guo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bob_guo/32/116843_2.png) [@Bob\_Guo](https://discuss.elastic.co/u/Bob_Guo)\
**Post date:** [February 8, 2023, 12:15am UTC](https://discuss.elastic.co/t/accent-with-edge-ngram-token-filter/324863/3 "2023-02-08T00:15:52Z")

</div>

Hi @RabBit_BR ,

Thanks for your reply.

I'm actually using the Thai analyzer on a different field. But the Thai analyzer itself isn't good enough especially when the user only searches for a partial keyword, for example, the Thai analyzer would produce the following tokens: `โจโฟน`, `หมด`, `สดชื่น` for the text `โจโฟน ได้หมดถ้าสดชื่น`. When the user searches for `โจ`, no results will be returned. So I created this custom analyzer with edge\_ngram filter to improve the search results. Below is the mappings:

```auto
{
  "properties": {
    "name": {
      "type": "text",
      "analyzer": "thai",
      "fields": {
        "auto": {
          "type": "text",
          "analyzer": "my_analyzer"
        }
      }
    }
  }
}

```

I've been trying to find a way to make the edge\_ngram filter accent sensitive but no luck. Not sure if this is achievable. I thought this would be a common problem for languages with diacritics but couldn't find anything helpful on the internet.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 8, 2023, 12:16am UTC](https://discuss.elastic.co/t/accent-with-edge-ngram-token-filter/324863/4 "2023-03-08T00:16:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
