# Accent insensitive search with search analyzer

**URL:** <https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037>\
**Category:** Elasticsearch\
**Created:** [December 22, 2017, 4:49pm UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037 "2017-12-22T16:49:21Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Valentin](https://avatars.discourse-cdn.com/v4/letter/v/f19dbf/32.png) [@Valentin](https://discuss.elastic.co/u/Valentin)\
**Post date:** [December 22, 2017, 4:49pm UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/1 "2017-12-22T16:49:21Z")

</div>

Hi,

I have some problems to configure my ES index.  
I want to perform a case-insensitive, accent-insensitive with special charachters search.

Here is the index creation script :

```
DELETE index-test
PUT index-test
{
  "settings": {
		"analysis": {
			"analyzer": {
				"SearchAnalyzer": {
					"type": "custom",
					"filter": ["lowercase", "asciifolding"],
					"tokenizer": "whitespace"
				},
				"IndexAnalyzer": {
					"type": "custom",
					"filter": ["lowercase", "asciifolding"],
					"tokenizer": "whitespace"
				}
			}
		}
	},
	"mappings": {
		"myType": {
			"properties": {
				"myField": {
					"type": "text",
					"analyzer": "IndexAnalyzer",
					"search_analyzer": "SearchAnalyzer"
				}
			}
		}
	}
}
PUT index-test/myType/1
{
  "myField": "aB-é"
}

```

Consider this search request :

```
GET index-test/myType/_search
{
	"query": {
		"bool": {
			"filter": [{
				"wildcard": {
					"myField": {
						"value": "*{charToSearch}*"
					}
				}
			}]
		}
	}
}

```

It works as expected if you replace {charToSearch} with these characters : "a", "A", "b", "B", "-", "e".  
Unfortunately, that does not work with "é".  
It's like the searchAnalyzer filter "asciifolding" is not performed.

EDIT : sorry for the poor formating of my first message and for the delay

Thanks for any help, Valentin.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 22, 2017, 5:11pm UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/2 "2017-12-22T17:11:25Z")

</div>

Please format your code using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will make your post more readable.

Or use markdown style like:

````
```
CODE
```

````

Could you provide a full recreation script as described in

> [@About the Elasticsearch category](https://discuss.elastic.co/t/about-the-elasticsearch-category/21):
>
> The heart of the free and open Elastic Stack Elasticsearch is a distributed, RESTful search and analytics engine capable of addressing a growing number of use cases. As the heart of the Elastic Stack, it centrally stores your data for lightning fast search, fine‑tuned relevancy, and powerful analytics that scale with ease. warning PLEASE READ THIS SECTION IF IT'S YOUR FIRST POST Some useful links: [elasticsearch reference guide](http://www.elastic.co/guide/en/elasticsearch/reference/current/index.html)[elasticsearch user guide](http://www.elastic.co/guide/en/elasticsearch/guide/current/index.html)[elasticsearch plugins](https://www.elastic.co/guide/en/elasticsearch/plugins/current/index.html)[elasticsearch cl…](https://www.elastic.co/guide/en/elasticsearch/client/index.html)

It will help to better understand what you are doing.  
Please, try to keep the example as simple as possible.

---

<div class="post-metadata">

**Author:** ![Valentin](https://avatars.discourse-cdn.com/v4/letter/v/f19dbf/32.png) [@Valentin](https://discuss.elastic.co/u/Valentin)\
**Post date:** [January 2, 2018, 8:28am UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/3 "2018-01-02T08:28:23Z")

</div>

Hi,

I added a recreation script and formatted the code.

Thanks for any help !

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 2, 2018, 9:37am UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/4 "2018-01-02T09:37:22Z")

</div>

Most likely it's because the wildcard query is not analyzed. See [https://www.elastic.co/guide/en/elasticsearch/reference/6.1/query-dsl-wildcard-query.html](https://www.elastic.co/guide/en/elasticsearch/reference/6.1/query-dsl-wildcard-query.html)

---

<div class="post-metadata">

**Author:** ![Valentin](https://avatars.discourse-cdn.com/v4/letter/v/f19dbf/32.png) [@Valentin](https://discuss.elastic.co/u/Valentin)\
**Post date:** [January 2, 2018, 10:11am UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/5 "2018-01-02T10:11:47Z")

</div>

I didn't see that..

An alternative could be to perform the "asciifolding" on the client side.  
Do you have a better idea to perform this search ?

---

<div class="post-metadata">

**Author:** ![val](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/val/32/138203_2.png) [@val](https://discuss.elastic.co/u/val)\
**Post date:** [January 2, 2018, 10:32am UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/6 "2018-01-02T10:32:45Z")

</div>

Do not use a wildcard query but include an [ngram token filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-ngram-tokenfilter.html) in `IndexAnalyzer` instead so that your text is sliced and diced into smaller text chunks all while being ascii-folded. No need to change the `SearchAnalyzer` though.

---

<div class="post-metadata">

**Author:** ![Valentin](https://avatars.discourse-cdn.com/v4/letter/v/f19dbf/32.png) [@Valentin](https://discuss.elastic.co/u/Valentin)\
**Post date:** [January 2, 2018, 10:35am UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/7 "2018-01-02T10:35:37Z")

</div>

Thanks for the suggestion, i will give it a try.

---

<div class="post-metadata">

**Author:** ![Valentin](https://avatars.discourse-cdn.com/v4/letter/v/f19dbf/32.png) [@Valentin](https://discuss.elastic.co/u/Valentin)\
**Post date:** [January 2, 2018, 12:49pm UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/8 "2018-01-02T12:49:00Z")

</div>

The searchAnalyzer seems to work like I wanted with the ngram token filter.

I need to search for long strings (a GUID for example -\> c163e2b5-5362-e556-490a-867a9cd63bc3), which is 36 charachters.

- What is the max "max\_gram" value to not exceed to avoid performance problem ?
- Is there a better solution than nGram to perform "contains query" for big values ?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 30, 2018, 12:49pm UTC](https://discuss.elastic.co/t/accent-insensitive-search-with-search-analyzer/113037/9 "2018-01-30T12:49:25Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
