# Problem searching queries with accents

**URL:** <https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401>\
**Category:** Elasticsearch\
**Created:** [January 16, 2012, 4:24pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401 "2012-01-16T16:24:06Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Felipe\_Hummel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_hummel/32/1263_2.png) [@Felipe\_Hummel](https://discuss.elastic.co/u/Felipe_Hummel)\
**Post date:** [January 16, 2012, 4:24pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/1 "2012-01-16T16:24:06Z")

</div>

Hi, I'm indexing brazilian portuguese text that contains accents. To remove  
them I used the asciifolding filter. My "test" index settings is as follows:

{

```
"test": {

    "settings": {

        "index.analysis.analyzer.default.filter.0": "standard",

        "index.analysis.analyzer.default.tokenizer": "standard",

        "index.analysis.analyzer.default.filter.1": "lowercase",

        "index.analysis.analyzer.default.filter.2": "stop",

        "index.analysis.analyzer.default.filter.3": "asciifolding",

        "index.number_of_shards": "1",

        "index.number_of_replicas": "0"

    }

}

```

}

I indexed the word "não". When I search "nao" (no accent) the document is  
retrieved. If I search for "não" no document is retrieved.

Something wrong with my configuration?

I'm using Curl to query elasticsearch.

Thanks

Felipe Hummel

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [January 16, 2012, 6:18pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/2 "2012-01-16T18:18:59Z")

</div>

> I indexed the word "nÃ£o". When I search "nao" (no accent) the document  
> is retrieved. If I search for "nÃ£o" no document is retrieved.

How are you searching? I bet you're using a 'term' query, which isn't  
analyzed. Change that to a 'text' query, and it should work

clint

---

<div class="post-metadata">

**Author:** ![Felipe\_Hummel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_hummel/32/1263_2.png) [@Felipe\_Hummel](https://discuss.elastic.co/u/Felipe_Hummel)\
**Post date:** [January 16, 2012, 6:59pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/3 "2012-01-16T18:59:37Z")

</div>

That is right!

Actually I was also testing with the form:

[http://localhost:9200/test/test1/\_search?q=não](http://localhost:9200/test/test1/_search?q=n%C3%A3o)

I suppose it just gets converted to a TermQuery. Because the following  
query:

[http://localhost:9200/test/teste1/\_search?q=não+something](http://localhost:9200/test/teste1/_search?q=n%C3%A3o+something)

yields the right results.

Thanks

Felipe Hummel

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [January 16, 2012, 7:05pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/4 "2012-01-16T19:05:55Z")

</div>

> Actually I was also testing with the form:

> [http://localhost:9200/test/test1/\_search?q=nÃ£o](http://localhost:9200/test/test1/_search?q=n%C3%83%C2%A3o)

Actually, that gets converted to a query\_string query against the \_all  
field, which should have worked.

I wonder if it was a problem with your encoding.

Does this work?

curl -XGET '[http://127.0.0.1:9200/test/test1/\_search?pretty=1&q=não](http://127.0.0.1:9200/test/test1/_search?pretty=1&q=n%C3%A3o)

clint

---

<div class="post-metadata">

**Author:** ![Felipe\_Hummel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_hummel/32/1263_2.png) [@Felipe\_Hummel](https://discuss.elastic.co/u/Felipe_Hummel)\
**Post date:** [January 16, 2012, 7:09pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/5 "2012-01-16T19:09:02Z")

</div>

You're right, it must be some encoding problem. The url encoded version  
works as expected.

Felipe Hummel

---

<div class="post-metadata">

**Author:** ![Frederic](https://avatars.discourse-cdn.com/v4/letter/f/4491bb/32.png) [@Frederic](https://discuss.elastic.co/u/Frederic)\
**Post date:** [January 16, 2012, 8:25pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/6 "2012-01-16T20:25:35Z")

</div>

Hi Clint, i take advantage of this thread for a quite similar  
question:

I've been performing some searches (ES 0.18.5) using accents as well  
(in spanish) but, instead of using the 'asciifolding' filter at a  
indexing time, I'd want to get similar results using fuzzy queries  
(apart from getting also results for similar words).

I want to search docs using one or more words, based on a free text  
field, called 'title', and I'd like to get the same results for both,  
for instance, "bateria" and "batería" words. The query is:

```
  "query" : {
    "text" : {
      "title" : {
        "query" : "batería",
        "type" : "boolean",
        "operator" : "AND",
        "fuzziness" : "0.7",
        "max_expansions" : 3
      }
    }
  }

```

AFAIK a word with one accented letter is at distance '1' from the same  
word with no accent, is this correct? If so, the query should consider  
all docs that contains, in this case "batería" and "batería", right?

Right now I'm getting different number of results and I'm not sure  
what could be the reason

Thanks in advance

Frederic

On 16 ene, 16:05, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> > Actually I was also testing with the form:  
> > [http://localhost:9200/test/test1/\_search?q=não](http://localhost:9200/test/test1/_search?q=n%C3%A3o)
> 
> Actually, that gets converted to a query\_string query against the \_all  
> field, which should have worked.
> 
> I wonder if it was a problem with your encoding.
> 
> Does this work?
> 
> curl -XGET '[http://127.0.0.1:9200/test/test1/\_search?pretty=1&q=não](http://127.0.0.1:9200/test/test1/_search?pretty=1&q=n%C3%A3o)
> 
> clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [January 17, 2012, 11:45am UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/7 "2012-01-17T11:45:04Z")

</div>

> I've been performing some searches (ES 0.18.5) using accents as well  
> (in spanish) but, instead of using the 'asciifolding' filter at a  
> indexing time, I'd want to get similar results using fuzzy queries  
> (apart from getting also results for similar words).
> 
> I want to search docs using one or more words, based on a free text  
> field, called 'title', and I'd like to get the same results for both,  
> for instance, "bateria" and "baterÃ­a" words. The query is:
> 
> ```
> "query" : {
> "text" : {
> "title" : {
> "query" : "baterÃ­a",
> "type" : "boolean",
> "operator" : "AND",
> "fuzziness" : "0.7",
> "max_expansions" : 3
> }
> }
> }
> 
> ```
> 
> AFAIK a word with one accented letter is at distance '1' from the same  
> word with no accent, is this correct? If so, the query should consider  
> all docs that contains, in this case "baterÃ­a" and "baterÃ­a", right?

If you change the "fuzziness" factor to 0.5, it will probably work. I  
don't understand exactly what that number represents, so can't give you  
more than a trial-and-error approach 🙂

That said, using a fuzzy query for this type of search is a lot heavier  
than analyzing your text properly at index time.

clint

---

<div class="post-metadata">

**Author:** ![Frederic](https://avatars.discourse-cdn.com/v4/letter/f/4491bb/32.png) [@Frederic](https://discuss.elastic.co/u/Frederic)\
**Post date:** [January 17, 2012, 4:05pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/8 "2012-01-17T16:05:27Z")

</div>

Thanks for your answer Clint, some comments:

> If you change the "fuzziness" factor to 0.5, it will probably work.  
> Not really actually as a factor of 0.7 should be enough for matching  
> words at a distance of 1.

> I don't understand exactly what that number represents, so can't give you  
> more than a trial-and-error approach 🙂  
> Just for the sake of providing info about this topic (this is what I  
> know so far, most likely Kimchy or some other Lucene expert will know  
> the right answer):

The 'fuzziness' factor refers to the 'minimunSimilarity' parameter of  
a Lucene FuzzyQuery ([Index of /\_\_root/docs.lucene.apache.org/core/3\_2\_0/api/all/org](http://lucene.apache.org/java/3_2_0/api/all/org/)  
apache/lucene/search/Query.html): for a minimumSimilarity of 0.7, a  
term of the same length as the query term is considered similar to the  
query term if the edit distance between both terms is less than  
length(term)\*(1-0.7)

Where the distance value is based on an implementation of the'  
Levenshtein Distance' algorithm ([http://www.merriampark.com/ld.htm](http://www.merriampark.com/ld.htm)).

Thus, LD between "bateria" and "batería" is 1 (just one char change)  
and length('batería')\*0.3 = 2.1 \> 1

> That said, using a fuzzy query for this type of search is a lot heavier  
> than analyzing your text properly at index time.

Totally agree, it's just that in my case I need to work in an already  
productive system with 50M docs indexed, so I cannot recreate the  
index for changing the 'title' field analyzer.  
The only idea I have so far, is to add another field to the type with  
an 'asciifolding' analyzer, populate that field for all docs and  
switch the field in which the searches are performing to the new one.

Thanks for your great support,

Frederic

On 17 ene, 08:45, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> > I've been performing some searches (ES 0.18.5) using accents as well  
> > (in spanish) but, instead of using the 'asciifolding' filter at a  
> > indexing time, I'd want to get similar results using fuzzy queries  
> > (apart from getting also results for similar words).
> 
> > I want to search docs using one or more words, based on a free text  
> > field, called 'title', and I'd like to get the same results for both,  
> > for instance, "bateria" and "batería" words. The query is:
> 
> > ```
> > "query" : {
> > "text" : {
> > "title" : {
> > "query" : "batería",
> > "type" : "boolean",
> > "operator" : "AND",
> > "fuzziness" : "0.7",
> > "max_expansions" : 3
> > }
> > }
> > }
> > 
> > ```
> 
> > AFAIK a word with one accented letter is at distance '1' from the same  
> > word with no accent, is this correct? If so, the query should consider  
> > all docs that contains, in this case "batería" and "batería", right?
> 
> If you change the "fuzziness" factor to 0.5, it will probably work. I  
> don't understand exactly what that number represents, so can't give you  
> more than a trial-and-error approach 🙂
> 
> That said, using a fuzzy query for this type of search is a lot heavier  
> than analyzing your text properly at index time.
> 
> clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [January 17, 2012, 4:26pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/9 "2012-01-17T16:26:06Z")

</div>

> Totally agree, it's just that in my case I need to work in an already  
> productive system with 50M docs indexed, so I cannot recreate the  
> index for changing the 'title' field analyzer.  
> The only idea I have so far, is to add another field to the type with  
> an 'asciifolding' analyzer, populate that field for all docs and  
> switch the field in which the searches are performing to the new one.

You may want to take a look at multi-fields:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

clint

---

<div class="post-metadata">

**Author:** ![Frederic](https://avatars.discourse-cdn.com/v4/letter/f/4491bb/32.png) [@Frederic](https://discuss.elastic.co/u/Frederic)\
**Post date:** [January 17, 2012, 6:37pm UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/10 "2012-01-17T18:37:51Z")

</div>

That's exactly what I need. Thanks a lot

Fred  
On 17 ene, 13:26, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> > Totally agree, it's just that in my case I need to work in an already  
> > productive system with 50M docs indexed, so I cannot recreate the  
> > index for changing the 'title' field analyzer.  
> > The only idea I have so far, is to add another field to the type with  
> > an 'asciifolding' analyzer, populate that field for all docs and  
> > switch the field in which the searches are performing to the new one.
> 
> You may want to take a look at multi-fields:
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/multi-field-type)...
> 
> clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:42am UTC](https://discuss.elastic.co/t/problem-searching-queries-with-accents/6401/11 "2017-07-06T03:42:26Z")

</div>


