# Searching for misspellings

**URL:** <https://discuss.elastic.co/t/searching-for-misspellings/5584>\
**Category:** Elasticsearch\
**Created:** [October 13, 2011, 5:01am UTC](https://discuss.elastic.co/t/searching-for-misspellings/5584 "2011-10-13T05:01:17Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 13, 2011, 5:01am UTC](https://discuss.elastic.co/t/searching-for-misspellings/5584/1 "2011-10-13T05:01:17Z")

</div>

Hey everyone. I'm trying to configure my index and mapping to return  
documents if the user misspells a word. Using Clinton Gormley's nGram  
example ([https://gist.github.com/961303](https://gist.github.com/961303)), I've created an index with a  
mapping on one field. Searching for correct spellings works, but not  
misspellings. Any idea what I'm doing wrong?

Here's a gist that can be run on the CLI easily:

> <https://gist.github.com/nickhoffman/1283380>

As far as I understand, when indexing the "optimus prime" document, "optimus  
prime" will be split into tokens "o", "op", ..., "opti", etc. Thus,  
shouldn't the "o" to "opti" tokens match "optius"?

If you have any advice, I'm all ears! Thanks,  
Nick

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [October 13, 2011, 7:37am UTC](https://discuss.elastic.co/t/searching-for-misspellings/5584/2 "2011-10-13T07:37:35Z")

</div>

for misspelling this is not the method of choice. opti just matches  
opti

You can create an phonetic analyzer:

> <https://stackoverflow.com/questions/6936256/elastic-search-implement-did-you-mean>

or wait for Lucene 4.0 (or use fuzzy queries now if not too many  
docs):

> <https://github.com/elastic/elasticsearch/issues/911>
>
> Google's "Did you mean" feature is very useful. Would be awesome if ES could imp…lement this.
> 
> Lucene has pulled in the \<a href="http://lucene.apache.org/java/3\_1\_0/api/all/org/apache/lucene/search/spell/SpellChecker.html"\>SpellChecker contrib\</a\>. Maybe ES could expose that?
> 
> Ex. if I specify suggestSimilar with some optional parameters in my search object I could get back an array with some suggestions.

On 13 Okt., 07:01, Nick Hoffman [n...@deadorange.com](mailto:n...@deadorange.com) wrote:

> Hey everyone. I'm trying to configure my index and mapping to return  
> documents if the user misspells a word. Using Clinton Gormley's nGram  
> example ([Ngram example · GitHub](https://gist.github.com/961303)), I've created an index with a  
> mapping on one field. Searching for correct spellings works, but not  
> misspellings. Any idea what I'm doing wrong?
> 
> Here's a gist that can be run on the CLI easily:[Why does a search for a misspelling ("optius" instead of "optimus") find no documents? · GitHub](https://gist.github.com/1283380)
> 
> As far as I understand, when indexing the "optimus prime" document, "optimus  
> prime" will be split into tokens "o", "op", ..., "opti", etc. Thus,  
> shouldn't the "o" to "opti" tokens match "optius"?
> 
> If you have any advice, I'm all ears! Thanks,  
> Nick

---

<div class="post-metadata">

**Author:** ![Jan\_Fiedler](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jan_fiedler/32/2518_2.png) [@Jan\_Fiedler](https://discuss.elastic.co/u/Jan_Fiedler)\
**Post date:** [October 13, 2011, 7:39am UTC](https://discuss.elastic.co/t/searching-for-misspellings/5584/3 "2011-10-13T07:39:43Z")

</div>

I think I can at least explain why you do not see the results you expect:  
From the gist it seems you are using an edge-ngram filter at indexing time  
but a pretty standard analyzer at query time. Lets look at what will happen  
for your example data:

Using the 'ascii\_edge\_ngram' at indexing time will index 'Optimus Prime'  
into something like: [o], [op], [opt], [opti], [optim], ...  
You can test this for yourself via the great analyzer API: curl -XGET  
'localhost:9200/test\_products/\_analyze?analyzer=ascii\_edge\_ngram&pretty=true'  
-d 'Optimus Prime'

At search time you are applying your 'ascii\_std' analyzer to a mispelled  
query like 'optius'. Using the analyzer API you can see this is broken down  
into a single token [optius]. You notice that there is not a single token in  
the indexed content that would match this. Therefore you will not have a  
match with the misspelled query.

I think you need to use the same (or at least similar) analyzers at indexing  
and search time. You may want to try an n-gram (not edge n-gram) and play  
with the side of your grams (ngrams of size 1 produce a lot of noise for  
this type of query).

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 13, 2011, 2:46pm UTC](https://discuss.elastic.co/t/searching-for-misspellings/5584/4 "2011-10-13T14:46:02Z")

</div>

I considered a phonetic analyzer, but many of the words and phrases that my  
app will be indexing are non-dictionary words with non-standard  
pronunciation.

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 13, 2011, 2:53pm UTC](https://discuss.elastic.co/t/searching-for-misspellings/5584/5 "2011-10-13T14:53:36Z")

</div>

Thanks for taking the time to examine the gist and write a detailed  
explanation, Jan. I really appreciate it. It was also very helpful. I  
changed the index and search analyzers from the 1-20-character front  
edgeNGram to a 3-8-character nGram, and my searches are returning results as  
I'd like.

I'm not sure if a 3-8-character nGram is optimal, though. Are there any  
recommendations for how to determine the min and max characters for an nGram  
filter?

Here's the latest version of the code. All of the searches return the  
expected results!

> <https://gist.github.com/nickhoffman/1283380/89b50d76bef849767a4d5980ebf042d1c309a2bf>

Thanks again

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:52am UTC](https://discuss.elastic.co/t/searching-for-misspellings/5584/6 "2017-07-06T03:52:05Z")

</div>


