# Advice on my approach to this search problem

**URL:** <https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606>\
**Category:** Elasticsearch\
**Created:** [October 15, 2011, 3:20pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606 "2011-10-15T15:20:25Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 15, 2011, 3:20pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/1 "2011-10-15T15:20:25Z")

</div>

Hey guys. I've been using ES for 1-2 weeks now, and love it. Being very new  
to it, though, I've been piecing together bits of knowledge as I go along. I  
have a semi-working solution to a problem, but I'm sure that it's nowhere  
close to an ideal solution. How would you approach this problem?

First, the data: The app has catalogs, products, and items, all of which are  
related. When a user performs a search, though, only products need to be  
found. For example, if there's a catalog named "Transformers" and the user  
searches for "Transformers", all products in that catalog should be  
returned. To accomplish this, when indexing products, I'm nesting the  
related catalog and item data inside the product data. Eg:  
{ name: "Optimus Prime", number: "TFG1S1-1",  
catalog: { name: "Series 1", number: "TFG1S1" },  
items: [{ name: "Optimus Prime" }, { name: "Instructions" }]  
}

When searching, some users will misspell a word (Eg: "Trasformers", missing  
the "n"), and some will provide a partial word (Eg: "Trans"). Despite this,  
users expect to receive search results. Eg: Products with a field that  
contains "Transformer" or "retransmit" could match.

My solution right now is this:

> <https://gist.github.com/nickhoffman/0cbe6892b4bc720bda92>
>
> There are more than three files. show original

Do you have any suggestions for improvements? Thanks for your help!  
Nick

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 17, 2011, 2:34pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/2 "2011-10-17T14:34:36Z")

</div>

Bump!

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [October 17, 2011, 2:56pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/3 "2011-10-17T14:56:54Z")

</div>

spell check can be done via a dictionary

> <https://github.com/elastic/elasticsearch/issues/646>
>
> Feature added by @uboness:
> 
> Added basic support for hunspell stemming. Hunspell …dictionaries will be picked up from a dedicated hunspell directory on the filesystem (defaults to \_\<path.conf\>\_/hunspell). Each dictionary is expected to have its own directory named after its associated locale (language). This dictionary directory is expected to hold both the \*.aff and \*.dic files (all of which will automatically be picked up). For example, assuming the default hunspell location is used, the following directory layout will define the \_en\_US\_ dictionary:
> 
> \`\`\`
> \- conf
> |-- hunspell 
> | |-- en\_US
> | | |-- en\_US.dic
> | | |-- en\_US.aff
> \`\`\`
> 
> The location of the hunspell directory can be configured using the \`indices.analysis.hunspell.dictionary.location\` settings in \_elasticsearch.yml\_.
> 
> Each dictionary can be configured with two settings:
> \- \`ignore\_case\` - If true, dictionary matching will be case insensitive (defaults to \`false\`)
> \- \`strict\_affix\_parsing\` - Determines whether errors while reading a affix rules file will cause exception or simple be ignored (defaults to \`true\`)
> 
> These settings can be configured globally in \`elasticsearch.yml\` using \`indices.analysis.hunspell.dictionary.ignore\_case\` and \`indices.analysis.hunspell.dictionary.strict\_affix\_parsing\`, or for specific dictionaries: \`indices.analysis.hunspell.dictionary.en\_US.ignore\_case\` and \`indices.analysis.hunspell.dictionary.en\_US.strict\_affix\_parsing\`.
> 
> It is also possible to add \`settings.yml\` file under the dictionary directory which holds these settings (this will override any other settings defined in the \`elasticsearch.yml\`).
> 
> One can use the hunspell stem filter by configuring it the analysis settings:
> 
> \`\`\` json
> {
> "analysis" : {
> "analyzer" : {
> "en" : {
> "tokenizer" : "standard",       
> "filter" : \["lowercase", "en\_US" \]
> }
> },
> "filter" : {
> "en\_US" : {
> "type" : "hunspell",
> "locale" : "en\_US",
> "dedup" : true
> }
> }
> }
> }
> \`\`\`
> \## Original Request:
> 
> Hunspell is a spell checker and morphological analyzer designed for languages with rich morphology and complex word compounding and character encoding.
> 
> \[1\] Wikipedia, http://en.wikipedia.org/wiki/Hunspell
> 
> \[2\] Source code, http://hunspell.sourceforge.net/
> 
> \[3\] Hunspell-Lucene integration, http://code.google.com/p/lucene-hunspell/
> 
> \[4\] presentation by Chris Male, EuroCon 2010, http://lucene-eurocon.org/slides/European-Language-Analysis-with-Hunspell\_Chris-Male.pdf (annotation of his talk can be found here: http://lucene-eurocon.org/sessions-track2-day2.html#5)

,lucene 4

> <https://github.com/elastic/elasticsearch/issues/911>
>
> Google's "Did you mean" feature is very useful. Would be awesome if ES could imp…lement this.
> 
> Lucene has pulled in the \<a href="http://lucene.apache.org/java/3\_1\_0/api/all/org/apache/lucene/search/spell/SpellChecker.html"\>SpellChecker contrib\</a\>. Maybe ES could expose that?
> 
> Ex. if I specify suggestSimilar with some optional parameters in my search object I could get back an array with some suggestions.

or with a phonetic analyzer

> <https://stackoverflow.com/questions/6936256/elastic-search-implement-did-you-mean>

is that what you were after?

On 17 Okt., 16:34, Nick Hoffman [n...@deadorange.com](mailto:n...@deadorange.com) wrote:

> Bump!

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 17, 2011, 3:07pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/4 "2011-10-17T15:07:33Z")

</div>

On Sat, 2011-10-15 at 08:20 -0700, Nick Hoffman wrote:

> Hey guys. I've been using ES for 1-2 weeks now, and love it. Being  
> very new to it, though, I've been piecing together bits of knowledge  
> as I go along. I have a semi-working solution to a problem, but I'm  
> sure that it's nowhere close to an ideal solution. How would you  
> approach this problem?

Hi Nick

The reason your misspellings work is because you are using ngrams for  
both your search and index analyzers.

This may, however, give your users weird results, eg the user searches  
for "slave" and gets a result for "lavatory" instead.

I would consider making a few changes:

1. use edge ngrams rather than ngrams ie s,sl,sla,slav,slave
2. use the edge ngram analyzers only as your search\_analyzer
3. for your misspellings, if you get no results, then retry  
the query using some fuzziness:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/text-query.html)

clint

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 17, 2011, 10:00pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/5 "2011-10-17T22:00:40Z")

</div>

On Monday, 17 October 2011 10:56:54 UTC-4, Karussell wrote:

> spell check can be done via a dictionary
> 
> [Analysis: Integration with Hunspell · Issue #646 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/646)
> 
> ,lucene 4
> 
> ["Did you mean" spellchecking · Issue #911 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/911)
> 
> or with a phonetic analyzer
> 
> [ruby on rails - Elastic Search - implement "Did you Mean" - Stack Overflow](http://stackoverflow.com/questions/6936256/elastic-search-implement-did-you-mean)
> 
> is that what you were after?

Thanks for the suggestions, mate. Unfortunately, I can't do spell checking  
with a dictionary because many of the words are unique names. Eg:  
Optimus Prime  
Megatron  
BE@RBRICK  
etc

I was going to try a phonetic analyzer, but a lot of the names in my data  
are pronounced strangely, and thus wouldn't match.

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 17, 2011, 10:07pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/6 "2011-10-17T22:07:45Z")

</div>

On Monday, 17 October 2011 11:07:33 UTC-4, Clinton Gormley wrote:

> The reason your misspellings work is because you are using ngrams for  
> both your search and index analyzers.

Yeah, I figured as much.

> This may, however, give your users weird results, eg the user searches  
> for "slave" and gets a result for "lavatory" instead.
> 
> I would consider making a few changes:
> 
> 1. use edge ngrams rather than ngrams ie s,sl,sla,slav,slave

Interesting. Why do you recommend that? I understand that it prevents the  
slave/lavatory example, which is great. However, it prevents mid-word  
matches. But then again, maybe that's a good thing...

> 1. use the edge ngram analyzers only as your search\_analyzer

So don't use them in any of the index analyzers?

> 1. for your misspellings, if you get no results, then retry  
> the query using some fuzziness:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/text-query.html)

A text query with the "fuzziness" option, or a fuzzy query[1]?

Thanks for your advice, Clint, and also for that example nGram gist. Very  
helpful!

[1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/fuzzy-query.html)

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 18, 2011, 7:27am UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/7 "2011-10-18T07:27:20Z")

</div>

Hiya

> ```
> This may, however, give your users weird results, eg the user
> searches
> for "slave" and gets a result for "lavatory" instead.
>     
> I would consider making a few changes:
> 1) use edge ngrams rather than ngrams ie s,sl,sla,slav,slave
> 
> ```
> 
> Interesting. Why do you recommend that? I understand that it prevents  
> the slave/lavatory example, which is great. However, it prevents  
> mid-word matches. But then again, maybe that's a good thing...

Yes exactly. Full ngrams are useful for some purposes, eg matching  
words in a URL, but in general, people start typing at one end of a word  
and expect the search results to reflect that.

> ```
> 2) use the edge ngram analyzers only as your search_analyzer
> 
> ```

> So don't use them in any of the index analyzers?

Apologies - I meant the other way around. Use them in your index  
analyzers, but use your ascii\_std analyzer for search analyzers.

> ```
> 3) for your misspellings, if you get no results, then retry
> the query using some fuzziness:  
>     
> http://www.elasticsearch.org/guide/reference/query-dsl/text-query.html
> 
> ```
> 
> A text query with the "fuzziness" option, or a fuzzy query[1]?

text with fuzziness. A fuzzy query is actually a term query - the  
search terms are not analyzed. However, a text query with fuzziness  
gives you the analysis plus the fuzzy behaviour.

> Thanks for your advice, Clint, and also for that example nGram gist.  
> Very helpful!

glad to hear it 🙂

clint

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 19, 2011, 3:48am UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/8 "2011-10-19T03:48:03Z")

</div>

> Apologies - I meant the other way around. Use them in your index  
> analyzers, but use your ascii\_std analyzer for search analyzers.

Thanks, Clint. That's working a lot better. However, I've noticed that I  
can't combine certain fields in the same text query. It seems to be related  
to nested fields. Any idea why that might be?

For example, ES accepts this:

curl -X GET -s  
"[http://localhost:9200/development\_products/product/\_search?pretty=true](http://localhost:9200/development_products/product/_search?pretty=true)" -d  
'{ query: { text: { "items.name": "optimus", "catalog.name": "optimus" } }  
}'

But adding the "name" field to the beginning or end:

curl -X GET -s  
"[http://localhost:9200/development\_products/product/\_search?pretty=true](http://localhost:9200/development_products/product/_search?pretty=true)" -d  
'{ query: { text: { "items.name": "optimus", "catalog.name": "optimus",  
"name": "optimus" } } }'

generates an error:

{  
"error" : "SearchPhaseExecutionException[Failed to execute phase [query],  
total failure; shardFailures  
{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][0]:  
SearchParseException[[development\_products][0]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to parse  
source [{ query: { text: { "items.name": "optimus", "catalog.name":  
"optimus", "name": "optimus" } } }]]]; nested:  
SearchParseException[[development\_products][0]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser for  
element [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][2]:  
SearchParseException[[development\_products][2]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to parse  
source [{ query: { text: { "items.name": "optimus", "catalog.name":  
"optimus", "name": "optimus" } } }]]]; nested:  
SearchParseException[[development\_products][2]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser for  
element [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][1]:  
SearchParseException[[development\_products][1]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to parse  
source [{ query: { text: { "items.name": "optimus", "catalog.name":  
"optimus", "name": "optimus" } } }]]]; nested:  
SearchParseException[[development\_products][1]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser for  
element [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][3]:  
SearchParseException[[development\_products][3]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to parse  
source [{ query: { text: { "items.name": "optimus", "catalog.name":  
"optimus", "name": "optimus" } } }]]]; nested:  
SearchParseException[[development\_products][3]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser for  
element [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][4]:  
SearchParseException[[development\_products][4]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to parse  
source [{ query: { text: { "items.name": "optimus", "catalog.name":  
"optimus", "name": "optimus" } } }]]]; nested:  
SearchParseException[[development\_products][4]:  
query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser for  
element [name]]]; }]",  
"status" : 500  
}

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 19, 2011, 10:04am UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/9 "2011-10-19T10:04:27Z")

</div>

> Thanks, Clint. That's working a lot better. However, I've noticed that  
> I can't combine certain fields in the same text query. It seems to be  
> related to nested fields. Any idea why that might be?

You can't pass multiple field/search\_text pairs to a single query. ES  
needs to know how to combine your various queries, so you need to have  
each as a separate 'text' query and combine them using either bool or  
dismax

clint

> For example, ES accepts this:
> 
> curl -X GET -s  
> "[http://localhost:9200/development\_products/product/\_search?pretty=true](http://localhost:9200/development_products/product/_search?pretty=true)" -d '{ query: { text: { "items.name": "optimus", "catalog.name": "optimus" } } }'
> 
> But adding the "name" field to the beginning or end:
> 
> curl -X GET -s  
> "[http://localhost:9200/development\_products/product/\_search?pretty=true](http://localhost:9200/development_products/product/_search?pretty=true)" -d '{ query: { text: { "items.name": "optimus", "catalog.name": "optimus", "name": "optimus" } } }'
> 
> generates an error:
> 
> {  
> "error" : "SearchPhaseExecutionException[Failed to execute phase  
> [query], total failure; shardFailures  
> {[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][0]:  
> SearchParseException[[development\_products][0]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to  
> parse source [{ query: { text: { "items.name": "optimus",  
> "catalog.name": "optimus", "name": "optimus" } } }]]]; nested:  
> SearchParseException[[development\_products][0]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser  
> for element  
> [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][2]:  
> SearchParseException[[development\_products][2]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to  
> parse source [{ query: { text: { "items.name": "optimus",  
> "catalog.name": "optimus", "name": "optimus" } } }]]]; nested:  
> SearchParseException[[development\_products][2]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser  
> for element  
> [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][1]:  
> SearchParseException[[development\_products][1]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to  
> parse source [{ query: { text: { "items.name": "optimus",  
> "catalog.name": "optimus", "name": "optimus" } } }]]]; nested:  
> SearchParseException[[development\_products][1]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser  
> for element  
> [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][3]:  
> SearchParseException[[development\_products][3]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to  
> parse source [{ query: { text: { "items.name": "optimus",  
> "catalog.name": "optimus", "name": "optimus" } } }]]]; nested:  
> SearchParseException[[development\_products][3]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser  
> for element  
> [name]]]; }{[V6gkYzvcSg6Gx-Ad9-hbOg][development\_products][4]:  
> SearchParseException[[development\_products][4]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [Failed to  
> parse source [{ query: { text: { "items.name": "optimus",  
> "catalog.name": "optimus", "name": "optimus" } } }]]]; nested:  
> SearchParseException[[development\_products][4]:  
> query[items.name:optimus],from[-1],size[-1]: Parse Failure [No parser  
> for element [name]]]; }]",  
> "status" : 500  
> }

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [October 19, 2011, 3:23pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/10 "2011-10-19T15:23:48Z")

</div>

That makes sense. Thanks. With dis\_max, though, it looks like the  
edge-ngrams aren't being used.

For example, there're 3 documents whose "name" field is "Optimus Primal",  
and 35 whose "name" field is "Optimus Prime". I figured that a dis\_max-text  
query for "primal" would match "Optimus Primal" and "Optimus Prime" docs.  
Unfortunately, only the 3 "Optimus Primal" docs matched. Why might that be?

curl -X GET -s  
"[http://localhost:9200/development\_products/product/\_search?pretty=true](http://localhost:9200/development_products/product/_search?pretty=true)" -d  
'  
{  
fields: ["name"],  
query: {  
dis\_max: {  
queries: [  
{ text: { "name" : "primal" } }  
]  
}  
}  
}  
'

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 19, 2011, 3:47pm UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/11 "2011-10-19T15:47:01Z")

</div>

> For example, there're 3 documents whose "name" field is "Optimus  
> Primal", and 35 whose "name" field is "Optimus Prime". I figured that  
> a dis\_max-text query for "primal" would match "Optimus Primal" and  
> "Optimus Prime" docs. Unfortunately, only the 3 "Optimus Primal" docs  
> matched. Why might that be?

Because you are no longer using ngrams on your search analyzer, so we're  
essentially doing a search for "primal\*"

Try the same thing but search for "prim" instead

clint

> curl -X GET -s  
> "[http://localhost:9200/development\_products/product/\_search?pretty=true](http://localhost:9200/development_products/product/_search?pretty=true)" -d '
> 
> {  
> fields: ["name"],  
> query: {  
> dis\_max: {  
> queries: [  
> { text: { "name" : "primal" } }  
> ]  
> }  
> }  
> }  
> '

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:51am UTC](https://discuss.elastic.co/t/advice-on-my-approach-to-this-search-problem/5606/12 "2017-07-06T03:51:24Z")

</div>


