# Completion Suggester and Analyzer

**URL:** <https://discuss.elastic.co/t/completion-suggester-and-analyzer/13867>\
**Category:** Elasticsearch\
**Created:** [October 7, 2013, 1:23pm UTC](https://discuss.elastic.co/t/completion-suggester-and-analyzer/13867 "2013-10-07T13:23:39Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pawel\_Mlynarczyk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pawel_mlynarczyk/32/1568_2.png) [@Pawel\_Mlynarczyk](https://discuss.elastic.co/u/Pawel_Mlynarczyk)\
**Post date:** [October 7, 2013, 1:23pm UTC](https://discuss.elastic.co/t/completion-suggester-and-analyzer/13867/1 "2013-10-07T13:23:39Z")

</div>

Hello

I'm trying out the new completion suggester feature.  
I'm using simple analyzer to analyze at both index and search time. I have  
"nirvana nevermind" as input for completion and still starting completion  
term with "never" does not return anything. I've expected this to work  
since analyzer splits "nirvana nevermind" into two separate tokens?

I'm using example data from elasticsearch website:

curl -X PUT localhost:9200/music  
curl -X PUT localhost:9200/music/song/\_mapping -d '{  
"song" : {  
"properties" : {  
"name" : { "type" : "string" },  
"suggest" : { "type" : "completion",  
"index\_analyzer" : "simple",  
"search\_analyzer" : "simple",  
"payloads" : true  
}  
}  
}  
}'

but I've changed the indexed item a bit:

curl -X PUT 'localhost:9200/music/song/1?refresh=true' -d '{  
"name" : "Nevermind",  
"suggest" : {  
"input": ["nirvana nevermind"],  
"output": "Nirvana - Nevermind"  
}  
}'

And this query doesn't return anything:

curl -X POST 'localhost:9200/music/\_suggest?pretty' -d '{  
"song-suggest" : {  
"text" : "never",  
"completion" : {  
"field" : "suggest"  
}  
}  
}'

I know I can handle this by just adding more inputs, but I am concerned  
about the size of the index, when the list of possible user inputs for an  
item goes huge...

Is there a way to analyze terms to match my expectations?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [October 7, 2013, 2:12pm UTC](https://discuss.elastic.co/t/completion-suggester-and-analyzer/13867/2 "2013-10-07T14:12:47Z")

</div>

Hey Pawel,

right now the suggester is a pure prefix suggester, this means the term you  
indexed was "Nirvana - Nevermind", so you only get suggestions back, when  
you enter "Nirv". So as a workaround you could index several inputs like  
"Nirvana" and "Nevermind". So

"input": ["Nirvana", "Nevermind"],  
"output" : "Nirvana - Nevermind"

would make your usecase work. Also in case you are afraid of the size, you  
can easily monitor by field using the nodes stats API, see more at:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

From a long term point of view it makes sense to support the  
AnalyzingInfixSuggester from Lucene as well.

Hope this helps.

--Alex

On Mon, Oct 7, 2013 at 3:23 PM, Paweł Młynarczyk [zwarios@gmail.com](mailto:zwarios@gmail.com) wrote:

> Hello
> 
> I'm trying out the new completion suggester feature.  
> I'm using simple analyzer to analyze at both index and search time. I have  
> "nirvana nevermind" as input for completion and still starting completion  
> term with "never" does not return anything. I've expected this to work  
> since analyzer splits "nirvana nevermind" into two separate tokens?
> 
> I'm using example data from elasticsearch website:
> 
> curl -X PUT localhost:9200/music  
> curl -X PUT localhost:9200/music/song/\_mapping -d '{  
> "song" : {  
> "properties" : {  
> "name" : { "type" : "string" },  
> "suggest" : { "type" : "completion",  
> "index\_analyzer" : "simple",  
> "search\_analyzer" : "simple",  
> "payloads" : true  
> }  
> }  
> }  
> }'
> 
> but I've changed the indexed item a bit:
> 
> curl -X PUT 'localhost:9200/music/song/1?refresh=true' -d '{  
> "name" : "Nevermind",  
> "suggest" : {  
> "input": ["nirvana nevermind"],  
> "output": "Nirvana - Nevermind"  
> }  
> }'
> 
> And this query doesn't return anything:
> 
> curl -X POST 'localhost:9200/music/\_suggest?pretty' -d '{  
> "song-suggest" : {  
> "text" : "never",  
> "completion" : {  
> "field" : "suggest"  
> }  
> }  
> }'
> 
> I know I can handle this by just adding more inputs, but I am concerned  
> about the size of the index, when the list of possible user inputs for an  
> item goes huge...
> 
> Is there a way to analyze terms to match my expectations?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [October 7, 2013, 2:13pm UTC](https://discuss.elastic.co/t/completion-suggester-and-analyzer/13867/3 "2013-10-07T14:13:18Z")

</div>

This suggester is in-fact a prefix suggester. it will only operate on the  
prefixes you are adding and it will complete them.  
You said you are afraid of the size of the index - I can assure you this  
one takes you extremely far without being an issue. The compression for  
this kind stuff is immense and I have personal experience with the exact  
same problems. Don't worry too much about the index size here unless you  
have tens of billions of records with many different prefixes. If you have  
stuff like \<song\_name\> you can easily have the combinations  
[-\<song\_name\>, \<song\_name\>, \<song\_name\>-] without issues.  
We are working on solutions that help with these situations but they won't  
use less space.

simon

On Monday, October 7, 2013 3:23:39 PM UTC+2, Paweł Młynarczyk wrote:

> Hello
> 
> I'm trying out the new completion suggester feature.  
> I'm using simple analyzer to analyze at both index and search time. I have  
> "nirvana nevermind" as input for completion and still starting completion  
> term with "never" does not return anything. I've expected this to work  
> since analyzer splits "nirvana nevermind" into two separate tokens?
> 
> I'm using example data from elasticsearch website:
> 
> curl -X PUT localhost:9200/music  
> curl -X PUT localhost:9200/music/song/\_mapping -d '{  
> "song" : {  
> "properties" : {  
> "name" : { "type" : "string" },  
> "suggest" : { "type" : "completion",  
> "index\_analyzer" : "simple",  
> "search\_analyzer" : "simple",  
> "payloads" : true  
> }  
> }  
> }  
> }'
> 
> but I've changed the indexed item a bit:
> 
> curl -X PUT 'localhost:9200/music/song/1?refresh=true' -d '{  
> "name" : "Nevermind",  
> "suggest" : {  
> "input": ["nirvana nevermind"],  
> "output": "Nirvana - Nevermind"  
> }  
> }'
> 
> And this query doesn't return anything:
> 
> curl -X POST 'localhost:9200/music/\_suggest?pretty' -d '{  
> "song-suggest" : {  
> "text" : "never",  
> "completion" : {  
> "field" : "suggest"  
> }  
> }  
> }'
> 
> I know I can handle this by just adding more inputs, but I am concerned  
> about the size of the index, when the list of possible user inputs for an  
> item goes huge...
> 
> Is there a way to analyze terms to match my expectations?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:13am UTC](https://discuss.elastic.co/t/completion-suggester-and-analyzer/13867/4 "2017-07-06T02:13:25Z")

</div>


