# Keyword tokenizer

**URL:** <https://discuss.elastic.co/t/keyword-tokenizer/15355>\
**Category:** Elasticsearch\
**Created:** [January 22, 2014, 11:17am UTC](https://discuss.elastic.co/t/keyword-tokenizer/15355 "2014-01-22T11:17:12Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![paul1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul1/32/46819_2.png) [@paul1](https://discuss.elastic.co/u/paul1)\
**Post date:** [January 22, 2014, 11:17am UTC](https://discuss.elastic.co/t/keyword-tokenizer/15355/1 "2014-01-22T11:17:12Z")

</div>

My mapping looks as below

"autocomplete\_index":{  
"type":"custom",  
"tokenizer":"keyword",  
"filter":[  
"lowercase",  
"syns\_filter",  
"my\_edgeNgram"  
]  
}

Now when i analyze the configuration using analyze api the word after space  
gets omitted . ie "university" is omitted

................../universityindextest2/\_analyze?analyzer=autocomplete\_index&text=yale%20university&pretty

## output

{ "tokens" : [{ "token" : "ya", "start\_offset" : 0, "end\_offset" : 15,"type" : "word","position" : 1}, {"token" : "yal","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 2}, {"token" : "yale","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 3}, {"token" : "yu","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 4}]  
}

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/d6bd7caa-b160-42ac-948c-6aab6884a51d%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/d6bd7caa-b160-42ac-948c-6aab6884a51d%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Binh\_Ly](https://avatars.discourse-cdn.com/v4/letter/b/ce7236/32.png) [@Binh\_Ly](https://discuss.elastic.co/u/Binh_Ly)\
**Post date:** [January 22, 2014, 5:30pm UTC](https://discuss.elastic.co/t/keyword-tokenizer/15355/2 "2014-01-22T17:30:31Z")

</div>

Paul, Is it possible that your "syns\_filter" is affecting your ngram  
filter? What happens when you remove the syns\_filter?

On Wednesday, January 22, 2014 6:17:12 AM UTC-5, paul wrote:

> My mapping looks as below
> 
> "autocomplete\_index":{  
> "type":"custom",  
> "tokenizer":"keyword",  
> "filter":[  
> "lowercase",  
> "syns\_filter",  
> "my\_edgeNgram"  
> ]  
> }
> 
> Now when i analyze the configuration using analyze api the word after  
> space gets omitted . ie "university" is omitted
> 
> ................../universityindextest2/\_analyze?analyzer=autocomplete\_index&text=yale%20university&pretty
> 
> ## output
> 
> { "tokens" : [{ "token" : "ya", "start\_offset" : 0, "end\_offset" : 15,"type" : "word","position" : 1}, {"token" : "yal","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 2}, {"token" : "yale","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 3}, {"token" : "yu","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 4}]  
> }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/423e6c0f-0aa2-4f48-a357-a313905fb8c0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/423e6c0f-0aa2-4f48-a357-a313905fb8c0%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![paul1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul1/32/46819_2.png) [@paul1](https://discuss.elastic.co/u/paul1)\
**Post date:** [January 24, 2014, 5:03am UTC](https://discuss.elastic.co/t/keyword-tokenizer/15355/3 "2014-01-24T05:03:34Z")

</div>

Binh , When i removed the syns\_filter its still the same but when i changed  
the "tokenizer":"keyword", to "whitespcae" it taking "university"  
into account. May be its a tokenizer problem , when there is a space the  
keyword tokenizer is omitting the word after space.

-paul

On Wed, Jan 22, 2014 at 11:00 PM, Binh Ly [binh@hibalo.com](mailto:binh@hibalo.com) wrote:

> Paul, Is it possible that your "syns\_filter" is affecting your ngram  
> filter? What happens when you remove the syns\_filter?
> 
> On Wednesday, January 22, 2014 6:17:12 AM UTC-5, paul wrote:
> 
> > My mapping looks as below
> > 
> > "autocomplete\_index":{  
> > "type":"custom",  
> > "tokenizer":"keyword",  
> > "filter":[  
> > "lowercase",  
> > "syns\_filter",  
> > "my\_edgeNgram"  
> > ]  
> > }
> > 
> > Now when i analyze the configuration using analyze api the word after  
> > space gets omitted . ie "university" is omitted
> > 
> > ................../universityindextest2/\_analyze?  
> > analyzer=autocomplete\_index&text=yale%20university&pretty
> > 
> > ## output
> > 
> > { "tokens" : [{ "token" : "ya", "start\_offset" : 0, "end\_offset" : 15,"type" : "word","position" : 1}, {"token" : "yal","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 2}, {"token" : "yale","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 3}, {"token" : "yu","start\_offset" : 0,"end\_offset" : 15,"type" : "word","position" : 4}]  
> > }
> > 
> > --  
> > You received this message because you are subscribed to a topic in the  
> > Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit  
> > [https://groups.google.com/d/topic/elasticsearch/inRyvJJDPpo/unsubscribe](https://groups.google.com/d/topic/elasticsearch/inRyvJJDPpo/unsubscribe).  
> > To unsubscribe from this group and all its topics, send an email to  
> > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/423e6c0f-0aa2-4f48-a357-a313905fb8c0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/423e6c0f-0aa2-4f48-a357-a313905fb8c0%40googlegroups.com)  
> > .
> 
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAO066G0Y%2BAoVt%2BN6q1bxr8KFN2A686U2Cp%3DyyEoHT\_s41\_vbzg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAO066G0Y%2BAoVt%2BN6q1bxr8KFN2A686U2Cp%3DyyEoHT_s41_vbzg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Binh\_Ly](https://avatars.discourse-cdn.com/v4/letter/b/ce7236/32.png) [@Binh\_Ly](https://discuss.elastic.co/u/Binh_Ly)\
**Post date:** [January 24, 2014, 4:57pm UTC](https://discuss.elastic.co/t/keyword-tokenizer/15355/4 "2014-01-24T16:57:02Z")

</div>

Paul, yes you are correct, I missed that. The keyword tokenizer will take  
your entire string and make it into a single token - that's why it is not  
ngramming "university".

On Friday, January 24, 2014 12:03:34 AM UTC-5, paul wrote:

> Binh , When i removed the syns\_filter its still the same but when i  
> changed the "tokenizer":"keyword", to "whitespcae" it taking  
> "university" into account. May be its a tokenizer problem , when there is a  
> space the keyword tokenizer is omitting the word after space.
> 
> -paul
> 
> On Wed, Jan 22, 2014 at 11:00 PM, Binh Ly \<[bi...@hibalo.com](mailto:bi...@hibalo.com) \<javascript:\>\>wrote:
> 
> > Paul, Is it possible that your "syns\_filter" is affecting your ngram  
> > filter? What happens when you remove the syns\_filter?
> > 
> > On Wednesday, January 22, 2014 6:17:12 AM UTC-5, paul wrote:
> > 
> > > My mapping looks as below
> > > 
> > > "autocomplete\_index":{  
> > > "type":"custom",  
> > > "tokenizer":"keyword",  
> > > "filter":[  
> > > "lowercase",  
> > > "syns\_filter",  
> > > "my\_edgeNgram"  
> > > ]  
> > > }
> > > 
> > > Now when i analyze the configuration using analyze api the word after  
> > > space gets omitted . ie "university" is omitted
> > > 
> > > ................../universityindextest2/\_analyze?  
> > > analyzer=autocomplete\_index&text=yale%20university&pretty
> > > 
> > > ## output
> > > 
> > > { "tokens" : [ { "token" : "ya", "start\_offset" : 0, "end\_offset" : 15,"type" : "word", "position"  
> > > : 1 }, { "token" : "yal", "start\_offset" : 0, "end\_offset" : 15, "type"  
> > > : "word", "position" : 2 }, { "token" : "yale", "start\_offset" : 0,"end\_offset" : 15, "type"  
> > > : "word", "position" : 3 }, { "token" : "yu", "start\_offset" : 0,"end\_offset" : 15,"type" : "word", "position"  
> > > : 4 } ]}
> > > 
> > > --  
> > > You received this message because you are subscribed to a topic in the  
> > > Google Groups "elasticsearch" group.  
> > > To unsubscribe from this topic, visit  
> > > [https://groups.google.com/d/topic/elasticsearch/inRyvJJDPpo/unsubscribe](https://groups.google.com/d/topic/elasticsearch/inRyvJJDPpo/unsubscribe).  
> > > To unsubscribe from this group and all its topics, send an email to  
> > > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > To view this discussion on the web visit  
> > > [https://groups.google.com/d/msgid/elasticsearch/423e6c0f-0aa2-4f48-a357-a313905fb8c0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/423e6c0f-0aa2-4f48-a357-a313905fb8c0%40googlegroups.com)  
> > > .
> > 
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/0bc9516f-1830-4f70-a25b-276a9b43ddac%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/0bc9516f-1830-4f70-a25b-276a9b43ddac%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:54am UTC](https://discuss.elastic.co/t/keyword-tokenizer/15355/5 "2017-07-06T01:54:56Z")

</div>


