# Ask for suggestion on what analyzer to use

**URL:** <https://discuss.elastic.co/t/ask-for-suggestion-on-what-analyzer-to-use/22038>\
**Category:** Elasticsearch\
**Created:** [February 9, 2015, 5:37am UTC](https://discuss.elastic.co/t/ask-for-suggestion-on-what-analyzer-to-use/22038 "2015-02-09T05:37:46Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Youxu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/youxu/32/117375_2.png) [@Youxu](https://discuss.elastic.co/u/Youxu)\
**Post date:** [February 9, 2015, 5:37am UTC](https://discuss.elastic.co/t/ask-for-suggestion-on-what-analyzer-to-use/22038/1 "2015-02-09T05:37:46Z")

</div>

I have document like this:

{"Title":"This is my test title"}

I want to return the doc with exactly matched title, but allowing case  
insensitive and redundant white space between words. That is, all of these  
queries: "this is my test TITLE", "this is my test title" and "this is  
my test title" will match the doc.

My initial idea is to define a custom analyzer with keyword tonenizer and  
lowercase filter:  
"settings" : {  
"analysis" : {  
"analyzer" : {  
"lowercase\_keyword" : {  
"type" : "custom",  
"tokenizer" : "keyword",  
"filter" : "lowercase"  
}  
}  
}  
}

and use lowercase\_keyword analyzer for title filed:  
"properties" : {  
"title" : {  
"type" : "string",  
"analyzer" : "lowercase\_keyword"  
}  
}

It works well for both "this is my test TITLE" and "this is my test title".  
But does not work for redundant whitespaces case like "this is my  
test title".

How can I define my cusom analyzer to achieve my goal?

Thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/fb146ea8-30d6-46c4-8f5b-d7412249a900%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/fb146ea8-30d6-46c4-8f5b-d7412249a900%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 9, 2015, 5:54am UTC](https://discuss.elastic.co/t/ask-for-suggestion-on-what-analyzer-to-use/22038/2 "2015-02-09T05:54:20Z")

</div>

May be keep the default analyzer but use a phrase search?

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

--  
David 😉  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

> Le 9 févr. 2015 à 06:37, Xudong You [xudong.you@gmail.com](mailto:xudong.you@gmail.com) a écrit :
> 
> I have document like this:
> 
> {"Title":"This is my test title"}
> 
> I want to return the doc with exactly matched title, but allowing case insensitive and redundant white space between words. That is, all of these queries: "this is my test TITLE", "this is my test title" and "this is my test title" will match the doc.
> 
> My initial idea is to define a custom analyzer with keyword tonenizer and lowercase filter:  
> "settings" : {  
> "analysis" : {  
> "analyzer" : {  
> "lowercase\_keyword" : {  
> "type" : "custom",  
> "tokenizer" : "keyword",  
> "filter" : "lowercase"  
> }  
> }  
> }  
> }
> 
> and use lowercase\_keyword analyzer for title filed:  
> "properties" : {  
> "title" : {  
> "type" : "string",  
> "analyzer" : "lowercase\_keyword"  
> }  
> }
> 
> It works well for both "this is my test TITLE" and "this is my test title". But does not work for redundant whitespaces case like "this is my test title".
> 
> How can I define my cusom analyzer to achieve my goal?
> 
> ## Thanks!
> 
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/fb146ea8-30d6-46c4-8f5b-d7412249a900%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/fb146ea8-30d6-46c4-8f5b-d7412249a900%40googlegroups.com).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/237D6305-79CC-4147-B11C-EF7FBE3C1A78%40pilato.fr](https://groups.google.com/d/msgid/elasticsearch/237D6305-79CC-4147-B11C-EF7FBE3C1A78%40pilato.fr).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Youxu1](https://avatars.discourse-cdn.com/v4/letter/y/b782af/32.png) [@Youxu1](https://discuss.elastic.co/u/Youxu1)\
**Post date:** [February 9, 2015, 6:10am UTC](https://discuss.elastic.co/t/ask-for-suggestion-on-what-analyzer-to-use/22038/3 "2015-02-09T06:10:17Z")

</div>

Thanks reply!  
Yes match\_phrase with default analyzer works somehow.  
But I would like to optimize it with a better solution that

1. Just index the title field with one single token, instead of multiple tokens with standard analyzer.
2. only match the doc when input query contains all words of title, that is, if search "this is my test", the doc won't match.

---

<div class="post-metadata">

**Author:** ![Youxu1](https://avatars.discourse-cdn.com/v4/letter/y/b782af/32.png) [@Youxu1](https://discuss.elastic.co/u/Youxu1)\
**Post date:** [February 10, 2015, 2:54am UTC](https://discuss.elastic.co/t/ask-for-suggestion-on-what-analyzer-to-use/22038/4 "2015-02-10T02:54:21Z")

</div>

Any one has good suggestions?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:34am UTC](https://discuss.elastic.co/t/ask-for-suggestion-on-what-analyzer-to-use/22038/5 "2017-07-06T00:34:02Z")

</div>


