# Is there a way to search terms lower cased?

**URL:** <https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056>\
**Category:** Elasticsearch\
**Created:** [June 30, 2010, 11:57am UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056 "2010-06-30T11:57:07Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![sezgin\_kucukkaraasla](https://avatars.discourse-cdn.com/v4/letter/s/9de0a6/32.png) [@sezgin\_kucukkaraasla](https://discuss.elastic.co/u/sezgin_kucukkaraasla)\
**Post date:** [June 30, 2010, 11:57am UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/1 "2010-06-30T11:57:07Z")

</div>

Hi,  
When using with default configuration and with no mapping, fields are  
analyzed with lowercase token filter. So when I index a field with value,  
let's say "ABC", it is tokenized as "abc". When I try to search it as I  
insert it with the following query I get no results:

{  
"query":{"term":"ABC"}  
}

It seems that only query strings supports analyzers during search. Is there  
a plan to add this feature to Elastic Search ?

Thanks in advance,  
Sezgin Kucukkaraaslan  
[www.ifountain.com](http://www.ifountain.com)

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [June 30, 2010, 12:19pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/2 "2010-06-30T12:19:58Z")

</div>

Hi,  
I did not try myself but it is possible to specify analyzer as a query  
parameter:  
[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/#Request\_Parameters](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/#Request_Parameters)  
[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/#Request\_Parameters](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/#Request_Parameters)So  
it should definitely work in JSON too. Also did you check  
[http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/analyzer/#Default\_Analyzers](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/#Default_Analyzers)  
?  
Lukas

2010/6/30 sezgin küçükkaraaslan [sezo104@gmail.com](mailto:sezo104@gmail.com)

> Hi,  
> When using with default configuration and with no mapping, fields are  
> analyzed with lowercase token filter. So when I index a field with value,  
> let's say "ABC", it is tokenized as "abc". When I try to search it as I  
> insert it with the following query I get no results:
> 
> {  
> "query":{"term":"ABC"}  
> }
> 
> It seems that only query strings supports analyzers during search. Is there  
> a plan to add this feature to Elastic Search ?
> 
> Thanks in advance,  
> Sezgin Kucukkaraaslan  
> [www.ifountain.com](http://www.ifountain.com)

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [June 30, 2010, 12:22pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/3 "2010-06-30T12:22:24Z")

</div>

Well... the term query is not analyzed, see:  
[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/query\_dsl/term\_query/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/query_dsl/term_query/)

On Wed, Jun 30, 2010 at 2:19 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> Hi,  
> I did not try myself but it is possible to specify analyzer as a query  
> parameter:  
> [http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/#Request\_Parameters](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/#Request_Parameters)  
> [http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/#Request\_Parameters](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/#Request_Parameters)So  
> it should definitely work in JSON too. Also did you check  
> [Analysis | Elasticsearch Guide [8.11] | Elastic](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/#Default_Analyzers)  
> ?  
> Lukas
> 
> 2010/6/30 sezgin küçükkaraaslan [sezo104@gmail.com](mailto:sezo104@gmail.com)
> 
> Hi,
> 
> > When using with default configuration and with no mapping, fields are  
> > analyzed with lowercase token filter. So when I index a field with value,  
> > let's say "ABC", it is tokenized as "abc". When I try to search it as I  
> > insert it with the following query I get no results:
> > 
> > {  
> > "query":{"term":"ABC"}  
> > }
> > 
> > It seems that only query strings supports analyzers during search. Is  
> > there a plan to add this feature to Elastic Search ?
> > 
> > Thanks in advance,  
> > Sezgin Kucukkaraaslan  
> > [www.ifountain.com](http://www.ifountain.com)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 30, 2010, 4:44pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/4 "2010-06-30T16:44:05Z")

</div>

If you want to have the text passed analyzed, then use the field query  
(which is a nice field level wrapper for the query\_string query):  
[http://www.elasticsearch.com/docs/elasticsearch/rest\_api/query\_dsl/field\_query/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/query_dsl/field_query/)  
.

Note that the analysis process can be a simple one as lowercasing, and can  
be more complex one that generates several terms for a single term analyzed.

-shay.banon

On Wed, Jun 30, 2010 at 3:22 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> Well... the term query is not analyzed, see:  
> [http://www.elasticsearch.com/docs/elasticsearch/rest\_api/query\_dsl/term\_query/](http://www.elasticsearch.com/docs/elasticsearch/rest_api/query_dsl/term_query/)
> 
> On Wed, Jun 30, 2010 at 2:19 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com)wrote:
> 
> > Hi,  
> > I did not try myself but it is possible to specify analyzer as a query  
> > parameter:  
> > [http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/#Request\_Parameters](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/#Request_Parameters)  
> > [http://www.elasticsearch.com/docs/elasticsearch/rest\_api/search/#Request\_Parameters](http://www.elasticsearch.com/docs/elasticsearch/rest_api/search/#Request_Parameters)So  
> > it should definitely work in JSON too. Also did you check  
> > [Analysis | Elasticsearch Guide [8.11] | Elastic](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/analyzer/#Default_Analyzers)  
> > ?  
> > Lukas
> > 
> > 2010/6/30 sezgin küçükkaraaslan [sezo104@gmail.com](mailto:sezo104@gmail.com)
> > 
> > Hi,
> > 
> > > When using with default configuration and with no mapping, fields are  
> > > analyzed with lowercase token filter. So when I index a field with value,  
> > > let's say "ABC", it is tokenized as "abc". When I try to search it as I  
> > > insert it with the following query I get no results:
> > > 
> > > {  
> > > "query":{"term":"ABC"}  
> > > }
> > > 
> > > It seems that only query strings supports analyzers during search. Is  
> > > there a plan to add this feature to Elastic Search ?
> > > 
> > > Thanks in advance,  
> > > Sezgin Kucukkaraaslan  
> > > [www.ifountain.com](http://www.ifountain.com)

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [June 30, 2010, 4:48pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/5 "2010-06-30T16:48:20Z")

</div>

For context - lukasvlcek had the conversation below in IRC, then left.

I'm answering him here

* * *

lukasvlcek:  
kimchy: I haven't been thinking about it before... what is the  
rationale of not allowing analyzer setup for term query when  
Query DSL is used? See  
[http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html](http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html)

```
    I am just curious why user has to search -exact- terms (Lower vs
    Upper case)

```

sam\_:  
the default analyzer if nothing is specified is standard isn't  
it?

lukasvlcek:  
I did not try this particular example but I am confused by the  
term query doc which explicitly says "not analyzed" (so even the  
default analyzer is not used?)

sam\_:  
if it is not analyzed then I would suspect you need to provide  
case  
an exact match  
the standard analyzer would result in it being converted

lukasvlcek:  
wouldn't it be useful to have ability to specify analyzer?

sam\_:  
you can  
well  
at least when you define the mappings  
the analyzer is used as part of the indexing  
as an alternative I would think you could provide your own  
parser implementation to which is what I'm trying to do  
but have been unsuccessful

lukasvlcek:  
but the point is if it is possible to specify analyzer when  
querying via URL parameters then why can not specify analyzer  
while using Query DSL  
Gotta go now... but I would appreciate if anybody (kimchy?) can  
follow up on that mail thread above (want to check that later)

ï»¿----------------------------------------------

Answer:

(Note - this is as I understand the situation - I'm open to correction)

All data stored in ElasticSearch/Lucene is stored as a 'term' which is  
atomic - it can't be broken down further.

So if you index {"text": "The quick brown fox jumped over the LAZY dog"}  
then the default analyzer would:

- remove stopwords
- lowercase all text
- split on whitespace and punctuation
- result in these terms:  
'quick', 'brown', 'fox', 'jumped','over', 'lazy', 'dog'

If you then do this search:  
{ "query\_string": { "query": "QUICK dOg"}}

Then the default analyzer would analyze your query string and return the  
following terms: "quick", "dog"

It then does a 'term' query for each of those terms and combines the  
results.

If you did this search:  
{ "wildcard": {"text": "_o_}}

Then it would first look at all terms, and find only those terms that  
match that pattern, ie: 'brown', 'fox', 'over', 'dog'.

It then does a 'term' query for each of those terms and combines the  
results.

So it doesn't make sense to analyze a 'term'. Terms are the result of  
analysis. If you need to analyse a search "phrase" then you should use a  
"query\_string" or "field" query.

For the same reason, you can't sort on an analyzed field because the  
original data doesn't exist. It is tokenised and stored as  
terms. ï»¿(unless the field is also stored? - not sure)

The analyzer used to analyze a search phrase is selected in this order:

- "analyzer" specified in the query DSL, eg:

- "search\_analyzer" specified in the mapping

- "analyzer" specified in the mapping

- the default\_search analyzer ï»¿specified in the index configuration

- the default analyzer specified in the index configuration

- the default\_search analyzer specified in the node configuration

- the default analyzer specified in the node configuration

- the "standard" analyzer

(I think that's right - I may have added a couple in there that don't  
actually exist)

Typically, it doesn't make sense to use a different analyzer at index  
and search time, because you may end up searching for terms that don't  
actually exist.

If a field is set to be 'not\_analyzed', then the whole value is treated  
as a term, so "ABC" and "abc" are different, and "abc" will not match  
"abc def".

hope this helps

Clint

--  
Web Announcements Limited is a company registered in England and Wales,  
with company number 05608868, with registered address at 10 Arvon Road,  
London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [June 30, 2010, 11:01pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/6 "2010-06-30T23:01:35Z")

</div>

Hey guys, thanks for keeping this conversation going. Appreciate this!  
Lukas

On Wed, Jun 30, 2010 at 6:48 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> For context - lukasvlcek had the conversation below in IRC, then left.
> 
> I'm answering him here
> 
> * * *
> 
> lukasvlcek:  
> kimchy: I haven't been thinking about it before... what is the  
> rationale of not allowing analyzer setup for term query when  
> Query DSL is used? See
> 
> [http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html](http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html)
> 
> ```
> I am just curious why user has to search -exact- terms (Lower vs
> Upper case)
> 
> ```
> 
> sam\_:  
> the default analyzer if nothing is specified is standard isn't  
> it?
> 
> lukasvlcek:  
> I did not try this particular example but I am confused by the  
> term query doc which explicitly says "not analyzed" (so even the  
> default analyzer is not used?)
> 
> sam\_:  
> if it is not analyzed then I would suspect you need to provide  
> case  
> an exact match  
> the standard analyzer would result in it being converted
> 
> lukasvlcek:  
> wouldn't it be useful to have ability to specify analyzer?
> 
> sam\_:  
> you can  
> well  
> at least when you define the mappings  
> the analyzer is used as part of the indexing  
> as an alternative I would think you could provide your own  
> parser implementation to which is what I'm trying to do  
> but have been unsuccessful
> 
> lukasvlcek:  
> but the point is if it is possible to specify analyzer when  
> querying via URL parameters then why can not specify analyzer  
> while using Query DSL  
> Gotta go now... but I would appreciate if anybody (kimchy?) can  
> follow up on that mail thread above (want to check that later)
> 
> ----------------------------------------------
> 
> Answer:
> 
> (Note - this is as I understand the situation - I'm open to correction)
> 
> All data stored in Elasticsearch/Lucene is stored as a 'term' which is  
> atomic - it can't be broken down further.
> 
> So if you index {"text": "The quick brown fox jumped over the LAZY dog"}  
> then the default analyzer would:
> 
> - remove stopwords
> - lowercase all text
> - split on whitespace and punctuation
> - result in these terms:  
> 'quick', 'brown', 'fox', 'jumped','over', 'lazy', 'dog'
> 
> If you then do this search:  
> { "query\_string": { "query": "QUICK dOg"}}
> 
> Then the default analyzer would analyze your query string and return the  
> following terms: "quick", "dog"
> 
> It then does a 'term' query for each of those terms and combines the  
> results.
> 
> If you did this search:  
> { "wildcard": {"text": "_o_}}
> 
> Then it would first look at all terms, and find only those terms that  
> match that pattern, ie: 'brown', 'fox', 'over', 'dog'.
> 
> It then does a 'term' query for each of those terms and combines the  
> results.
> 
> So it doesn't make sense to analyze a 'term'. Terms are the result of  
> analysis. If you need to analyse a search "phrase" then you should use a  
> "query\_string" or "field" query.
> 
> For the same reason, you can't sort on an analyzed field because the  
> original data doesn't exist. It is tokenised and stored as  
> terms. ﻿(unless the field is also stored? - not sure)
> 
> The analyzer used to analyze a search phrase is selected in this order:
> 
> - "analyzer" specified in the query DSL, eg:
> 
> - "search\_analyzer" specified in the mapping
> 
> - "analyzer" specified in the mapping
> 
> - the default\_search analyzer ﻿specified in the index configuration
> 
> - the default analyzer specified in the index configuration
> 
> - the default\_search analyzer specified in the node configuration
> 
> - the default analyzer specified in the node configuration
> 
> - the "standard" analyzer
> 
> (I think that's right - I may have added a couple in there that don't  
> actually exist)
> 
> Typically, it doesn't make sense to use a different analyzer at index  
> and search time, because you may end up searching for terms that don't  
> actually exist.
> 
> If a field is set to be 'not\_analyzed', then the whole value is treated  
> as a term, so "ABC" and "abc" are different, and "abc" will not match  
> "abc def".
> 
> hope this helps
> 
> Clint
> 
> --  
> Web Announcements Limited is a company registered in England and Wales,  
> with company number 05608868, with registered address at 10 Arvon Road,  
> London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![sezgin\_kucukkaraasla](https://avatars.discourse-cdn.com/v4/letter/s/9de0a6/32.png) [@sezgin\_kucukkaraasla](https://discuss.elastic.co/u/sezgin_kucukkaraasla)\
**Post date:** [July 1, 2010, 8:54am UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/7 "2010-07-01T08:54:55Z")

</div>

Thanks for the replies..  
I think I'd better to explain what I'm trying to do. I'm working on a IT  
event management application and want to store my data on Elastic Search to  
leverage it's clustering and redundancy features. The requirement is to  
index data and give the operators the flexibility to search events case  
insensitively from the UI. To gain from performance I don't want all fields  
in my event model to be analyzed. For example I want to keep fields like  
"identifier", which I know that it will consist of one word, as  
"not\_analyzed". The problem with field search here is that I can't search  
these kind of properties with it. (It gives zero result.). So I will not be  
able to use it all the time. After some thinking, I decide to use two kinds  
of analyzers for my fields, which are:

myAnalyzer1 :  
filter: [lowercase]  
tokenizer: keyword

for the fields like "identifier", and :

myAnalyzer2:  
filter:[lowercase]  
tokenizer: whitespace

for the fields like "description", which can consist of multiple words.

I can take some advices here, am I in the right path? Is there any  
performance loss that I will bear by using the first analyzer instead of  
keeping it as "not\_analyzed"?  
Thank you very much again...

Sezgin Kucukkaraaslan  
[www.ifountain.com](http://www.ifountain.com)

On Thu, Jul 1, 2010 at 2:01 AM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> Hey guys, thanks for keeping this conversation going. Appreciate this!  
> Lukas
> 
> On Wed, Jun 30, 2010 at 6:48 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:
> 
> > For context - lukasvlcek had the conversation below in IRC, then left.
> > 
> > I'm answering him here
> > 
> > * * *
> > 
> > lukasvlcek:  
> > kimchy: I haven't been thinking about it before... what is the  
> > rationale of not allowing analyzer setup for term query when  
> > Query DSL is used? See
> > 
> > [http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html](http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html)
> > 
> > ```
> > I am just curious why user has to search -exact- terms (Lower vs
> > Upper case)
> > 
> > ```
> > 
> > sam\_:  
> > the default analyzer if nothing is specified is standard isn't  
> > it?
> > 
> > lukasvlcek:  
> > I did not try this particular example but I am confused by the  
> > term query doc which explicitly says "not analyzed" (so even the  
> > default analyzer is not used?)
> > 
> > sam\_:  
> > if it is not analyzed then I would suspect you need to provide  
> > case  
> > an exact match  
> > the standard analyzer would result in it being converted
> > 
> > lukasvlcek:  
> > wouldn't it be useful to have ability to specify analyzer?
> > 
> > sam\_:  
> > you can  
> > well  
> > at least when you define the mappings  
> > the analyzer is used as part of the indexing  
> > as an alternative I would think you could provide your own  
> > parser implementation to which is what I'm trying to do  
> > but have been unsuccessful
> > 
> > lukasvlcek:  
> > but the point is if it is possible to specify analyzer when  
> > querying via URL parameters then why can not specify analyzer  
> > while using Query DSL  
> > Gotta go now... but I would appreciate if anybody (kimchy?) can  
> > follow up on that mail thread above (want to check that later)
> > 
> > ----------------------------------------------
> > 
> > Answer:
> > 
> > (Note - this is as I understand the situation - I'm open to correction)
> > 
> > All data stored in Elasticsearch/Lucene is stored as a 'term' which is  
> > atomic - it can't be broken down further.
> > 
> > So if you index {"text": "The quick brown fox jumped over the LAZY dog"}  
> > then the default analyzer would:
> > 
> > - remove stopwords
> > - lowercase all text
> > - split on whitespace and punctuation
> > - result in these terms:  
> > 'quick', 'brown', 'fox', 'jumped','over', 'lazy', 'dog'
> > 
> > If you then do this search:  
> > { "query\_string": { "query": "QUICK dOg"}}
> > 
> > Then the default analyzer would analyze your query string and return the  
> > following terms: "quick", "dog"
> > 
> > It then does a 'term' query for each of those terms and combines the  
> > results.
> > 
> > If you did this search:  
> > { "wildcard": {"text": "_o_}}
> > 
> > Then it would first look at all terms, and find only those terms that  
> > match that pattern, ie: 'brown', 'fox', 'over', 'dog'.
> > 
> > It then does a 'term' query for each of those terms and combines the  
> > results.
> > 
> > So it doesn't make sense to analyze a 'term'. Terms are the result of  
> > analysis. If you need to analyse a search "phrase" then you should use a  
> > "query\_string" or "field" query.
> > 
> > For the same reason, you can't sort on an analyzed field because the  
> > original data doesn't exist. It is tokenised and stored as  
> > terms. ﻿(unless the field is also stored? - not sure)
> > 
> > The analyzer used to analyze a search phrase is selected in this order:
> > 
> > - "analyzer" specified in the query DSL, eg:
> > 
> > - "search\_analyzer" specified in the mapping
> > 
> > - "analyzer" specified in the mapping
> > 
> > - the default\_search analyzer ﻿specified in the index configuration
> > 
> > - the default analyzer specified in the index configuration
> > 
> > - the default\_search analyzer specified in the node configuration
> > 
> > - the default analyzer specified in the node configuration
> > 
> > - the "standard" analyzer
> > 
> > (I think that's right - I may have added a couple in there that don't  
> > actually exist)
> > 
> > Typically, it doesn't make sense to use a different analyzer at index  
> > and search time, because you may end up searching for terms that don't  
> > actually exist.
> > 
> > If a field is set to be 'not\_analyzed', then the whole value is treated  
> > as a term, so "ABC" and "abc" are different, and "abc" will not match  
> > "abc def".
> > 
> > hope this helps
> > 
> > Clint
> > 
> > --  
> > Web Announcements Limited is a company registered in England and Wales,  
> > with company number 05608868, with registered address at 10 Arvon Road,  
> > London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 2, 2010, 1:34pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/8 "2010-07-02T13:34:55Z")

</div>

Its a good way to solve what you are trying. You shouldn't notice the  
performance difference in indexing time with this compared to  
`not_analyzed`.

-shay.banon

2010/7/1 sezgin küçükkaraaslan [sezo104@gmail.com](mailto:sezo104@gmail.com)

> Thanks for the replies..  
> I think I'd better to explain what I'm trying to do. I'm working on a IT  
> event management application and want to store my data on Elastic Search to  
> leverage it's clustering and redundancy features. The requirement is to  
> index data and give the operators the flexibility to search events case  
> insensitively from the UI. To gain from performance I don't want all fields  
> in my event model to be analyzed. For example I want to keep fields like  
> "identifier", which I know that it will consist of one word, as  
> "not\_analyzed". The problem with field search here is that I can't search  
> these kind of properties with it. (It gives zero result.). So I will not be  
> able to use it all the time. After some thinking, I decide to use two kinds  
> of analyzers for my fields, which are:
> 
> myAnalyzer1 :  
> filter: [lowercase]  
> tokenizer: keyword
> 
> for the fields like "identifier", and :
> 
> myAnalyzer2:  
> filter:[lowercase]  
> tokenizer: whitespace
> 
> for the fields like "description", which can consist of multiple words.
> 
> I can take some advices here, am I in the right path? Is there any  
> performance loss that I will bear by using the first analyzer instead of  
> keeping it as "not\_analyzed"?  
> Thank you very much again...
> 
> Sezgin Kucukkaraaslan  
> [www.ifountain.com](http://www.ifountain.com)
> 
> On Thu, Jul 1, 2010 at 2:01 AM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:
> 
> > Hey guys, thanks for keeping this conversation going. Appreciate this!  
> > Lukas
> > 
> > On Wed, Jun 30, 2010 at 6:48 PM, Clinton Gormley \<[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)
> > 
> > > wrote:
> > 
> > > For context - lukasvlcek had the conversation below in IRC, then left.
> > > 
> > > I'm answering him here
> > > 
> > > * * *
> > > 
> > > lukasvlcek:  
> > > kimchy: I haven't been thinking about it before... what is the  
> > > rationale of not allowing analyzer setup for term query when  
> > > Query DSL is used? See
> > > 
> > > [http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html](http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html)
> > > 
> > > ```
> > > I am just curious why user has to search -exact- terms (Lower vs
> > > Upper case)
> > > 
> > > ```
> > > 
> > > sam\_:  
> > > the default analyzer if nothing is specified is standard isn't  
> > > it?
> > > 
> > > lukasvlcek:  
> > > I did not try this particular example but I am confused by the  
> > > term query doc which explicitly says "not analyzed" (so even the  
> > > default analyzer is not used?)
> > > 
> > > sam\_:  
> > > if it is not analyzed then I would suspect you need to provide  
> > > case  
> > > an exact match  
> > > the standard analyzer would result in it being converted
> > > 
> > > lukasvlcek:  
> > > wouldn't it be useful to have ability to specify analyzer?
> > > 
> > > sam\_:  
> > > you can  
> > > well  
> > > at least when you define the mappings  
> > > the analyzer is used as part of the indexing  
> > > as an alternative I would think you could provide your own  
> > > parser implementation to which is what I'm trying to do  
> > > but have been unsuccessful
> > > 
> > > lukasvlcek:  
> > > but the point is if it is possible to specify analyzer when  
> > > querying via URL parameters then why can not specify analyzer  
> > > while using Query DSL  
> > > Gotta go now... but I would appreciate if anybody (kimchy?) can  
> > > follow up on that mail thread above (want to check that later)
> > > 
> > > ----------------------------------------------
> > > 
> > > Answer:
> > > 
> > > (Note - this is as I understand the situation - I'm open to correction)
> > > 
> > > All data stored in Elasticsearch/Lucene is stored as a 'term' which is  
> > > atomic - it can't be broken down further.
> > > 
> > > So if you index {"text": "The quick brown fox jumped over the LAZY dog"}  
> > > then the default analyzer would:
> > > 
> > > - remove stopwords
> > > - lowercase all text
> > > - split on whitespace and punctuation
> > > - result in these terms:  
> > > 'quick', 'brown', 'fox', 'jumped','over', 'lazy', 'dog'
> > > 
> > > If you then do this search:  
> > > { "query\_string": { "query": "QUICK dOg"}}
> > > 
> > > Then the default analyzer would analyze your query string and return the  
> > > following terms: "quick", "dog"
> > > 
> > > It then does a 'term' query for each of those terms and combines the  
> > > results.
> > > 
> > > If you did this search:  
> > > { "wildcard": {"text": "_o_}}
> > > 
> > > Then it would first look at all terms, and find only those terms that  
> > > match that pattern, ie: 'brown', 'fox', 'over', 'dog'.
> > > 
> > > It then does a 'term' query for each of those terms and combines the  
> > > results.
> > > 
> > > So it doesn't make sense to analyze a 'term'. Terms are the result of  
> > > analysis. If you need to analyse a search "phrase" then you should use a  
> > > "query\_string" or "field" query.
> > > 
> > > For the same reason, you can't sort on an analyzed field because the  
> > > original data doesn't exist. It is tokenised and stored as  
> > > terms. ﻿(unless the field is also stored? - not sure)
> > > 
> > > The analyzer used to analyze a search phrase is selected in this order:
> > > 
> > > - "analyzer" specified in the query DSL, eg:
> > > 
> > > - "search\_analyzer" specified in the mapping
> > > 
> > > - "analyzer" specified in the mapping
> > > 
> > > - the default\_search analyzer ﻿specified in the index configuration
> > > 
> > > - the default analyzer specified in the index configuration
> > > 
> > > - the default\_search analyzer specified in the node configuration
> > > 
> > > - the default analyzer specified in the node configuration
> > > 
> > > - the "standard" analyzer
> > > 
> > > (I think that's right - I may have added a couple in there that don't  
> > > actually exist)
> > > 
> > > Typically, it doesn't make sense to use a different analyzer at index  
> > > and search time, because you may end up searching for terms that don't  
> > > actually exist.
> > > 
> > > If a field is set to be 'not\_analyzed', then the whole value is treated  
> > > as a term, so "ABC" and "abc" are different, and "abc" will not match  
> > > "abc def".
> > > 
> > > hope this helps
> > > 
> > > Clint
> > > 
> > > --  
> > > Web Announcements Limited is a company registered in England and Wales,  
> > > with company number 05608868, with registered address at 10 Arvon Road,  
> > > London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![sezgin\_kucukkaraasla](https://avatars.discourse-cdn.com/v4/letter/s/9de0a6/32.png) [@sezgin\_kucukkaraasla](https://discuss.elastic.co/u/sezgin_kucukkaraasla)\
**Post date:** [July 7, 2010, 11:25am UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/9 "2010-07-07T11:25:43Z")

</div>

Thanks...

On Fri, Jul 2, 2010 at 4:34 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Its a good way to solve what you are trying. You shouldn't notice the  
> performance difference in indexing time with this compared to  
> `not_analyzed`.
> 
> -shay.banon
> 
> 2010/7/1 sezgin küçükkaraaslan [sezo104@gmail.com](mailto:sezo104@gmail.com)
> 
> Thanks for the replies..
> 
> > I think I'd better to explain what I'm trying to do. I'm working on a IT  
> > event management application and want to store my data on Elastic Search to  
> > leverage it's clustering and redundancy features. The requirement is to  
> > index data and give the operators the flexibility to search events case  
> > insensitively from the UI. To gain from performance I don't want all fields  
> > in my event model to be analyzed. For example I want to keep fields like  
> > "identifier", which I know that it will consist of one word, as  
> > "not\_analyzed". The problem with field search here is that I can't search  
> > these kind of properties with it. (It gives zero result.). So I will not be  
> > able to use it all the time. After some thinking, I decide to use two kinds  
> > of analyzers for my fields, which are:
> > 
> > myAnalyzer1 :  
> > filter: [lowercase]  
> > tokenizer: keyword
> > 
> > for the fields like "identifier", and :
> > 
> > myAnalyzer2:  
> > filter:[lowercase]  
> > tokenizer: whitespace
> > 
> > for the fields like "description", which can consist of multiple words.
> > 
> > I can take some advices here, am I in the right path? Is there any  
> > performance loss that I will bear by using the first analyzer instead of  
> > keeping it as "not\_analyzed"?  
> > Thank you very much again...
> > 
> > Sezgin Kucukkaraaslan  
> > [www.ifountain.com](http://www.ifountain.com)
> > 
> > On Thu, Jul 1, 2010 at 2:01 AM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com)wrote:
> > 
> > > Hey guys, thanks for keeping this conversation going. Appreciate this!  
> > > Lukas
> > > 
> > > On Wed, Jun 30, 2010 at 6:48 PM, Clinton Gormley \<  
> > > [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)\> wrote:
> > > 
> > > > For context - lukasvlcek had the conversation below in IRC, then left.
> > > > 
> > > > I'm answering him here
> > > > 
> > > > * * *
> > > > 
> > > > lukasvlcek:  
> > > > kimchy: I haven't been thinking about it before... what is the  
> > > > rationale of not allowing analyzer setup for term query when  
> > > > Query DSL is used? See
> > > > 
> > > > [http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html](http://elasticsearch-users.115913.n3.nabble.com/Is-there-a-way-to-search-terms-lower-cased-tp932996.html)
> > > > 
> > > > ```
> > > > I am just curious why user has to search -exact- terms (Lower vs
> > > > Upper case)
> > > > 
> > > > ```
> > > > 
> > > > sam\_:  
> > > > the default analyzer if nothing is specified is standard isn't  
> > > > it?
> > > > 
> > > > lukasvlcek:  
> > > > I did not try this particular example but I am confused by the  
> > > > term query doc which explicitly says "not analyzed" (so even the  
> > > > default analyzer is not used?)
> > > > 
> > > > sam\_:  
> > > > if it is not analyzed then I would suspect you need to provide  
> > > > case  
> > > > an exact match  
> > > > the standard analyzer would result in it being converted
> > > > 
> > > > lukasvlcek:  
> > > > wouldn't it be useful to have ability to specify analyzer?
> > > > 
> > > > sam\_:  
> > > > you can  
> > > > well  
> > > > at least when you define the mappings  
> > > > the analyzer is used as part of the indexing  
> > > > as an alternative I would think you could provide your own  
> > > > parser implementation to which is what I'm trying to do  
> > > > but have been unsuccessful
> > > > 
> > > > lukasvlcek:  
> > > > but the point is if it is possible to specify analyzer when  
> > > > querying via URL parameters then why can not specify analyzer  
> > > > while using Query DSL  
> > > > Gotta go now... but I would appreciate if anybody (kimchy?) can  
> > > > follow up on that mail thread above (want to check that later)
> > > > 
> > > > ----------------------------------------------
> > > > 
> > > > Answer:
> > > > 
> > > > (Note - this is as I understand the situation - I'm open to correction)
> > > > 
> > > > All data stored in Elasticsearch/Lucene is stored as a 'term' which is  
> > > > atomic - it can't be broken down further.
> > > > 
> > > > So if you index {"text": "The quick brown fox jumped over the LAZY dog"}  
> > > > then the default analyzer would:
> > > > 
> > > > - remove stopwords
> > > > - lowercase all text
> > > > - split on whitespace and punctuation
> > > > - result in these terms:  
> > > > 'quick', 'brown', 'fox', 'jumped','over', 'lazy', 'dog'
> > > > 
> > > > If you then do this search:  
> > > > { "query\_string": { "query": "QUICK dOg"}}
> > > > 
> > > > Then the default analyzer would analyze your query string and return the  
> > > > following terms: "quick", "dog"
> > > > 
> > > > It then does a 'term' query for each of those terms and combines the  
> > > > results.
> > > > 
> > > > If you did this search:  
> > > > { "wildcard": {"text": "_o_}}
> > > > 
> > > > Then it would first look at all terms, and find only those terms that  
> > > > match that pattern, ie: 'brown', 'fox', 'over', 'dog'.
> > > > 
> > > > It then does a 'term' query for each of those terms and combines the  
> > > > results.
> > > > 
> > > > So it doesn't make sense to analyze a 'term'. Terms are the result of  
> > > > analysis. If you need to analyse a search "phrase" then you should use a  
> > > > "query\_string" or "field" query.
> > > > 
> > > > For the same reason, you can't sort on an analyzed field because the  
> > > > original data doesn't exist. It is tokenised and stored as  
> > > > terms. ﻿(unless the field is also stored? - not sure)
> > > > 
> > > > The analyzer used to analyze a search phrase is selected in this order:
> > > > 
> > > > - "analyzer" specified in the query DSL, eg:
> > > > 
> > > > - "search\_analyzer" specified in the mapping
> > > > 
> > > > - "analyzer" specified in the mapping
> > > > 
> > > > - the default\_search analyzer ﻿specified in the index configuration
> > > > 
> > > > - the default analyzer specified in the index configuration
> > > > 
> > > > - the default\_search analyzer specified in the node configuration
> > > > 
> > > > - the default analyzer specified in the node configuration
> > > > 
> > > > - the "standard" analyzer
> > > > 
> > > > (I think that's right - I may have added a couple in there that don't  
> > > > actually exist)
> > > > 
> > > > Typically, it doesn't make sense to use a different analyzer at index  
> > > > and search time, because you may end up searching for terms that don't  
> > > > actually exist.
> > > > 
> > > > If a field is set to be 'not\_analyzed', then the whole value is treated  
> > > > as a term, so "ABC" and "abc" are different, and "abc" will not match  
> > > > "abc def".
> > > > 
> > > > hope this helps
> > > > 
> > > > Clint
> > > > 
> > > > --  
> > > > Web Announcements Limited is a company registered in England and Wales,  
> > > > with company number 05608868, with registered address at 10 Arvon Road,  
> > > > London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:22am UTC](https://discuss.elastic.co/t/is-there-a-way-to-search-terms-lower-cased/3056/10 "2017-07-06T04:22:37Z")

</div>


