# Searching an index with 2 types using a keyword tokenizer

**URL:** <https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811>\
**Category:** Elasticsearch\
**Created:** [July 16, 2013, 4:21pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811 "2013-07-16T16:21:59Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 16, 2013, 4:21pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/1 "2013-07-16T16:21:59Z")

</div>

Hi,

I'm getting strange results trying to search on an index with 2 types using  
a keyword tokenizer.

Using ElasticSearch 0.90.2 this :  
curl -XGET localhost:9200/myindex/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
f"}}}'  
Returns a result containing "Carex" alone (unexpected behaviour)

curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
f"}}}'  
Will return an expected results of "Carex feta" and not "Carex" alone.

If I do the same thing using ElasticSearch 0.90.1, the 2 queries above will  
return the expected results. This could be related to different  
configuration but I am using the default configurations on both versions.

So, I would like to know what is the ElasticSearch expected behavior for  
the first query?  
Could it be related to ES using a default tokenizer (Standard) when we use  
multiple types?

Here are the current settings:

curl -XPOST "localhost:9200/myindex" -d '  
{  
"settings":{  
"index":{  
"analysis":{  
"filter" : {  
"name\_nGram" : {  
"max\_gram" : 100,  
"min\_gram" : 2,  
"type" : "edge\_ngram"  
}  
},  
"analyzer":{  
"name\_index" : {  
"filter" : [  
"lowercase","asciifolding","name\_nGram"  
],  
"tokenizer" : "standard"  
},  
"full\_name\_index" : {  
"filter" : [  
"lowercase","asciifolding"  
],  
"tokenizer" : "keyword"  
},  
"scientificname\_index" : {  
"filter" : [  
"lowercase","asciifolding","name\_nGram"  
],  
"tokenizer" : "keyword"  
},  
"name\_search" : {  
"filter" : [  
"lowercase","asciifolding"  
],  
"tokenizer" : "keyword"  
}  
}  
}  
}  
},  
"mappings" : {  
"taxon" : {  
"properties" : {  
"name" : {  
"type" : "multi\_field",  
"fields":{  
"name":{  
"type" : "string",  
"index\_analyzer" : "full\_name\_index",  
"search\_analyzer" : "name\_search"  
},  
"ngrams":{  
"type" : "string",  
"index\_analyzer" : "scientificname\_index",  
"search\_analyzer" : "name\_search"  
}  
}  
},  
"status":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"namehtml":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"namehtmlauthor":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"rankname":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"parentid":{  
"index" : "not\_analyzed",  
"type" : "integer"  
},  
"parentnamehtml":{  
"index" : "not\_analyzed",  
"type" : "string"  
}  
}  
},  
"vernacular" : {  
"properties" : {  
"name" : {  
"type" : "multi\_field",  
"fields":{  
"name":{  
"type" : "string",  
"index" : "not\_analyzed"  
},  
"ngrams":{  
"type" : "string",  
"search\_analyzer" : "name\_search",  
"index\_analyzer" : "name\_index"  
}  
}  
},  
"taxonid":{  
"index" : "not\_analyzed",  
"type" : "integer"  
},  
"status":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"language":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"taxonnamehtml":{  
"index" : "not\_analyzed",  
"type" : "string"  
}  
}  
}  
}  
}'

Add some data:

curl -XPUT '[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)' -d '{  
"name" : "carex"  
}'  
curl -XPUT '[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)' -d '{  
"name" : "carex feta"  
}'

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [July 17, 2013, 12:20pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/2 "2013-07-17T12:20:28Z")

</div>

Hi Christian,  
I just tested it with elasticsearch 0.90.2 and I got the expected results:  
1 results querying for 'carex', one querying for 'carex feta'.

Your weird results seem to be caused by indexing ngrams, since you get back  
partial results. On the other hand the recreation that you posted works  
fine. I would check again the field you're querying on and your mapping,  
maybe you're doing something slightly different from your curl example?

On Tuesday, July 16, 2013 6:21:59 PM UTC+2, Christian Gendreau wrote:

> Hi,
> 
> I'm getting strange results trying to search on an index with 2 types  
> using a keyword tokenizer.
> 
> Using Elasticsearch 0.90.2 this :  
> curl -XGET localhost:9200/myindex/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> f"}}}'  
> Returns a result containing "Carex" alone (unexpected behaviour)
> 
> curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> f"}}}'  
> Will return an expected results of "Carex feta" and not "Carex" alone.
> 
> If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> will return the expected results. This could be related to different  
> configuration but I am using the default configurations on both versions.
> 
> So, I would like to know what is the Elasticsearch expected behavior for  
> the first query?  
> Could it be related to ES using a default tokenizer (Standard) when we use  
> multiple types?
> 
> Here are the current settings:
> 
> curl -XPOST "localhost:9200/myindex" -d '  
> {  
> "settings":{  
> "index":{  
> "analysis":{  
> "filter" : {  
> "name\_nGram" : {  
> "max\_gram" : 100,  
> "min\_gram" : 2,  
> "type" : "edge\_ngram"  
> }  
> },  
> "analyzer":{  
> "name\_index" : {  
> "filter" : [  
> "lowercase","asciifolding","name\_nGram"  
> ],  
> "tokenizer" : "standard"  
> },  
> "full\_name\_index" : {  
> "filter" : [  
> "lowercase","asciifolding"  
> ],  
> "tokenizer" : "keyword"  
> },  
> "scientificname\_index" : {  
> "filter" : [  
> "lowercase","asciifolding","name\_nGram"  
> ],  
> "tokenizer" : "keyword"  
> },  
> "name\_search" : {  
> "filter" : [  
> "lowercase","asciifolding"  
> ],  
> "tokenizer" : "keyword"  
> }  
> }  
> }  
> }  
> },  
> "mappings" : {  
> "taxon" : {  
> "properties" : {  
> "name" : {  
> "type" : "multi\_field",  
> "fields":{  
> "name":{  
> "type" : "string",  
> "index\_analyzer" : "full\_name\_index",  
> "search\_analyzer" : "name\_search"  
> },  
> "ngrams":{  
> "type" : "string",  
> "index\_analyzer" : "scientificname\_index",  
> "search\_analyzer" : "name\_search"  
> }  
> }  
> },  
> "status":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "namehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "namehtmlauthor":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "rankname":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "parentid":{  
> "index" : "not\_analyzed",  
> "type" : "integer"  
> },  
> "parentnamehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> }  
> }  
> },  
> "vernacular" : {  
> "properties" : {  
> "name" : {  
> "type" : "multi\_field",  
> "fields":{  
> "name":{  
> "type" : "string",  
> "index" : "not\_analyzed"  
> },  
> "ngrams":{  
> "type" : "string",  
> "search\_analyzer" : "name\_search",  
> "index\_analyzer" : "name\_index"  
> }  
> }  
> },  
> "taxonid":{  
> "index" : "not\_analyzed",  
> "type" : "integer"  
> },  
> "status":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "language":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "taxonnamehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> }  
> }  
> }  
> }  
> }'
> 
> Add some data:
> 
> curl -XPUT '[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)' -d '{  
> "name" : "carex"  
> }'  
> curl -XPUT '[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)' -d '{  
> "name" : "carex feta"  
> }'

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 17, 2013, 3:36pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/3 "2013-07-17T15:36:15Z")

</div>

Hi Luca,

Sorry there is an error in my post:  
curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
f"}}}'  
Will return an expected result of : _nothing_

Since I'm matching the field "name", using a tokenizer "keyword", I would  
expect to get no match for "carex f".

Am I wrong?

On Wednesday, July 17, 2013 8:20:28 AM UTC-4, Luca Cavanna wrote:

> Hi Christian,  
> I just tested it with elasticsearch 0.90.2 and I got the expected results:  
> 1 results querying for 'carex', one querying for 'carex feta'.
> 
> Your weird results seem to be caused by indexing ngrams, since you get  
> back partial results. On the other hand the recreation that you posted  
> works fine. I would check again the field you're querying on and your  
> mapping, maybe you're doing something slightly different from your curl  
> example?
> 
> On Tuesday, July 16, 2013 6:21:59 PM UTC+2, Christian Gendreau wrote:
> 
> > Hi,
> > 
> > I'm getting strange results trying to search on an index with 2 types  
> > using a keyword tokenizer.
> > 
> > Using Elasticsearch 0.90.2 this :  
> > curl -XGET localhost:9200/myindex/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> > f"}}}'  
> > Returns a result containing "Carex" alone (unexpected behaviour)
> > 
> > curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> > f"}}}'  
> > Will return an expected results of "Carex feta" and not "Carex" alone.
> > 
> > If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> > will return the expected results. This could be related to different  
> > configuration but I am using the default configurations on both versions.
> > 
> > So, I would like to know what is the Elasticsearch expected behavior for  
> > the first query?  
> > Could it be related to ES using a default tokenizer (Standard) when we  
> > use multiple types?
> > 
> > Here are the current settings:
> > 
> > curl -XPOST "localhost:9200/myindex" -d '  
> > {  
> > "settings":{  
> > "index":{  
> > "analysis":{  
> > "filter" : {  
> > "name\_nGram" : {  
> > "max\_gram" : 100,  
> > "min\_gram" : 2,  
> > "type" : "edge\_ngram"  
> > }  
> > },  
> > "analyzer":{  
> > "name\_index" : {  
> > "filter" : [  
> > "lowercase","asciifolding","name\_nGram"  
> > ],  
> > "tokenizer" : "standard"  
> > },  
> > "full\_name\_index" : {  
> > "filter" : [  
> > "lowercase","asciifolding"  
> > ],  
> > "tokenizer" : "keyword"  
> > },  
> > "scientificname\_index" : {  
> > "filter" : [  
> > "lowercase","asciifolding","name\_nGram"  
> > ],  
> > "tokenizer" : "keyword"  
> > },  
> > "name\_search" : {  
> > "filter" : [  
> > "lowercase","asciifolding"  
> > ],  
> > "tokenizer" : "keyword"  
> > }  
> > }  
> > }  
> > }  
> > },  
> > "mappings" : {  
> > "taxon" : {  
> > "properties" : {  
> > "name" : {  
> > "type" : "multi\_field",  
> > "fields":{  
> > "name":{  
> > "type" : "string",  
> > "index\_analyzer" : "full\_name\_index",  
> > "search\_analyzer" : "name\_search"  
> > },  
> > "ngrams":{  
> > "type" : "string",  
> > "index\_analyzer" : "scientificname\_index",  
> > "search\_analyzer" : "name\_search"  
> > }  
> > }  
> > },  
> > "status":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "namehtml":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "namehtmlauthor":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "rankname":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "parentid":{  
> > "index" : "not\_analyzed",  
> > "type" : "integer"  
> > },  
> > "parentnamehtml":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > }  
> > }  
> > },  
> > "vernacular" : {  
> > "properties" : {  
> > "name" : {  
> > "type" : "multi\_field",  
> > "fields":{  
> > "name":{  
> > "type" : "string",  
> > "index" : "not\_analyzed"  
> > },  
> > "ngrams":{  
> > "type" : "string",  
> > "search\_analyzer" : "name\_search",  
> > "index\_analyzer" : "name\_index"  
> > }  
> > }  
> > },  
> > "taxonid":{  
> > "index" : "not\_analyzed",  
> > "type" : "integer"  
> > },  
> > "status":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "language":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "taxonnamehtml":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > }  
> > }  
> > }  
> > }  
> > }'
> > 
> > Add some data:
> > 
> > curl -XPUT '[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)' -d '{  
> > "name" : "carex"  
> > }'  
> > curl -XPUT '[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)' -d '{  
> > "name" : "carex feta"  
> > }'

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [July 17, 2013, 3:41pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/4 "2013-07-17T15:41:36Z")

</div>

Hi Christian,  
exactly you can't find partial matches when indexing with keyword  
tokenizer. Only querying for 'carex' or 'carex feta' would find a match  
since that's what you indexed, with no tokenization.  
I didn't get whether you're still having an unexpected behaviour or not  
though. Could you please clarify that?

Cheers  
Luca

On Wed, Jul 17, 2013 at 5:36 PM, Christian Gendreau \<  
[christiangendreau@gmail.com](mailto:christiangendreau@gmail.com)\> wrote:

> Hi Luca,
> 
> Sorry there is an error in my post:  
> curl -XGET localhost:9200/myindex/taxon/\_\*\*search?pretty=1 -d  
> '{"query":{"match":{"name":"\*\*carex f"}}}'  
> Will return an expected result of : _nothing_
> 
> Since I'm matching the field "name", using a tokenizer "keyword", I would  
> expect to get no match for "carex f".
> 
> Am I wrong?
> 
> On Wednesday, July 17, 2013 8:20:28 AM UTC-4, Luca Cavanna wrote:
> 
> > Hi Christian,  
> > I just tested it with elasticsearch 0.90.2 and I got the expected  
> > results: 1 results querying for 'carex', one querying for 'carex feta'.
> > 
> > Your weird results seem to be caused by indexing ngrams, since you get  
> > back partial results. On the other hand the recreation that you posted  
> > works fine. I would check again the field you're querying on and your  
> > mapping, maybe you're doing something slightly different from your curl  
> > example?
> > 
> > On Tuesday, July 16, 2013 6:21:59 PM UTC+2, Christian Gendreau wrote:
> > 
> > > Hi,
> > > 
> > > I'm getting strange results trying to search on an index with 2 types  
> > > using a keyword tokenizer.
> > > 
> > > Using Elasticsearch 0.90.2 this :  
> > > curl -XGET localhost:9200/myindex/\_search\*\*?pretty=1 -d  
> > > '{"query":{"match":{"name":"\*\*carex f"}}}'  
> > > Returns a result containing "Carex" alone (unexpected behaviour)
> > > 
> > > curl -XGET localhost:9200/myindex/taxon/\_\*\*search?pretty=1 -d  
> > > '{"query":{"match":{"name":"\*\*carex f"}}}'  
> > > Will return an expected results of "Carex feta" and not "Carex" alone.
> > > 
> > > If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> > > will return the expected results. This could be related to different  
> > > configuration but I am using the default configurations on both versions.
> > > 
> > > So, I would like to know what is the Elasticsearch expected behavior for  
> > > the first query?  
> > > Could it be related to ES using a default tokenizer (Standard) when we  
> > > use multiple types?
> > > 
> > > Here are the current settings:
> > > 
> > > curl -XPOST "localhost:9200/myindex" -d '  
> > > {  
> > > "settings":{  
> > > "index":{  
> > > "analysis":{  
> > > "filter" : {  
> > > "name\_nGram" : {  
> > > "max\_gram" : 100,  
> > > "min\_gram" : 2,  
> > > "type" : "edge\_ngram"  
> > > }  
> > > },  
> > > "analyzer":{  
> > > "name\_index" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding","\*\*name\_nGram"  
> > > ],  
> > > "tokenizer" : "standard"  
> > > },  
> > > "full\_name\_index" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding"  
> > > ],  
> > > "tokenizer" : "keyword"  
> > > },  
> > > "scientificname\_index" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding","\*\*name\_nGram"  
> > > ],  
> > > "tokenizer" : "keyword"  
> > > },  
> > > "name\_search" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding"  
> > > ],  
> > > "tokenizer" : "keyword"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > },  
> > > "mappings" : {  
> > > "taxon" : {  
> > > "properties" : {  
> > > "name" : {  
> > > "type" : "multi\_field",  
> > > "fields":{  
> > > "name":{  
> > > "type" : "string",  
> > > "index\_analyzer" : "full\_name\_index",  
> > > "search\_analyzer" : "name\_search"  
> > > },  
> > > "ngrams":{  
> > > "type" : "string",  
> > > "index\_analyzer" : "scientificname\_index",  
> > > "search\_analyzer" : "name\_search"  
> > > }  
> > > }  
> > > },  
> > > "status":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "namehtml":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "namehtmlauthor":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "rankname":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "parentid":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "integer"  
> > > },  
> > > "parentnamehtml":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > }  
> > > }  
> > > },  
> > > "vernacular" : {  
> > > "properties" : {  
> > > "name" : {  
> > > "type" : "multi\_field",  
> > > "fields":{  
> > > "name":{  
> > > "type" : "string",  
> > > "index" : "not\_analyzed"  
> > > },  
> > > "ngrams":{  
> > > "type" : "string",  
> > > "search\_analyzer" : "name\_search",  
> > > "index\_analyzer" : "name\_index"  
> > > }  
> > > }  
> > > },  
> > > "taxonid":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "integer"  
> > > },  
> > > "status":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "language":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "taxonnamehtml":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }'
> > > 
> > > Add some data:
> > > 
> > > curl -XPUT '[http://localhost:9200/\*\*myindex/taxon/1](http://localhost:9200/**myindex/taxon/1)[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)  
> > > ' -d '{  
> > > "name" : "carex"  
> > > }'  
> > > curl -XPUT '[http://localhost:9200/\*\*myindex/taxon/2](http://localhost:9200/**myindex/taxon/2)[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)  
> > > ' -d '{  
> > > "name" : "carex feta"  
> > > }'
> > > 
> > > --  
> > > You received this message because you are subscribed to a topic in the  
> > > Google Groups "elasticsearch" group.  
> > > To unsubscribe from this topic, visit  
> > > [https://groups.google.com/d/topic/elasticsearch/8EWZmY\_PZyE/unsubscribe](https://groups.google.com/d/topic/elasticsearch/8EWZmY_PZyE/unsubscribe).  
> > > To unsubscribe from this group and all its topics, send an email to  
> > > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 18, 2013, 3:44pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/5 "2013-07-18T15:44:13Z")

</div>

Indeed, I didn't make it clear.  
I rebuilt the whole index from scratch and it is working properly. This is  
a little bit scary since the response of "...\_mapping" was identical on  
both installations but I guess I made a mistake somewhere.  
Anyway, thank you for confirming that it should work.

The remaining issue (that I'm solving here by keeping the untouched field)  
I have is already in another thread  
([http://elasticsearch-users.115913.n3.nabble.com/boosting-exact-matches-in-edgengram-search-td4035244.html](http://elasticsearch-users.115913.n3.nabble.com/boosting-exact-matches-in-edgengram-search-td4035244.html)).

Thanks again

Christian

On Wednesday, July 17, 2013 11:41:36 AM UTC-4, Luca Cavanna wrote:

> Hi Christian,  
> exactly you can't find partial matches when indexing with keyword  
> tokenizer. Only querying for 'carex' or 'carex feta' would find a match  
> since that's what you indexed, with no tokenization.  
> I didn't get whether you're still having an unexpected behaviour or not  
> though. Could you please clarify that?
> 
> Cheers  
> Luca
> 
> On Wed, Jul 17, 2013 at 5:36 PM, Christian Gendreau \<[christia...@gmail.com](mailto:christia...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > Hi Luca,
> > 
> > Sorry there is an error in my post:  
> > curl -XGET localhost:9200/myindex/taxon/\_\*\*search?pretty=1 -d  
> > '{"query":{"match":{"name":"\*\*carex f"}}}'  
> > Will return an expected result of : _nothing_
> > 
> > Since I'm matching the field "name", using a tokenizer "keyword", I would  
> > expect to get no match for "carex f".
> > 
> > Am I wrong?
> > 
> > On Wednesday, July 17, 2013 8:20:28 AM UTC-4, Luca Cavanna wrote:
> > 
> > > Hi Christian,  
> > > I just tested it with elasticsearch 0.90.2 and I got the expected  
> > > results: 1 results querying for 'carex', one querying for 'carex feta'.
> > > 
> > > Your weird results seem to be caused by indexing ngrams, since you get  
> > > back partial results. On the other hand the recreation that you posted  
> > > works fine. I would check again the field you're querying on and your  
> > > mapping, maybe you're doing something slightly different from your curl  
> > > example?
> > > 
> > > On Tuesday, July 16, 2013 6:21:59 PM UTC+2, Christian Gendreau wrote:
> > > 
> > > > Hi,
> > > > 
> > > > I'm getting strange results trying to search on an index with 2 types  
> > > > using a keyword tokenizer.
> > > > 
> > > > Using Elasticsearch 0.90.2 this :  
> > > > curl -XGET localhost:9200/myindex/\_search\*\*?pretty=1 -d  
> > > > '{"query":{"match":{"name":"\*\*carex f"}}}'  
> > > > Returns a result containing "Carex" alone (unexpected behaviour)
> > > > 
> > > > curl -XGET localhost:9200/myindex/taxon/\_\*\*search?pretty=1 -d  
> > > > '{"query":{"match":{"name":"\*\*carex f"}}}'  
> > > > Will return an expected results of "Carex feta" and not "Carex" alone.
> > > > 
> > > > If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> > > > will return the expected results. This could be related to different  
> > > > configuration but I am using the default configurations on both versions.
> > > > 
> > > > So, I would like to know what is the Elasticsearch expected behavior  
> > > > for the first query?  
> > > > Could it be related to ES using a default tokenizer (Standard) when we  
> > > > use multiple types?
> > > > 
> > > > Here are the current settings:
> > > > 
> > > > curl -XPOST "localhost:9200/myindex" -d '  
> > > > {  
> > > > "settings":{  
> > > > "index":{  
> > > > "analysis":{  
> > > > "filter" : {  
> > > > "name\_nGram" : {  
> > > > "max\_gram" : 100,  
> > > > "min\_gram" : 2,  
> > > > "type" : "edge\_ngram"  
> > > > }  
> > > > },  
> > > > "analyzer":{  
> > > > "name\_index" : {  
> > > > "filter" : [  
> > > > "lowercase","asciifolding","\*\*name\_nGram"  
> > > > ],  
> > > > "tokenizer" : "standard"  
> > > > },  
> > > > "full\_name\_index" : {  
> > > > "filter" : [  
> > > > "lowercase","asciifolding"  
> > > > ],  
> > > > "tokenizer" : "keyword"  
> > > > },  
> > > > "scientificname\_index" : {  
> > > > "filter" : [  
> > > > "lowercase","asciifolding","\*\*name\_nGram"  
> > > > ],  
> > > > "tokenizer" : "keyword"  
> > > > },  
> > > > "name\_search" : {  
> > > > "filter" : [  
> > > > "lowercase","asciifolding"  
> > > > ],  
> > > > "tokenizer" : "keyword"  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > },  
> > > > "mappings" : {  
> > > > "taxon" : {  
> > > > "properties" : {  
> > > > "name" : {  
> > > > "type" : "multi\_field",  
> > > > "fields":{  
> > > > "name":{  
> > > > "type" : "string",  
> > > > "index\_analyzer" : "full\_name\_index",  
> > > > "search\_analyzer" : "name\_search"  
> > > > },  
> > > > "ngrams":{  
> > > > "type" : "string",  
> > > > "index\_analyzer" : "scientificname\_index",  
> > > > "search\_analyzer" : "name\_search"  
> > > > }  
> > > > }  
> > > > },  
> > > > "status":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > },  
> > > > "namehtml":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > },  
> > > > "namehtmlauthor":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > },  
> > > > "rankname":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > },  
> > > > "parentid":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "integer"  
> > > > },  
> > > > "parentnamehtml":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > }  
> > > > }  
> > > > },  
> > > > "vernacular" : {  
> > > > "properties" : {  
> > > > "name" : {  
> > > > "type" : "multi\_field",  
> > > > "fields":{  
> > > > "name":{  
> > > > "type" : "string",  
> > > > "index" : "not\_analyzed"  
> > > > },  
> > > > "ngrams":{  
> > > > "type" : "string",  
> > > > "search\_analyzer" : "name\_search",  
> > > > "index\_analyzer" : "name\_index"  
> > > > }  
> > > > }  
> > > > },  
> > > > "taxonid":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "integer"  
> > > > },  
> > > > "status":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > },  
> > > > "language":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > },  
> > > > "taxonnamehtml":{  
> > > > "index" : "not\_analyzed",  
> > > > "type" : "string"  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }'
> > > > 
> > > > Add some data:
> > > > 
> > > > curl -XPUT '[http://localhost:9200/\*\*myindex/taxon/1](http://localhost:9200/**myindex/taxon/1)[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)  
> > > > ' -d '{  
> > > > "name" : "carex"  
> > > > }'  
> > > > curl -XPUT '[http://localhost:9200/\*\*myindex/taxon/2](http://localhost:9200/**myindex/taxon/2)[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)  
> > > > ' -d '{  
> > > > "name" : "carex feta"  
> > > > }'
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to a topic in the  
> > > > Google Groups "elasticsearch" group.  
> > > > To unsubscribe from this topic, visit  
> > > > [https://groups.google.com/d/topic/elasticsearch/8EWZmY\_PZyE/unsubscribe](https://groups.google.com/d/topic/elasticsearch/8EWZmY_PZyE/unsubscribe).  
> > > > To unsubscribe from this group and all its topics, send an email to  
> > > > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 19, 2013, 7:34pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/6 "2013-07-19T19:34:35Z")

</div>

I still have this issue and I can now reproduce it:

Settings and mapping  
curl -XPOST "localhost:9200/myindex" -d '  
{  
"settings":{  
"index":{  
"analysis":{  
"filter" : {  
"name\_nGram" : {  
"max\_gram" : 100,  
"min\_gram" : 2,  
"type" : "edge\_ngram"  
},  
"strip\_hybrid\_sign\_filter":{  
"pattern":"\u00D7",  
"replacement":"",  
"type": "pattern\_replace"  
}  
},  
"analyzer":{  
"name\_index" : {  
"filter" : [  
"lowercase","asciifolding","name\_nGram"  
],  
"tokenizer" : "standard"  
},  
"full\_name\_index" : {  
"filter" : [  
"lowercase","asciifolding"  
],  
"tokenizer" : "keyword"  
},  
"scientificname\_index" : {  
"filter" : [  
"lowercase","asciifolding","strip\_hybrid\_sign\_filter","name\_nGram"  
],  
"tokenizer" : "keyword"  
},  
"name\_search" : {  
"filter" : [  
"lowercase","asciifolding"  
],  
"tokenizer" : "keyword"  
},  
"scientificname\_search" : {  
"filter" : [  
"lowercase","asciifolding","strip\_hybrid\_sign\_filter"  
],  
"tokenizer" : "keyword"  
}  
}  
}  
}  
},  
"mappings" : {  
"taxon" : {  
"properties" : {  
"name" : {  
"type" : "multi\_field",  
"fields":{  
"name":{  
"type" : "string",  
"index\_analyzer" : "full\_name\_index",  
"search\_analyzer" : "name\_search"  
},  
"ngrams":{  
"type" : "string",  
"index\_analyzer" : "scientificname\_index",  
"search\_analyzer" : "scientificname\_search"  
}  
}  
},  
"status":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"namehtml":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"namehtmlauthor":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"rankname":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"parentid":{  
"index" : "not\_analyzed",  
"type" : "integer"  
},  
"parentnamehtml":{  
"index" : "not\_analyzed",  
"type" : "string"  
}  
}  
},  
"vernacular" : {  
"properties" : {  
"name" : {  
"type" : "multi\_field",  
"fields":{  
"name":{  
"type" : "string",  
"index" : "not\_analyzed"  
},  
"ngrams":{  
"type" : "string",  
"search\_analyzer" : "name\_search",  
"index\_analyzer" : "name\_index"  
}  
}  
},  
"taxonid":{  
"index" : "not\_analyzed",  
"type" : "integer"  
},  
"status":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"language":{  
"index" : "not\_analyzed",  
"type" : "string"  
},  
"taxonnamehtml":{  
"index" : "not\_analyzed",  
"type" : "string"  
}  
}  
}  
}  
}'

Add data:  
curl -XPUT '[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)' -d '{  
"name" : "×Achnella"  
}'  
curl -XPUT '[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)' -d '{  
"name" : "carex feta"  
}'

I get no result for this query:  
curl -XGET localhost:9200/myindex/\_search?pretty=1 -d  
'{"query":{"match":{"name":"×Achnella"}}}'

But I do get the row if I use this query:  
curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d  
'{"query":{"match":{"name":"×Achnella"}}}'

Maybe it's related to my "strip\_hybrid\_sign\_filter"?  
But the result of this command looks good:  
curl -XGET'localhost:9200/myindex/\_analyze?field=name&pretty=1' -d  
"×Achnella"

Any idea?

Thanks

On Tuesday, July 16, 2013 12:21:59 PM UTC-4, Christian Gendreau wrote:

> Hi,
> 
> I'm getting strange results trying to search on an index with 2 types  
> using a keyword tokenizer.
> 
> Using Elasticsearch 0.90.2 this :  
> curl -XGET localhost:9200/myindex/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> f"}}}'  
> Returns a result containing "Carex" alone (unexpected behaviour)
> 
> curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> f"}}}'  
> Will return an expected results of "Carex feta" and not "Carex" alone.
> 
> If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> will return the expected results. This could be related to different  
> configuration but I am using the default configurations on both versions.
> 
> So, I would like to know what is the Elasticsearch expected behavior for  
> the first query?  
> Could it be related to ES using a default tokenizer (Standard) when we use  
> multiple types?
> 
> Here are the current settings:
> 
> curl -XPOST "localhost:9200/myindex" -d '  
> {  
> "settings":{  
> "index":{  
> "analysis":{  
> "filter" : {  
> "name\_nGram" : {  
> "max\_gram" : 100,  
> "min\_gram" : 2,  
> "type" : "edge\_ngram"  
> }  
> },  
> "analyzer":{  
> "name\_index" : {  
> "filter" : [  
> "lowercase","asciifolding","name\_nGram"  
> ],  
> "tokenizer" : "standard"  
> },  
> "full\_name\_index" : {  
> "filter" : [  
> "lowercase","asciifolding"  
> ],  
> "tokenizer" : "keyword"  
> },  
> "scientificname\_index" : {  
> "filter" : [  
> "lowercase","asciifolding","name\_nGram"  
> ],  
> "tokenizer" : "keyword"  
> },  
> "name\_search" : {  
> "filter" : [  
> "lowercase","asciifolding"  
> ],  
> "tokenizer" : "keyword"  
> }  
> }  
> }  
> }  
> },  
> "mappings" : {  
> "taxon" : {  
> "properties" : {  
> "name" : {  
> "type" : "multi\_field",  
> "fields":{  
> "name":{  
> "type" : "string",  
> "index\_analyzer" : "full\_name\_index",  
> "search\_analyzer" : "name\_search"  
> },  
> "ngrams":{  
> "type" : "string",  
> "index\_analyzer" : "scientificname\_index",  
> "search\_analyzer" : "name\_search"  
> }  
> }  
> },  
> "status":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "namehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "namehtmlauthor":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "rankname":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "parentid":{  
> "index" : "not\_analyzed",  
> "type" : "integer"  
> },  
> "parentnamehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> }  
> }  
> },  
> "vernacular" : {  
> "properties" : {  
> "name" : {  
> "type" : "multi\_field",  
> "fields":{  
> "name":{  
> "type" : "string",  
> "index" : "not\_analyzed"  
> },  
> "ngrams":{  
> "type" : "string",  
> "search\_analyzer" : "name\_search",  
> "index\_analyzer" : "name\_index"  
> }  
> }  
> },  
> "taxonid":{  
> "index" : "not\_analyzed",  
> "type" : "integer"  
> },  
> "status":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "language":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "taxonnamehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> }  
> }  
> }  
> }  
> }'
> 
> Add some data:
> 
> curl -XPUT '[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)' -d '{  
> "name" : "carex"  
> }'  
> curl -XPUT '[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)' -d '{  
> "name" : "carex feta"  
> }'

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [July 22, 2013, 2:52pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/7 "2013-07-22T14:52:39Z")

</div>

Hi Christian,  
I was able to reproduce your issue.

The mapping for the field name is different in the two types that you have.  
You used once keyword tokenizer + lowercase filter etc., while on the other  
one the field is not analyzed.

Elasticsearch has to pick an analyzer here for the query, and it picks the  
wrong one in your case unfortunately. In fact it ends up not analyzing the  
query and querying for exactly the same term you use in the query, while in  
the index you have the lowercased version, thus there's no match.

When you specify the type the problem doesn't exist since there's only one  
field called name under that type and there's only one analyzer, thus no  
choice to be made.

I think it was just an error in your mapping, but if you do want to have  
fields with same name and different analyzers under the same index, well  
that's not a good idea. You'd better go for two separate indices since the  
query would be analyzed differently then per index.

Hope this clarifies things for you

Cheers  
Luca

On Tuesday, July 16, 2013 6:21:59 PM UTC+2, Christian Gendreau wrote:

> Hi,
> 
> I'm getting strange results trying to search on an index with 2 types  
> using a keyword tokenizer.
> 
> Using Elasticsearch 0.90.2 this :  
> curl -XGET localhost:9200/myindex/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> f"}}}'  
> Returns a result containing "Carex" alone (unexpected behaviour)
> 
> curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> f"}}}'  
> Will return an expected results of "Carex feta" and not "Carex" alone.
> 
> If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> will return the expected results. This could be related to different  
> configuration but I am using the default configurations on both versions.
> 
> So, I would like to know what is the Elasticsearch expected behavior for  
> the first query?  
> Could it be related to ES using a default tokenizer (Standard) when we use  
> multiple types?
> 
> Here are the current settings:
> 
> curl -XPOST "localhost:9200/myindex" -d '  
> {  
> "settings":{  
> "index":{  
> "analysis":{  
> "filter" : {  
> "name\_nGram" : {  
> "max\_gram" : 100,  
> "min\_gram" : 2,  
> "type" : "edge\_ngram"  
> }  
> },  
> "analyzer":{  
> "name\_index" : {  
> "filter" : [  
> "lowercase","asciifolding","name\_nGram"  
> ],  
> "tokenizer" : "standard"  
> },  
> "full\_name\_index" : {  
> "filter" : [  
> "lowercase","asciifolding"  
> ],  
> "tokenizer" : "keyword"  
> },  
> "scientificname\_index" : {  
> "filter" : [  
> "lowercase","asciifolding","name\_nGram"  
> ],  
> "tokenizer" : "keyword"  
> },  
> "name\_search" : {  
> "filter" : [  
> "lowercase","asciifolding"  
> ],  
> "tokenizer" : "keyword"  
> }  
> }  
> }  
> }  
> },  
> "mappings" : {  
> "taxon" : {  
> "properties" : {  
> "name" : {  
> "type" : "multi\_field",  
> "fields":{  
> "name":{  
> "type" : "string",  
> "index\_analyzer" : "full\_name\_index",  
> "search\_analyzer" : "name\_search"  
> },  
> "ngrams":{  
> "type" : "string",  
> "index\_analyzer" : "scientificname\_index",  
> "search\_analyzer" : "name\_search"  
> }  
> }  
> },  
> "status":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "namehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "namehtmlauthor":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "rankname":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "parentid":{  
> "index" : "not\_analyzed",  
> "type" : "integer"  
> },  
> "parentnamehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> }  
> }  
> },  
> "vernacular" : {  
> "properties" : {  
> "name" : {  
> "type" : "multi\_field",  
> "fields":{  
> "name":{  
> "type" : "string",  
> "index" : "not\_analyzed"  
> },  
> "ngrams":{  
> "type" : "string",  
> "search\_analyzer" : "name\_search",  
> "index\_analyzer" : "name\_index"  
> }  
> }  
> },  
> "taxonid":{  
> "index" : "not\_analyzed",  
> "type" : "integer"  
> },  
> "status":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "language":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> },  
> "taxonnamehtml":{  
> "index" : "not\_analyzed",  
> "type" : "string"  
> }  
> }  
> }  
> }  
> }'
> 
> Add some data:
> 
> curl -XPUT '[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)' -d '{  
> "name" : "carex"  
> }'  
> curl -XPUT '[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)' -d '{  
> "name" : "carex feta"  
> }'

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![cgendreau](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cgendreau/32/2204_2.png) [@cgendreau](https://discuss.elastic.co/u/cgendreau)\
**Post date:** [July 22, 2013, 3:49pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/8 "2013-07-22T15:49:08Z")

</div>

Hi Luca,  
Thanks for the explanations, that helps a lot to understand what is going  
on.

Yes, the mapping is intentionally different for the 2 types.  
The goal was to get merged results from the 2 types, with one query, using  
the same field name. Since my 2 types are different "kind" of data, I have  
2 different analyzers.

Then, I though ES would take the analyzer mapped to the field depending on  
the type if I do not specify a specific type in my query.  
I guess this is where I was wrong.

What would be my best option?  
Considering I want to merge the result of my 2 types when I send a query,  
should I

1. rename my fields and use a multi\_match query
2. use 2 separate index and an indices query

Regards,

Christian

On Monday, July 22, 2013 10:52:39 AM UTC-4, Luca Cavanna wrote:

> Hi Christian,  
> I was able to reproduce your issue.
> 
> The mapping for the field name is different in the two types that you  
> have. You used once keyword tokenizer + lowercase filter etc., while on the  
> other one the field is not analyzed.
> 
> Elasticsearch has to pick an analyzer here for the query, and it picks the  
> wrong one in your case unfortunately. In fact it ends up not analyzing the  
> query and querying for exactly the same term you use in the query, while in  
> the index you have the lowercased version, thus there's no match.
> 
> When you specify the type the problem doesn't exist since there's only one  
> field called name under that type and there's only one analyzer, thus no  
> choice to be made.
> 
> I think it was just an error in your mapping, but if you do want to have  
> fields with same name and different analyzers under the same index, well  
> that's not a good idea. You'd better go for two separate indices since the  
> query would be analyzed differently then per index.
> 
> Hope this clarifies things for you
> 
> Cheers  
> Luca
> 
> On Tuesday, July 16, 2013 6:21:59 PM UTC+2, Christian Gendreau wrote:
> 
> > Hi,
> > 
> > I'm getting strange results trying to search on an index with 2 types  
> > using a keyword tokenizer.
> > 
> > Using Elasticsearch 0.90.2 this :  
> > curl -XGET localhost:9200/myindex/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> > f"}}}'  
> > Returns a result containing "Carex" alone (unexpected behaviour)
> > 
> > curl -XGET localhost:9200/myindex/taxon/\_search?pretty=1 -d '{"query":{"match":{"name":"carex  
> > f"}}}'  
> > Will return an expected results of "Carex feta" and not "Carex" alone.
> > 
> > If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> > will return the expected results. This could be related to different  
> > configuration but I am using the default configurations on both versions.
> > 
> > So, I would like to know what is the Elasticsearch expected behavior for  
> > the first query?  
> > Could it be related to ES using a default tokenizer (Standard) when we  
> > use multiple types?
> > 
> > Here are the current settings:
> > 
> > curl -XPOST "localhost:9200/myindex" -d '  
> > {  
> > "settings":{  
> > "index":{  
> > "analysis":{  
> > "filter" : {  
> > "name\_nGram" : {  
> > "max\_gram" : 100,  
> > "min\_gram" : 2,  
> > "type" : "edge\_ngram"  
> > }  
> > },  
> > "analyzer":{  
> > "name\_index" : {  
> > "filter" : [  
> > "lowercase","asciifolding","name\_nGram"  
> > ],  
> > "tokenizer" : "standard"  
> > },  
> > "full\_name\_index" : {  
> > "filter" : [  
> > "lowercase","asciifolding"  
> > ],  
> > "tokenizer" : "keyword"  
> > },  
> > "scientificname\_index" : {  
> > "filter" : [  
> > "lowercase","asciifolding","name\_nGram"  
> > ],  
> > "tokenizer" : "keyword"  
> > },  
> > "name\_search" : {  
> > "filter" : [  
> > "lowercase","asciifolding"  
> > ],  
> > "tokenizer" : "keyword"  
> > }  
> > }  
> > }  
> > }  
> > },  
> > "mappings" : {  
> > "taxon" : {  
> > "properties" : {  
> > "name" : {  
> > "type" : "multi\_field",  
> > "fields":{  
> > "name":{  
> > "type" : "string",  
> > "index\_analyzer" : "full\_name\_index",  
> > "search\_analyzer" : "name\_search"  
> > },  
> > "ngrams":{  
> > "type" : "string",  
> > "index\_analyzer" : "scientificname\_index",  
> > "search\_analyzer" : "name\_search"  
> > }  
> > }  
> > },  
> > "status":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "namehtml":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "namehtmlauthor":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "rankname":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "parentid":{  
> > "index" : "not\_analyzed",  
> > "type" : "integer"  
> > },  
> > "parentnamehtml":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > }  
> > }  
> > },  
> > "vernacular" : {  
> > "properties" : {  
> > "name" : {  
> > "type" : "multi\_field",  
> > "fields":{  
> > "name":{  
> > "type" : "string",  
> > "index" : "not\_analyzed"  
> > },  
> > "ngrams":{  
> > "type" : "string",  
> > "search\_analyzer" : "name\_search",  
> > "index\_analyzer" : "name\_index"  
> > }  
> > }  
> > },  
> > "taxonid":{  
> > "index" : "not\_analyzed",  
> > "type" : "integer"  
> > },  
> > "status":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "language":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > },  
> > "taxonnamehtml":{  
> > "index" : "not\_analyzed",  
> > "type" : "string"  
> > }  
> > }  
> > }  
> > }  
> > }'
> > 
> > Add some data:
> > 
> > curl -XPUT '[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)' -d '{  
> > "name" : "carex"  
> > }'  
> > curl -XPUT '[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)' -d '{  
> > "name" : "carex feta"  
> > }'

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [July 22, 2013, 3:58pm UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/9 "2013-07-22T15:58:10Z")

</div>

Hey Christian,  
ok the reason why it doesn't work under the same type is that we execute a  
single lucene query per index (which is composed of one more shards, that  
are effectivey the lucene indices). Thus making it work would mean  
executing multiple lucene queries (analyzed differently) on the same shard,  
which is not really what we want.

Both the options to solve the problem look good. I'd say it depends on how  
much it's important for you to keep the two datasets on the same index, or  
how much it bothers you to use two different names for the two fields.  
Don't know enough about your domain to make any choice but both ways are  
fine.

Cheers  
Luca

On Mon, Jul 22, 2013 at 5:49 PM, Christian Gendreau \<  
[christiangendreau@gmail.com](mailto:christiangendreau@gmail.com)\> wrote:

> Hi Luca,  
> Thanks for the explanations, that helps a lot to understand what is going  
> on.
> 
> Yes, the mapping is intentionally different for the 2 types.  
> The goal was to get merged results from the 2 types, with one query, using  
> the same field name. Since my 2 types are different "kind" of data, I have  
> 2 different analyzers.
> 
> Then, I though ES would take the analyzer mapped to the field depending on  
> the type if I do not specify a specific type in my query.  
> I guess this is where I was wrong.
> 
> What would be my best option?  
> Considering I want to merge the result of my 2 types when I send a query,  
> should I
> 
> 1. rename my fields and use a multi\_match query
> 2. use 2 separate index and an indices query
> 
> Regards,
> 
> Christian
> 
> On Monday, July 22, 2013 10:52:39 AM UTC-4, Luca Cavanna wrote:
> 
> > Hi Christian,  
> > I was able to reproduce your issue.
> > 
> > The mapping for the field name is different in the two types that you  
> > have. You used once keyword tokenizer + lowercase filter etc., while on the  
> > other one the field is not analyzed.
> > 
> > Elasticsearch has to pick an analyzer here for the query, and it picks  
> > the wrong one in your case unfortunately. In fact it ends up not analyzing  
> > the query and querying for exactly the same term you use in the query,  
> > while in the index you have the lowercased version, thus there's no match.
> > 
> > When you specify the type the problem doesn't exist since there's only  
> > one field called name under that type and there's only one analyzer, thus  
> > no choice to be made.
> > 
> > I think it was just an error in your mapping, but if you do want to have  
> > fields with same name and different analyzers under the same index, well  
> > that's not a good idea. You'd better go for two separate indices since the  
> > query would be analyzed differently then per index.
> > 
> > Hope this clarifies things for you
> > 
> > Cheers  
> > Luca
> > 
> > On Tuesday, July 16, 2013 6:21:59 PM UTC+2, Christian Gendreau wrote:
> > 
> > > Hi,
> > > 
> > > I'm getting strange results trying to search on an index with 2 types  
> > > using a keyword tokenizer.
> > > 
> > > Using Elasticsearch 0.90.2 this :  
> > > curl -XGET localhost:9200/myindex/\_search\*\*?pretty=1 -d  
> > > '{"query":{"match":{"name":"\*\*carex f"}}}'  
> > > Returns a result containing "Carex" alone (unexpected behaviour)
> > > 
> > > curl -XGET localhost:9200/myindex/taxon/\_\*\*search?pretty=1 -d  
> > > '{"query":{"match":{"name":"\*\*carex f"}}}'  
> > > Will return an expected results of "Carex feta" and not "Carex" alone.
> > > 
> > > If I do the same thing using Elasticsearch 0.90.1, the 2 queries above  
> > > will return the expected results. This could be related to different  
> > > configuration but I am using the default configurations on both versions.
> > > 
> > > So, I would like to know what is the Elasticsearch expected behavior for  
> > > the first query?  
> > > Could it be related to ES using a default tokenizer (Standard) when we  
> > > use multiple types?
> > > 
> > > Here are the current settings:
> > > 
> > > curl -XPOST "localhost:9200/myindex" -d '  
> > > {  
> > > "settings":{  
> > > "index":{  
> > > "analysis":{  
> > > "filter" : {  
> > > "name\_nGram" : {  
> > > "max\_gram" : 100,  
> > > "min\_gram" : 2,  
> > > "type" : "edge\_ngram"  
> > > }  
> > > },  
> > > "analyzer":{  
> > > "name\_index" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding","\*\*name\_nGram"  
> > > ],  
> > > "tokenizer" : "standard"  
> > > },  
> > > "full\_name\_index" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding"  
> > > ],  
> > > "tokenizer" : "keyword"  
> > > },  
> > > "scientificname\_index" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding","\*\*name\_nGram"  
> > > ],  
> > > "tokenizer" : "keyword"  
> > > },  
> > > "name\_search" : {  
> > > "filter" : [  
> > > "lowercase","asciifolding"  
> > > ],  
> > > "tokenizer" : "keyword"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > },  
> > > "mappings" : {  
> > > "taxon" : {  
> > > "properties" : {  
> > > "name" : {  
> > > "type" : "multi\_field",  
> > > "fields":{  
> > > "name":{  
> > > "type" : "string",  
> > > "index\_analyzer" : "full\_name\_index",  
> > > "search\_analyzer" : "name\_search"  
> > > },  
> > > "ngrams":{  
> > > "type" : "string",  
> > > "index\_analyzer" : "scientificname\_index",  
> > > "search\_analyzer" : "name\_search"  
> > > }  
> > > }  
> > > },  
> > > "status":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "namehtml":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "namehtmlauthor":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "rankname":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "parentid":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "integer"  
> > > },  
> > > "parentnamehtml":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > }  
> > > }  
> > > },  
> > > "vernacular" : {  
> > > "properties" : {  
> > > "name" : {  
> > > "type" : "multi\_field",  
> > > "fields":{  
> > > "name":{  
> > > "type" : "string",  
> > > "index" : "not\_analyzed"  
> > > },  
> > > "ngrams":{  
> > > "type" : "string",  
> > > "search\_analyzer" : "name\_search",  
> > > "index\_analyzer" : "name\_index"  
> > > }  
> > > }  
> > > },  
> > > "taxonid":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "integer"  
> > > },  
> > > "status":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "language":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > },  
> > > "taxonnamehtml":{  
> > > "index" : "not\_analyzed",  
> > > "type" : "string"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }'
> > > 
> > > Add some data:
> > > 
> > > curl -XPUT '[http://localhost:9200/\*\*myindex/taxon/1](http://localhost:9200/**myindex/taxon/1)[http://localhost:9200/myindex/taxon/1](http://localhost:9200/myindex/taxon/1)  
> > > ' -d '{  
> > > "name" : "carex"  
> > > }'  
> > > curl -XPUT '[http://localhost:9200/\*\*myindex/taxon/2](http://localhost:9200/**myindex/taxon/2)[http://localhost:9200/myindex/taxon/2](http://localhost:9200/myindex/taxon/2)  
> > > ' -d '{  
> > > "name" : "carex feta"  
> > > }'
> > > 
> > > --  
> > > You received this message because you are subscribed to a topic in the  
> > > Google Groups "elasticsearch" group.  
> > > To unsubscribe from this topic, visit  
> > > [https://groups.google.com/d/topic/elasticsearch/8EWZmY\_PZyE/unsubscribe](https://groups.google.com/d/topic/elasticsearch/8EWZmY_PZyE/unsubscribe).  
> > > To unsubscribe from this group and all its topics, send an email to  
> > > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:25am UTC](https://discuss.elastic.co/t/searching-an-index-with-2-types-using-a-keyword-tokenizer/12811/10 "2017-07-06T02:25:10Z")

</div>


