# Confused about query\_string and the use of wildcards

**URL:** <https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103>\
**Category:** Elasticsearch\
**Created:** [March 15, 2011, 7:52pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103 "2011-03-15T19:52:52Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 15, 2011, 7:52pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/1 "2011-03-15T19:52:52Z")

</div>

Hi,

I'm struggling on why a simple query like this one:

"query": {  
"query\_string": {  
"default\_operator": "AND",  
"query": "\*phone"  
}  
}

does not return any results, whereas this one (notice the extra '\*' as a  
suffix to the query):

"query": {  
"query\_string": {  
"default\_operator": "AND",  
"query": "_phone_"  
}  
}

does return some results. But cannot understand why, to be honest...

Thanks.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 15, 2011, 10:56pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/2 "2011-03-15T22:56:38Z")

</div>

Maybe because you have terms that don't end with phone?  
On Tuesday, March 15, 2011 at 9:52 PM, Enrique Medina Montenegro wrote:

> Hi,
> 
> I'm struggling on why a simple query like this one:
> 
> "query": {  
> "query\_string": {  
> "default\_operator": "AND",  
> "query": "\*phone"  
> }  
> }
> 
> does not return any results, whereas this one (notice the extra '\*' as a suffix to the query):
> 
> "query": {  
> "query\_string": {  
> "default\_operator": "AND",  
> "query": "_phone_"  
> }  
> }
> 
> does return some results. But cannot understand why, to be honest...
> 
> Thanks.

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 9:08am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/3 "2011-03-16T09:08:46Z")

</div>

Shay,

But what about iPhone? Shouldn't it be included as part of the results for  
"\*phone"?

Or maybe it's just that the '\*' doesn't work here as a real wildcard, as in  
SQL a '%'?

Thanks.

On Tue, Mar 15, 2011 at 11:56 PM, Shay Banon  
[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Maybe because you have terms that don't end with phone?
> 
> On Tuesday, March 15, 2011 at 9:52 PM, Enrique Medina Montenegro wrote:
> 
> Hi,
> 
> I'm struggling on why a simple query like this one:
> 
> "query": {  
> "query\_string": {  
> "default\_operator": "AND",  
> "query": "\*phone"  
> }  
> }
> 
> does not return any results, whereas this one (notice the extra '\*' as a  
> suffix to the query):
> 
> "query": {  
> "query\_string": {  
> "default\_operator": "AND",  
> "query": "_phone_"  
> }  
> }
> 
> does return some results. But cannot understand why, to be honest...
> 
> Thanks.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 16, 2011, 10:02am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/4 "2011-03-16T10:02:58Z")

</div>

Hi Enrique

> But what about iPhone? Shouldn't it be included as part of the results  
> for "\*phone"?

> Or maybe it's just that the '\*' doesn't work here as a real wildcard,  
> as in SQL a '%'?

It is the same as % in SQL, and your example works for me.

I suggest you gist a complete curl recreation, from index creation, data  
indexing, and searching to demonstrate the problem.

clint

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 10:36am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/5 "2011-03-16T10:36:18Z")

</div>

Clinton,

I found the issue and it was on my default Spanish analyzer. For some  
reason, iPhone gets analyzed in Spanish like this:

[http://localhost:9200/mytest/\_analyze?text=iPhone+4](http://localhost:9200/mytest/_analyze?text=iPhone+4)

{"tokens":[{"token":"iphon","start\_offset":0,"end\_offset":6,"type":"","position":1},{"token":"4","start\_offset":7,"end\_offset":8,"type":"","position":2}]}

whereas in the case of the default analyzer it gets like this:

[http://localhost:9200/mytest/\_analyze?text=iPhone+4&analyzer=standard](http://localhost:9200/mytest/_analyze?text=iPhone+4&analyzer=standard)

{"tokens":[{"token":"iphone","start\_offset":0,"end\_offset":6,"type":"","position":1},{"token":"4","start\_offset":7,"end\_offset":8,"type":"","position":2}]}

Hence, the token "iphon" in Spanish was not matching the "_phone_", but  
matches "_phon_".

What do you recommend in these particular cases? Adding iPhone as a stop  
word?

Thanks.

On Wed, Mar 16, 2011 at 11:02 AM, Clinton Gormley  
[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> Hi Enrique
> 
> > But what about iPhone? Shouldn't it be included as part of the results  
> > for "\*phone"?
> 
> > Or maybe it's just that the '\*' doesn't work here as a real wildcard,  
> > as in SQL a '%'?
> 
> It is the same as % in SQL, and your example works for me.
> 
> I suggest you gist a complete curl recreation, from index creation, data  
> indexing, and searching to demonstrate the problem.
> 
> clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 16, 2011, 10:51am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/6 "2011-03-16T10:51:17Z")

</div>

Hi Enrique

> I found the issue and it was on my default Spanish analyzer. For some  
> reason, iPhone gets analyzed in Spanish like this:
> 
> [http://localhost:9200/mytest/\_analyze?text=iPhone+4](http://localhost:9200/mytest/_analyze?text=iPhone+4)  
> {"tokens":[{"token":"iphon","start\_offset":0,"end\_offset":6,"type":"","position":1},{"token":"4","start\_offset":7,"end\_offset":8,"type":"","position":2}]}

Presumably you're using the snowball stemmer? It analyzes 'iphone' as  
'iphon' to be able to recognise eg "cansada" and "cansado" as the same  
stem.

All you need to do is to be sure that you're using the same analyzer at  
index time as at search time.

You have a few options here:

1. you're searching on a field (eg product\_name) and you set the  
'analyzer' for that field to be the spanish stemmer, when  
you put the mapping

2. you're searching on the '\_all' field (which is the default)  
and you can set the analyzer for the '\_all' field to  
be the spanish stemmer when you put the mapping

3. you can't determine at mapping time which language you're  
going to be using at search time, and you specify the  
analyzer in the query\_string query itself:

4. you could do something wizzy per document with the \_analyzer field

clint

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 11:03am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/7 "2011-03-16T11:03:25Z")

</div>

Clinton,

Based on some other discussion with Shay, I defined this in the  
elasticsearch.yml config file:

index:  
analysis:  
analyzer:  
default:  
type: es.cuestamenos.lucene.analizadores.SpanishAnalyzerProvider

And the analyzer is this one:

> <https://gist.github.com/emedina/872322>

Shouldn't that be enough both for index and search?

Thanks.

On Wed, Mar 16, 2011 at 11:51 AM, Clinton Gormley  
[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> Hi Enrique
> 
> > I found the issue and it was on my default Spanish analyzer. For some  
> > reason, iPhone gets analyzed in Spanish like this:
> > 
> > [http://localhost:9200/mytest/\_analyze?text=iPhone+4](http://localhost:9200/mytest/_analyze?text=iPhone+4)
> 
> {"tokens":[{"token":"iphon","start\_offset":0,"end\_offset":6,"type":"","position":1},{"token":"4","start\_offset":7,"end\_offset":8,"type":"","position":2}]}
> 
> Presumably you're using the snowball stemmer? It analyzes 'iphone' as  
> 'iphon' to be able to recognise eg "cansada" and "cansado" as the same  
> stem.
> 
> All you need to do is to be sure that you're using the same analyzer at  
> index time as at search time.
> 
> You have a few options here:
> 
> 1. you're searching on a field (eg product\_name) and you set the  
> 'analyzer' for that field to be the spanish stemmer, when  
> you put the mapping
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> 
> 1. you're searching on the '\_all' field (which is the default)  
> and you can set the analyzer for the '\_all' field to  
> be the spanish stemmer when you put the mapping
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/all-field.html)
> 
> (this is probably not what you want, as the \_all field will contain  
> some fields which shouldn't have the stemmer applied)
> 
> 1. you can't determine at mapping time which language you're  
> going to be using at search time, and you specify the  
> analyzer in the query\_string query itself:
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/query-string-query.html)
> 
> 1. you could do something wizzy per document with the \_analyzer field
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/analyzer-field.html)
> 
> clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 16, 2011, 11:10am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/8 "2011-03-16T11:10:21Z")

</div>

Hi Enrique

> Based on some other discussion with Shay, I defined this in the  
> elasticsearch.yml config file:
> 
> index:  
> analysis:  
> analyzer:  
> default:  
> type:  
> es.cuestamenos.lucene.analizadores.SpanishAnalyzerProvider

> Shouldn't that be enough both for index and search?

I would have thought so. But it doesn't appear to be applied at search  
time. Are you searching against a specific field, or against \_all?

If it works against a specific field, but not against \_all, then perhaps  
there is a bug.

A complete curl recreation would be useful

clint

>

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 11:23am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/9 "2011-03-16T11:23:50Z")

</div>

Clinton,

I'm searching against "\_all", which is the default.

I get consistent results (even the lack of results) when adding a specific  
field or specific analyzer:

{  
"query": {  
"query\_string": {  
"default\_operator": "AND",  
"query": "\*phon",  
"default\_field": "name",  
"analyzer": "default"  
}  
}  
}

So I guess it's not a bug, but as explained in my previous email, the fact  
that the Spanish analyzer created a token = "iphon" for iPhone so no matter  
how I search, it will never match "\*phone", right?

Regards.

On Wed, Mar 16, 2011 at 12:10 PM, Clinton Gormley  
[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> Hi Enrique
> 
> > Based on some other discussion with Shay, I defined this in the  
> > elasticsearch.yml config file:
> > 
> > index:  
> > analysis:  
> > analyzer:  
> > default:  
> > type:  
> > es.cuestamenos.lucene.analizadores.SpanishAnalyzerProvider
> 
> > Shouldn't that be enough both for index and search?
> 
> I would have thought so. But it doesn't appear to be applied at search  
> time. Are you searching against a specific field, or against \_all?
> 
> If it works against a specific field, but not against \_all, then perhaps  
> there is a bug.
> 
> A complete curl recreation would be useful
> 
> clint
> 
> >

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 16, 2011, 11:33am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/10 "2011-03-16T11:33:15Z")

</div>

hi Enrique

> 

I've just remembered your original question, which was:

```
"*phone"

```

vs  
"_phone_"

As I understand it, the way this wildcard search works is that Lucene  
looks up all matching terms, and searches against each of these.

So for some reason, "\*phone" doesn't find the the right term, but  
"_phone_" does.

> I get consistent results (even the lack of results) when adding a  
> specific field or specific analyzer:

You mean, you see the same thing?

> So I guess it's not a bug, but as explained in my previous email, the  
> fact that the Spanish analyzer created a token = "iphon" for iPhone so  
> no matter how I search, it will never match "\*phone", right?

No. This should work. For instance, using the default analyzer, if you  
index "The Quick BROWN fox" you end up with the terms  
"quick","brown","fox"

If you then search for "The Quick BROWN fox", it performs the same  
analysis, resulting in the same terms, and searches for those.

So to me (and I'm ignorant of the Lucene internals) it sounds like a  
potential bug in the lucene query parser syntax.

A complete recreation would be very useful for debugging.

clint

---

<div class="post-metadata">

**Author:** ![Joaquin\_Cuenca\_Abela](https://avatars.discourse-cdn.com/v4/letter/j/e480ec/32.png) [@Joaquin\_Cuenca\_Abela](https://discuss.elastic.co/u/Joaquin_Cuenca_Abela)\
**Post date:** [March 16, 2011, 11:52am UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/11 "2011-03-16T11:52:43Z")

</div>

I don't know why your pattern search is not working, but on this  
analyzer you're ascii folding the terms before you remove the stop  
words, and your stop words contain non ascii letters. You should  
either ascii fold your stop words or remove them before the ascii  
folding step.

Cheers,

On Wed, Mar 16, 2011 at 12:03 PM, Enrique Medina Montenegro  
[e.medina.m@gmail.com](mailto:e.medina.m@gmail.com) wrote:

> Clinton,  
> Based on some other discussion with Shay, I defined this in the  
> elasticsearch.yml config file:  
> index:  
> analysis:  
> analyzer:  
> default:  
> type: es.cuestamenos.lucene.analizadores.SpanishAnalyzerProvider  
> And the analyzer is this one:  
> [Custom analyzer for Spanish · GitHub](https://gist.github.com/872322)  
> Shouldn't that be enough both for index and search?  
> Thanks.
> 
> On Wed, Mar 16, 2011 at 11:51 AM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)  
> wrote:
> 
> > Hi Enrique
> > 
> > > I found the issue and it was on my default Spanish analyzer. For some  
> > > reason, iPhone gets analyzed in Spanish like this:
> > > 
> > > [http://localhost:9200/mytest/\_analyze?text=iPhone+4](http://localhost:9200/mytest/_analyze?text=iPhone+4)
> > > 
> > > {"tokens":[{"token":"iphon","start\_offset":0,"end\_offset":6,"type":"","position":1},{"token":"4","start\_offset":7,"end\_offset":8,"type":"","position":2}]}
> > 
> > Presumably you're using the snowball stemmer? It analyzes 'iphone' as  
> > 'iphon' to be able to recognise eg "cansada" and "cansado" as the same  
> > stem.
> > 
> > All you need to do is to be sure that you're using the same analyzer at  
> > index time as at search time.
> > 
> > You have a few options here:
> > 
> > 1. you're searching on a field (eg product\_name) and you set the  
> > 'analyzer' for that field to be the spanish stemmer, when  
> > you put the mapping
> > 
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> > 
> > 1. you're searching on the '\_all' field (which is the default)  
> > and you can set the analyzer for the '\_all' field to  
> > be the spanish stemmer when you put the mapping
> > 
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/all-field.html)
> > 
> > (this is probably not what you want, as the \_all field will contain  
> > some fields which shouldn't have the stemmer applied)
> > 
> > 1. you can't determine at mapping time which language you're  
> > going to be using at search time, and you specify the  
> > analyzer in the query\_string query itself:
> > 
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/query-string-query.html)
> > 
> > 1. you could do something wizzy per document with the \_analyzer field
> > 
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/analyzer-field.html)
> > 
> > clint

--  
Joaquin Cuenca Abela -- [presspeople.com](http://presspeople.com): Fuentes de prensa y comunicados

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 1:21pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/12 "2011-03-16T13:21:22Z")

</div>

Nice catch, Joaquín.

I'll fix it and try to recreate the issue for Clinton.

On Wed, Mar 16, 2011 at 12:52 PM, Joaquin Cuenca Abela \<  
[joaquin@cuencaabela.com](mailto:joaquin@cuencaabela.com)\> wrote:

> I don't know why your pattern search is not working, but on this  
> analyzer you're ascii folding the terms before you remove the stop  
> words, and your stop words contain non ascii letters. You should  
> either ascii fold your stop words or remove them before the ascii  
> folding step.
> 
> Cheers,
> 
> On Wed, Mar 16, 2011 at 12:03 PM, Enrique Medina Montenegro  
> [e.medina.m@gmail.com](mailto:e.medina.m@gmail.com) wrote:
> 
> > Clinton,  
> > Based on some other discussion with Shay, I defined this in the  
> > elasticsearch.yml config file:  
> > index:  
> > analysis:  
> > analyzer:  
> > default:  
> > type: es.cuestamenos.lucene.analizadores.SpanishAnalyzerProvider  
> > And the analyzer is this one:  
> > [Custom analyzer for Spanish · GitHub](https://gist.github.com/872322)  
> > Shouldn't that be enough both for index and search?  
> > Thanks.
> > 
> > On Wed, Mar 16, 2011 at 11:51 AM, Clinton Gormley \<  
> > [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)\>  
> > wrote:
> > 
> > > Hi Enrique
> > > 
> > > > I found the issue and it was on my default Spanish analyzer. For some  
> > > > reason, iPhone gets analyzed in Spanish like this:
> > > > 
> > > > [http://localhost:9200/mytest/\_analyze?text=iPhone+4](http://localhost:9200/mytest/_analyze?text=iPhone+4)
> 
> {"tokens":[{"token":"iphon","start\_offset":0,"end\_offset":6,"type":"","position":1},{"token":"4","start\_offset":7,"end\_offset":8,"type":"","position":2}]}
> 
> > > Presumably you're using the snowball stemmer? It analyzes 'iphone' as  
> > > 'iphon' to be able to recognise eg "cansada" and "cansado" as the same  
> > > stem.
> > > 
> > > All you need to do is to be sure that you're using the same analyzer at  
> > > index time as at search time.
> > > 
> > > You have a few options here:
> > > 
> > > 1. you're searching on a field (eg product\_name) and you set the  
> > > 'analyzer' for that field to be the spanish stemmer, when  
> > > you put the mapping
> > > 
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> > > 
> > > 1. you're searching on the '\_all' field (which is the default)  
> > > and you can set the analyzer for the '\_all' field to  
> > > be the spanish stemmer when you put the mapping
> > > 
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/all-field.html)
> > > 
> > > (this is probably not what you want, as the \_all field will contain  
> > > some fields which shouldn't have the stemmer applied)
> > > 
> > > 1. you can't determine at mapping time which language you're  
> > > going to be using at search time, and you specify the  
> > > analyzer in the query\_string query itself:
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/query-string-query.html)
> 
> > > 1. you could do something wizzy per document with the \_analyzer field
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/analyzer-field.html)
> 
> > > clint
> 
> --  
> Joaquin Cuenca Abela -- [presspeople.com](http://presspeople.com): Fuentes de prensa y comunicados

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 1:38pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/13 "2011-03-16T13:38:33Z")

</div>

I think I found the issue without having to do a full recreation...

If I search using this:

{  
"query": {  
"query\_string": {  
"default\_operator": "AND",  
"query": "iphone"  
}  
}  
}

then it works as expected, I do get the expected results with the word  
"iPhone".

However, if I use:

{  
"query": {  
"query\_string": {  
"default\_operator": "AND",  
"query": "\*phone"  
}  
}  
}

then I don't get them. It seems that when you specify a wildcard in the  
query, it's not being properly analyzed like it should:

[http://localhost:9200/mytest/\_analyze?text=\*phone](http://localhost:9200/mytest/_analyze?text=*phone)

{"tokens":[{"token":"phon","start\_offset":0,"end\_offset":5,"type":"","position":1}]}

Therefore the wildcard is lost when tokenizing it and the search  
doesn't return any results, as "iPhone" doesn't match the token  
"phon".

Does this make sense now?

On Wed, Mar 16, 2011 at 12:33 PM, Clinton Gormley  
[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> hi Enrique
> 
> > 
> 
> I've just remembered your original question, which was:
> 
> "\*phone"  
> vs  
> "_phone_"
> 
> As I understand it, the way this wildcard search works is that Lucene  
> looks up all matching terms, and searches against each of these.
> 
> So for some reason, "\*phone" doesn't find the the right term, but  
> "_phone_" does.
> 
> > I get consistent results (even the lack of results) when adding a  
> > specific field or specific analyzer:
> 
> You mean, you see the same thing?
> 
> > So I guess it's not a bug, but as explained in my previous email, the  
> > fact that the Spanish analyzer created a token = "iphon" for iPhone so  
> > no matter how I search, it will never match "\*phone", right?
> 
> No. This should work. For instance, using the default analyzer, if you  
> index "The Quick BROWN fox" you end up with the terms  
> "quick","brown","fox"
> 
> If you then search for "The Quick BROWN fox", it performs the same  
> analysis, resulting in the same terms, and searches for those.
> 
> So to me (and I'm ignorant of the Lucene internals) it sounds like a  
> potential bug in the lucene query parser syntax.
> 
> A complete recreation would be very useful for debugging.
> 
> clint

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 1:43pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/14 "2011-03-16T13:43:27Z")

</div>

Which makes me think, is the '\*' actually acting as a wildcard in the query  
or is it interpreted by Lucene as just another character that has to be  
analyzed, as explained in my previous email, therefore losing all the  
wildcard information for the search?

On Wed, Mar 16, 2011 at 2:38 PM, Enrique Medina Montenegro \<  
[e.medina.m@gmail.com](mailto:e.medina.m@gmail.com)\> wrote:

> I think I found the issue without having to do a full recreation...
> 
> If I search using this:
> 
> {  
> "query": {  
> "query\_string": {  
> "default\_operator": "AND",  
> "query": "iphone"  
> }  
> }  
> }
> 
> then it works as expected, I do get the expected results with the word  
> "iPhone".
> 
> However, if I use:
> 
> {  
> "query": {  
> "query\_string": {  
> "default\_operator": "AND",  
> "query": "\*phone"  
> }  
> }  
> }
> 
> then I don't get them. It seems that when you specify a wildcard in the  
> query, it's not being properly analyzed like it should:
> 
> [http://localhost:9200/mytest/\_analyze?text=\*phone](http://localhost:9200/mytest/_analyze?text=*phone)
> 
> {"tokens":[{"token":"phon","start\_offset":0,"end\_offset":5,"type":"","position":1}]}
> 
> Therefore the wildcard is lost when tokenizing it and the search doesn't return any results, as "iPhone" doesn't match the token "phon".
> 
> Does this make sense now?
> 
> On Wed, Mar 16, 2011 at 12:33 PM, Clinton Gormley \<[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)
> 
> > wrote:
> 
> > hi Enrique
> > 
> > > 
> > 
> > I've just remembered your original question, which was:
> > 
> > "\*phone"  
> > vs  
> > "_phone_"
> > 
> > As I understand it, the way this wildcard search works is that Lucene  
> > looks up all matching terms, and searches against each of these.
> > 
> > So for some reason, "\*phone" doesn't find the the right term, but  
> > "_phone_" does.
> > 
> > > I get consistent results (even the lack of results) when adding a  
> > > specific field or specific analyzer:
> > 
> > You mean, you see the same thing?
> > 
> > > So I guess it's not a bug, but as explained in my previous email, the  
> > > fact that the Spanish analyzer created a token = "iphon" for iPhone so  
> > > no matter how I search, it will never match "\*phone", right?
> > 
> > No. This should work. For instance, using the default analyzer, if you  
> > index "The Quick BROWN fox" you end up with the terms  
> > "quick","brown","fox"
> > 
> > If you then search for "The Quick BROWN fox", it performs the same  
> > analysis, resulting in the same terms, and searches for those.
> > 
> > So to me (and I'm ignorant of the Lucene internals) it sounds like a  
> > potential bug in the lucene query parser syntax.
> > 
> > A complete recreation would be very useful for debugging.
> > 
> > clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 16, 2011, 1:47pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/15 "2011-03-16T13:47:25Z")

</div>

Hi Enrique

On Wed, 2011-03-16 at 14:38 +0100, Enrique Medina Montenegro wrote:

> I think I found the issue without having to do a full recreation...

The reason I keep asking for a complete recreation is so that Shay has  
got a test case to figure out where the bug is. The easier you make  
things for him, the more likely your bug will get attended to.

> then I don't get them. It seems that when you specify a wildcard in  
> the query, it's not being properly analyzed like it should:

yes, i agree

> [http://localhost:9200/mytest/\_analyze?text=\*phone](http://localhost:9200/mytest/_analyze?text=*phone)
> 
> {"tokens":[{"token":"phon","start\_offset":0,"end\_offset":5,"type":"","position":1}]}
> 
> Therefore the wildcard is lost when tokenizing it and the search  
> doesn't return any results, as "iPhone" doesn't match the token  
> "phon".

Not quite - the analyze API is just one part of this. What you're not  
seeing is the lucene query parser in action. That's where I think the  
bug is.

I suggest that you gist a complete recreation and post an issue to

> **[Issues · elastic/elasticsearch](https://github.com/elastic/elasticsearch/issues)**
>
> Free and Open, Distributed, RESTful Search Engine. Contribute to elastic/elasticsearch development by creating an account on GitHub.

ta

clint

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 1:58pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/16 "2011-03-16T13:58:31Z")

</div>

Yes, I will post the recreation right after this email.

I did some more testing with the wildcard, and it seems that wildcards do  
not match blank spaces, so if you specify "_iphone_" it will not match a  
name of "iPhone 4", but something like "iPad/iPhone 4/iPod".

Is this the expected behaviour or am I missing something?

On Wed, Mar 16, 2011 at 2:47 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> Hi Enrique
> 
> On Wed, 2011-03-16 at 14:38 +0100, Enrique Medina Montenegro wrote:
> 
> > I think I found the issue without having to do a full recreation...
> 
> The reason I keep asking for a complete recreation is so that Shay has  
> got a test case to figure out where the bug is. The easier you make  
> things for him, the more likely your bug will get attended to.
> 
> > then I don't get them. It seems that when you specify a wildcard in  
> > the query, it's not being properly analyzed like it should:
> 
> yes, i agree
> 
> > [http://localhost:9200/mytest/\_analyze?text=\*phone](http://localhost:9200/mytest/_analyze?text=*phone)
> 
> {"tokens":[{"token":"phon","start\_offset":0,"end\_offset":5,"type":"","position":1}]}
> 
> > Therefore the wildcard is lost when tokenizing it and the search  
> > doesn't return any results, as "iPhone" doesn't match the token  
> > "phon".
> 
> Not quite - the analyze API is just one part of this. What you're not  
> seeing is the lucene query parser in action. That's where I think the  
> bug is.
> 
> I suggest that you gist a complete recreation and post an issue to  
> [Issues · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues)
> 
> ta
> 
> clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 16, 2011, 2:11pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/17 "2011-03-16T14:11:23Z")

</div>

On Wed, 2011-03-16 at 14:58 +0100, Enrique Medina Montenegro wrote:

> Yes, I will post the recreation right after this email.

thanks 🙂

> I did some more testing with the wildcard, and it seems that wildcards  
> do not match blank spaces, so if you specify "_iphone_" it will not  
> match a name of "iPhone 4", but something like "iPad/iPhone 4/iPod".
> 
> Is this the expected behaviour or am I missing something?

This is correct - so I was wrong is saying that \* is equivalent to % in  
SQL. It works only on a per-word basis.

Also searching for '"ipho\*"' (ie in double quotes) would not work, as  
the \* would be interpreted literally, rather than as a wildcard.

clint

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 2:26pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/18 "2011-03-16T14:26:17Z")

</div>

Then it's definitively clear that it's not a bug, but a side effect of my  
Spanish analyzer tokenizing "iPhone" as "iphon", therefore not matching  
"phone" (token is different) or "\*phone" (wildcard takes word as a term, not  
its token).

I wonder if there's any sort of query in Lucene that acts as a '%' SQL  
wildcard, so for instance, when I specify "\*phone", instead of matching  
tokens with that literal "phone" and something before, it could first  
tokenize the literal, i.e. "phon", and then perform the search, which would  
definitively match my "iPhone"...

Maybe the solution is to tokenize the words entered by the user before  
applying the wildcard, and then passing the tokenized version to the query  
eventually...

On Wed, Mar 16, 2011 at 3:11 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> On Wed, 2011-03-16 at 14:58 +0100, Enrique Medina Montenegro wrote:
> 
> > Yes, I will post the recreation right after this email.
> 
> thanks 🙂
> 
> > I did some more testing with the wildcard, and it seems that wildcards  
> > do not match blank spaces, so if you specify "_iphone_" it will not  
> > match a name of "iPhone 4", but something like "iPad/iPhone 4/iPod".
> > 
> > Is this the expected behaviour or am I missing something?
> 
> This is correct - so I was wrong is saying that \* is equivalent to % in  
> SQL. It works only on a per-word basis.
> 
> Also searching for '"ipho\*"' (ie in double quotes) would not work, as  
> the \* would be interpreted literally, rather than as a wildcard.
> 
> clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 16, 2011, 3:07pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/19 "2011-03-16T15:07:42Z")

</div>

On Wed, 2011-03-16 at 15:26 +0100, Enrique Medina Montenegro wrote:

> Then it's definitively clear that it's not a bug, but a side effect of  
> my Spanish analyzer tokenizing "iPhone" as "iphon", therefore not  
> matching "phone" (token is different) or "\*phone" (wildcard takes word  
> as a term, not its token).

OK, I'm tired of trying to convince you that this is a bug. So i've  
opened the issue for you, with a recreation:

> <https://github.com/elastic/elasticsearch/issues/784>
>
> Hiya
> 
> A query string search with a wildcard on a field that has passed through t…he snowball stemmer is not working correctly.
> 
> Text indexed: \`I have an iPhone\`
> Search for: \`iphone\` works, but search for \`iphone\*\` doesn't
> 
> \`\`\`
> \# \[Wed Mar 16 16:02:34 2011\] Protocol: http, Server: 192.168.5.103:9200
> curl -XPUT 'http://127.0.0.1:9200/test/?pretty=1' -d '
> {
> "mappings" : {
> "doc" : {
> "properties" : {
> "text" : {
> "type" : "string",
> "analyzer" : "spanish"
> }
> }
> }
> },
> "settings" : {
> "analysis" : {
> "analyzer" : {
> "spanish" : {
> "language" : "Spanish",
> "type" : "snowball"
> }
> }
> }
> }
> }
> '
> 
> \# \[Wed Mar 16 16:02:34 2011\] Response:
> \# {
> \# "ok" : true,
> \# "acknowledged" : true
> \# }
> 
> \# \[Wed Mar 16 16:02:39 2011\] Protocol: http, Server: 192.168.5.103:9200
> curl -XPOST 'http://127.0.0.1:9200/test/doc?pretty=1' -d '
> {
> "text" : "I have an iPhone"
> }
> '
> 
> \# \[Wed Mar 16 16:02:39 2011\] Response:
> \# {
> \# "ok" : true,
> \# "\_index" : "test",
> \# "\_id" : "imN8\_G5rTwGwuESyTEz8pg",
> \# "\_type" : "doc",
> \# "\_version" : 1
> \# }
> 
> \# \[Wed Mar 16 16:02:45 2011\] Protocol: http, Server: 192.168.5.103:9200
> curl -XGET 'http://127.0.0.1:9200/test/doc/\_search?pretty=1' -d '
> {
> "query" : {
> "field" : {
> "text" : "iphone"
> }
> }
> }
> '
> 
> \# \[Wed Mar 16 16:02:45 2011\] Response:
> \# {
> \# "hits" : {
> \# "hits" : \[
> \# {
> \# "\_source" : {
> \# "text" : "I have an iPhone"
> \# },
> \# "\_score" : 0.15342641,
> \# "\_index" : "test",
> \# "\_id" : "imN8\_G5rTwGwuESyTEz8pg",
> \# "\_type" : "doc"
> \# }
> \# \],
> \# "max\_score" : 0.15342641,
> \# "total" : 1
> \# },
> \# "timed\_out" : false,
> \# "\_shards" : {
> \# "failed" : 0,
> \# "successful" : 5,
> \# "total" : 5
> \# },
> \# "took" : 2
> \# }
> 
> \# \[Wed Mar 16 16:02:48 2011\] Protocol: http, Server: 192.168.5.103:9200
> curl -XGET 'http://127.0.0.1:9200/test/doc/\_search?pretty=1' -d '
> {
> "query" : {
> "field" : {
> "text" : "iphone\*"
> }
> }
> }
> '
> 
> \# \[Wed Mar 16 16:02:48 2011\] Response:
> \# {
> \# "hits" : {
> \# "hits" : \[\],
> \# "max\_score" : null,
> \# "total" : 0
> \# },
> \# "timed\_out" : false,
> \# "\_shards" : {
> \# "failed" : 0,
> \# "successful" : 5,
> \# "total" : 5
> \# },
> \# "took" : 1
> \# }
> \`\`\`

clint

---

<div class="post-metadata">

**Author:** ![Enrique\_Medina\_Monte](https://avatars.discourse-cdn.com/v4/letter/e/4491bb/32.png) [@Enrique\_Medina\_Monte](https://discuss.elastic.co/u/Enrique_Medina_Monte)\
**Post date:** [March 16, 2011, 3:19pm UTC](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103/20 "2011-03-16T15:19:19Z")

</div>

I was already working on the recreation, but if you already did, it's done.

All my confusion is about the expected behaviour of the wildcard queries. So  
let's say, if my user wants to search for "iPhone" and I run a query with  
wildcards "_iPhone_" then:

1. As it is working now, it will not analyze the search term, but just use  
iPhone as the token itself, therefore not finding "iPhone" which has a token  
of "iphon".

2. As I expected it to be, i.e. the "_iPhone_" is analyzed into "_iphon_"  
and then executed the search, and "iPhone" results are returned.

So if current behaviour is 1), it's not a bug, but just a misunderstanding  
on my side. If current behaviour should be 2), then there's definitively a  
bug.

Looking forward to Shay's feedback on it.

On Wed, Mar 16, 2011 at 4:07 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> On Wed, 2011-03-16 at 15:26 +0100, Enrique Medina Montenegro wrote:
> 
> > Then it's definitively clear that it's not a bug, but a side effect of  
> > my Spanish analyzer tokenizing "iPhone" as "iphon", therefore not  
> > matching "phone" (token is different) or "\*phone" (wildcard takes word  
> > as a term, not its token).
> 
> OK, I'm tired of trying to convince you that this is a bug. So i've  
> opened the issue for you, with a recreation:
> 
> [WIldcard not working with snowball stemmer · Issue #784 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/784)
> 
> clint

[Next page](https://discuss.elastic.co/t/confused-about-query-string-and-the-use-of-wildcards/4103.md?page=2)
