# Wildcard and slashes

**URL:** <https://discuss.elastic.co/t/wildcard-and-slashes/7088>\
**Category:** Elasticsearch\
**Created:** [March 21, 2012, 10:29pm UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088 "2012-03-21T22:29:49Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![chimingc](https://avatars.discourse-cdn.com/v4/letter/c/ba9def/32.png) [@chimingc](https://discuss.elastic.co/u/chimingc)\
**Post date:** [March 21, 2012, 10:29pm UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/1 "2012-03-21T22:29:49Z")

</div>

I've indexed this: "24/account".  
I understand that it's been tokenized into "24" and "account", which is not a problem for me.

However, when I query "24/a\*", it finds no match.

Then I tried the following cases:

1. This works  
"query\_string": {  
"analyze\_wildcard": true,  
"query":"24/a\*"  
}

2. This works  
"query\_string": {  
"query":"24/account"  
}

3. This works  
"query\_string": {  
"query":"24 / a\*"  
}

4. This works  
"query\_string": {  
"query":"a\*"  
}

5. This doesn't work  
"query\_string": {  
"query":"24/a\*"  
}

I can't explain why 5 doesn't work. Perhaps without setting analyze\_wildcard to true, elasticsearch simply removes the slash and searches for "24a\*"?

What exactly does analyze\_wildcard do when set to true?  
As you can see 3 and 4 work without setting analyze\_wildcard to true. So when do we need to set it to true?

Thanks,  
jimmy

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [March 22, 2012, 2:11am UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/2 "2012-03-22T02:11:08Z")

</div>

Assuming that you are using standard analyzer, this is what these 5 queries  
are translated into on Lucene level:

1: \_all:24\* - prefix query for terms that start with "24"  
2: \_all:24 \_all:account - query for the term "24" or the term "account"  
3: \_all:24 \_all:a\* - query for the term "24" or prefix query for terms  
that start with "a"  
4: \_all:a\* - prefix query for terms that start with "a"  
5: \_all:24/a\* - prefix query for terms that start with "24/a"

Cases 2-4 are obvious, but cases 1 and 5, probably, require some  
explanation. By default, wildcard terms are not analyzed. This is why 5th  
case is getting translated into prefix query with the prefix "24/a". As you  
correctly noticed, "24/account" is indexed as two tokens "24" and  
"account". So, there are no tokens in the index that start with 24/a and  
therefore 5th case doesn't return any results. In the case 1, wildcard  
terms are analyzed and "24/a" is getting translated into two tokens "24"  
and "a". The token "a\*" is a stopword and it's getting dropped and the  
query is getting translated into prefix query for terms that start with 24.

On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:

> I've indexed this: "24/account".  
> I understand that it's been tokenized into "24" and "account", which is not  
> a problem for me.
> 
> However, when I query "24/a\*", it finds no match.
> 
> Then I tried the following cases:
> 
> 1. This works  
> "query\_string": {  
> "analyze\_wildcard": true,  
> "query":"24/a\*"  
> }
> 
> 2. This works  
> "query\_string": {  
> "query":"24/account"  
> }
> 
> 3. This works  
> "query\_string": {  
> "query":"24 / a\*"  
> }
> 
> 4. This works  
> "query\_string": {  
> "query":"a\*"  
> }
> 
> 5. This doesn't work  
> "query\_string": {  
> "query":"24/a\*"  
> }
> 
> I can't explain why 5 doesn't work. Perhaps without setting  
> analyze\_wildcard  
> to true, elasticsearch simply removes the slash and searches for "24a\*"?
> 
> What exactly does analyze\_wildcard do when set to true?  
> As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
> when do we need to set it to true?
> 
> Thanks,  
> jimmy
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![Gregory\_Rice](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gregory_rice/32/2884_2.png) [@Gregory\_Rice](https://discuss.elastic.co/u/Gregory_Rice)\
**Post date:** [March 22, 2012, 4:18am UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/3 "2012-03-22T04:18:18Z")

</div>

Igor and friends,

I've got a similar situation as far as special characters in a string,  
and I've instructed Elasticsearch to index a field containing the  
string "frag-mpm" with the following analyzer:

index :  
analysis :  
analyzer :  
string\_lowercase:  
tokenizer: keyword  
filter: lowercase

using the following mapping:

{  
"clientlog" : {  
"\_analyzer" : {  
"path" : "analyzer"  
},  
"\_source" : {  
"enabled" : true,  
"compress" : true  
},  
"properties" : {  
"analyzer" : {  
"type" : "string",  
"index" : "no"  
},  
"@fields" : {  
"dynamic" : "true",  
"type" : "object"  
},  
"@timestamp" : {  
"format" : "dateOptionalTime",  
"type" : "date"  
},  
"@message" : {  
"type" : "string",  
"analyzer" :"string\_lowercase"  
},  
"@source" : {  
"type" : "string"  
},  
"@type" : {  
"type" : "string"  
},  
"@tags" : {  
"type" : "string"  
},  
"@source\_host" : {  
"type" : "string"  
},  
"@source\_path" : {  
"type" : "string"  
}  
}  
}  
}

I'm still not seeing any search results for the whole string, "frag-  
mpm". I see stuff that contains the string "frag", but is it lucene  
itself splitting it, even though the analyzer is indexing it with a  
keyword tokenizer?

What am I configuring wrong?

Thanks,  
Greg Rice  
MobiTV

On Mar 21, 7:11 pm, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:

> Assuming that you are using standard analyzer, this is what these 5 queries  
> are translated into on Lucene level:
> 
> 1: \_all:24\* - prefix query for terms that start with "24"  
> 2: \_all:24 \_all:account - query for the term "24" or the term "account"  
> 3: \_all:24 \_all:a\* - query for the term "24" or prefix query for terms  
> that start with "a"  
> 4: \_all:a\* - prefix query for terms that start with "a"  
> 5: \_all:24/a\* - prefix query for terms that start with "24/a"
> 
> Cases 2-4 are obvious, but cases 1 and 5, probably, require some  
> explanation. By default, wildcard terms are not analyzed. This is why 5th  
> case is getting translated into prefix query with the prefix "24/a". As you  
> correctly noticed, "24/account" is indexed as two tokens "24" and  
> "account". So, there are no tokens in the index that start with 24/a and  
> therefore 5th case doesn't return any results. In the case 1, wildcard  
> terms are analyzed and "24/a" is getting translated into two tokens "24"  
> and "a". The token "a\*" is a stopword and it's getting dropped and the  
> query is getting translated into prefix query for terms that start with 24.
> 
> On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:
> 
> > I've indexed this: "24/account".  
> > I understand that it's been tokenized into "24" and "account", which is not  
> > a problem for me.
> 
> > However, when I query "24/a\*", it finds no match.
> 
> > Then I tried the following cases:
> 
> > 1. This works  
> > "query\_string": {  
> > "analyze\_wildcard": true,  
> > "query":"24/a\*"  
> > }
> 
> > 1. This works  
> > "query\_string": {  
> > "query":"24/account"  
> > }
> 
> > 1. This works  
> > "query\_string": {  
> > "query":"24 / a\*"  
> > }
> 
> > 1. This works  
> > "query\_string": {  
> > "query":"a\*"  
> > }
> 
> > 1. This doesn't work  
> > "query\_string": {  
> > "query":"24/a\*"  
> > }
> 
> > I can't explain why 5 doesn't work. Perhaps without setting  
> > analyze\_wildcard  
> > to true, elasticsearch simply removes the slash and searches for "24a\*"?
> 
> > What exactly does analyze\_wildcard do when set to true?  
> > As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
> > when do we need to set it to true?
> 
> > Thanks,  
> > jimmy
> 
> > --  
> > View this message in context:  
> > [http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-)...  
> > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 22, 2012, 7:08am UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/4 "2012-03-22T07:08:48Z")

</div>

Hi there,

It could depend on how you query.  
If you want to find "frag-mpm", you have to query on specific field  
containing this value.  
By default, you search in \_all which have a default analyzer so this field  
is broken in tokens.

HTH  
David

> -----Message d'origine-----  
> De : [elasticsearch@googlegroups.com](mailto:elasticsearch@googlegroups.com)  
> [[mailto:elasticsearch@googlegroups.com](mailto:elasticsearch@googlegroups.com)] De la part de Gregory Rice  
> Envoyé : jeudi 22 mars 2012 05:18  
> À : elasticsearch  
> Objet : Re: wildcard and slashes
> 
> Igor and friends,
> 
> I've got a similar situation as far as special characters in a string,  
> and I've instructed Elasticsearch to index a field containing the  
> string "frag-mpm" with the following analyzer:
> 
> index :  
> analysis :  
> analyzer :  
> string\_lowercase:  
> tokenizer: keyword  
> filter: lowercase
> 
> using the following mapping:
> 
> {  
> "clientlog" : {  
> "\_analyzer" : {  
> "path" : "analyzer"  
> },  
> "\_source" : {  
> "enabled" : true,  
> "compress" : true  
> },  
> "properties" : {  
> "analyzer" : {  
> "type" : "string",  
> "index" : "no"  
> },  
> "@fields" : {  
> "dynamic" : "true",  
> "type" : "object"  
> },  
> "@timestamp" : {  
> "format" : "dateOptionalTime",  
> "type" : "date"  
> },  
> "@message" : {  
> "type" : "string",  
> "analyzer" :"string\_lowercase"  
> },  
> "@source" : {  
> "type" : "string"  
> },  
> "@type" : {  
> "type" : "string"  
> },  
> "@tags" : {  
> "type" : "string"  
> },  
> "@source\_host" : {  
> "type" : "string"  
> },  
> "@source\_path" : {  
> "type" : "string"  
> }  
> }  
> }  
> }
> 
> I'm still not seeing any search results for the whole string, "frag-  
> mpm". I see stuff that contains the string "frag", but is it lucene  
> itself splitting it, even though the analyzer is indexing it with a  
> keyword tokenizer?
> 
> What am I configuring wrong?
> 
> Thanks,  
> Greg Rice  
> MobiTV
> 
> On Mar 21, 7:11 pm, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> 
> > Assuming that you are using standard analyzer, this is what these 5  
> > queries are translated into on Lucene level:
> > 
> > 1: \_all:24\* - prefix query for terms that start with "24"  
> > 2: \_all:24 \_all:account - query for the term "24" or the term  
> > "account"  
> > 3: \_all:24 \_all:a\* - query for the term "24" or prefix query for  
> > terms that start with "a"  
> > 4: \_all:a\* - prefix query for terms that start with "a"  
> > 5: \_all:24/a\* - prefix query for terms that start with "24/a"
> > 
> > Cases 2-4 are obvious, but cases 1 and 5, probably, require some  
> > explanation. By default, wildcard terms are not analyzed. This is why  
> > 5th case is getting translated into prefix query with the prefix  
> > "24/a". As you correctly noticed, "24/account" is indexed as two  
> > tokens "24" and "account". So, there are no tokens in the index that  
> > start with 24/a and therefore 5th case doesn't return any results. In  
> > the case 1, wildcard terms are analyzed and "24/a" is getting  
> > translated into two tokens "24"  
> > and "a". The token "a\*" is a stopword and it's getting dropped and  
> > the  
> > query is getting translated into prefix query for terms that start  
> > with 24.
> > 
> > On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:
> > 
> > > I've indexed this: "24/account".  
> > > I understand that it's been tokenized into "24" and "account",  
> > > which  
> > > is not a problem for me.
> > 
> > > However, when I query "24/a\*", it finds no match.
> > 
> > > Then I tried the following cases:
> > 
> > > 1. This works  
> > > "query\_string": {  
> > > "analyze\_wildcard": true,  
> > > "query":"24/a\*"  
> > > }
> > 
> > > 1. This works  
> > > "query\_string": {  
> > > "query":"24/account"  
> > > }
> > 
> > > 1. This works  
> > > "query\_string": {  
> > > "query":"24 / a\*"  
> > > }
> > 
> > > 1. This works  
> > > "query\_string": {  
> > > "query":"a\*"  
> > > }
> > 
> > > 1. This doesn't work  
> > > "query\_string": {  
> > > "query":"24/a\*"  
> > > }
> > 
> > > I can't explain why 5 doesn't work. Perhaps without setting  
> > > analyze\_wildcard to true, elasticsearch simply removes the slash  
> > > and  
> > > searches for "24a\*"?
> > 
> > > What exactly does analyze\_wildcard do when set to true?  
> > > As you can see 3 and 4 work without setting analyze\_wildcard to  
> > > true. So when do we need to set it to true?
> > 
> > > Thanks,  
> > > jimmy
> > 
> > > --  
> > > View this message in context:  
> > > [http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-)  
> > > slashes-...  
> > > Sent from the Elasticsearch Users mailing list archive at  
> > > [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![chimingc](https://avatars.discourse-cdn.com/v4/letter/c/ba9def/32.png) [@chimingc](https://discuss.elastic.co/u/chimingc)\
**Post date:** [March 22, 2012, 5:23pm UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/5 "2012-03-22T17:23:10Z")

</div>

Igor,

Thanks for the response. Really helpful.  
I have a few more questions though.

Neither query 3 nor 5 is being analyzed, so why does 3 get broken down into 2 tokens but 5 doesn't?

Also, how did you get the translated queries? Anyway to query elastic to get them? I think it's very helpful to know what the query strings eventually become.

Thanks again,  
jimmy

From: "Igor Motov-3 [via ElasticSearch Users]" \<[ml-node+s115913n3847358h2@n3.nabble.com](mailto:ml-node+s115913n3847358h2@n3.nabble.com)[mailto:ml-node+s115913n3847358h2@n3.nabble.com](mailto:ml-node+s115913n3847358h2@n3.nabble.com)\>  
Date: Wed, 21 Mar 2012 21:11:28 -0500  
To: Jimmy Chen \<[jchen@sugarcrm.com](mailto:jchen@sugarcrm.com)[mailto:jchen@sugarcrm.com](mailto:jchen@sugarcrm.com)\>  
Subject: Re: wildcard and slashes

Assuming that you are using standard analyzer, this is what these 5 queries are translated into on Lucene level:

1: \_all:24\* - prefix query for terms that start with "24"  
2: \_all:24 \_all:account - query for the term "24" or the term "account"  
3: \_all:24 \_all:a\* - query for the term "24" or prefix query for terms that start with "a"  
4: \_all:a\* - prefix query for terms that start with "a"  
5: \_all:24/a\* - prefix query for terms that start with "24/a"

Cases 2-4 are obvious, but cases 1 and 5, probably, require some explanation. By default, wildcard terms are not analyzed. This is why 5th case is getting translated into prefix query with the prefix "24/a". As you correctly noticed, "24/account" is indexed as two tokens "24" and "account". So, there are no tokens in the index that start with 24/a and therefore 5th case doesn't return any results. In the case 1, wildcard terms are analyzed and "24/a" is getting translated into two tokens "24" and "a". The token "a\*" is a stopword and it's getting dropped and the query is getting translated into prefix query for terms that start with 24.

On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:  
I've indexed this: "24/account".  
I understand that it's been tokenized into "24" and "account", which is not  
a problem for me.

However, when I query "24/a\*", it finds no match.

Then I tried the following cases:

1. This works  
"query\_string": {  
"analyze\_wildcard": true,  
"query":"24/a\*"  
}

2. This works  
"query\_string": {  
"query":"24/account"  
}

3. This works  
"query\_string": {  
"query":"24 / a\*"  
}

4. This works  
"query\_string": {  
"query":"a\*"  
}

5. This doesn't work  
"query\_string": {  
"query":"24/a\*"  
}

I can't explain why 5 doesn't work. Perhaps without setting analyze\_wildcard  
to true, elasticsearch simply removes the slash and searches for "24a\*"?

What exactly does analyze\_wildcard do when set to true?  
As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
when do we need to set it to true?

Thanks,  
jimmy

--  
View this message in context: [http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html)  
Sent from the ElasticSearch Users mailing list archive at [Nabble.com](http://Nabble.com).

* * *

If you reply to this email, your message will be added to the discussion below:  
[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html)  
To unsubscribe from wildcard and slashes, click here[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=unsubscribe\_by\_code&node=3847057&code=amNoZW5Ac3VnYXJjcm0uY29tfDM4NDcwNTd8LTcxNDM5MzQ0Nw==](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=unsubscribe_by_code&node=3847057&code=amNoZW5Ac3VnYXJjcm0uY29tfDM4NDcwNTd8LTcxNDM5MzQ0Nw==).  
NAML[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro\_viewer&id=instant\_html!nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify\_subscribers!nabble%3Aemail.naml-instant\_emails!nabble%3Aemail.naml-send\_instant\_email!nabble%3Aemail.naml](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro_viewer&id=instant_html%21nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify_subscribers%21nabble%3Aemail.naml-instant_emails%21nabble%3Aemail.naml-send_instant_email%21nabble%3Aemail.naml)

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [March 22, 2012, 7:52pm UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/6 "2012-03-22T19:52:32Z")

</div>

3 gets broken into queries by query parser.

I am not aware of any simple way to get the translated queries. When I need  
to figure out what's actually going on with my queries I just start  
elasticsearch under debugger, place breakpoint  
here [https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/query/QueryPhase.java#L176](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/query/QueryPhase.java#L176)  
and execute my search. The query variable there points to the actual  
Lucene query.

On Thursday, March 22, 2012 2:02:31 PM UTC-4, chimingc wrote:

> Igor,
> 
> Thanks for the response. Really helpful.  
> I have a few more questions though.
> 
> Neither query 3 nor 5 is being analyzed, so why does 3 get broken down  
> into 2 tokens but 5 doesn't?
> 
> Also, how did you get the translated queries? Anyway to query elastic to  
> get them? I think it's very helpful to know what the query strings  
> eventually become.
> 
> Thanks again,  
> jimmy
> 
> From: "Igor Motov-3 [via Elasticsearch Users]" \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=0](http://user/SendEmail.jtp?type=node&node=3849083&i=0)
> 
> > 
> 
> Date: Wed, 21 Mar 2012 21:11:28 -0500  
> To: Jimmy Chen \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=1](http://user/SendEmail.jtp?type=node&node=3849083&i=1)
> 
> > 
> 
> Subject: Re: wildcard and slashes
> 
> Assuming that you are using standard analyzer, this is what these 5  
> queries are translated into on Lucene level:
> 
> 1: \_all:24\* - prefix query for terms that start with "24"  
> 2: \_all:24 \_all:account - query for the term "24" or the term "account"  
> 3: \_all:24 \_all:a\* - query for the term "24" or prefix query for terms  
> that start with "a"  
> 4: \_all:a\* - prefix query for terms that start with "a"  
> 5: \_all:24/a\* - prefix query for terms that start with "24/a"
> 
> Cases 2-4 are obvious, but cases 1 and 5, probably, require some  
> explanation. By default, wildcard terms are not analyzed. This is why 5th  
> case is getting translated into prefix query with the prefix "24/a". As you  
> correctly noticed, "24/account" is indexed as two tokens "24" and  
> "account". So, there are no tokens in the index that start with 24/a and  
> therefore 5th case doesn't return any results. In the case 1, wildcard  
> terms are analyzed and "24/a" is getting translated into two tokens "24"  
> and "a". The token "a\*" is a stopword and it's getting dropped and the  
> query is getting translated into prefix query for terms that start with 24.
> 
> On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:
> 
> > I've indexed this: "24/account".  
> > I understand that it's been tokenized into "24" and "account", which is  
> > not  
> > a problem for me.
> > 
> > However, when I query "24/a\*", it finds no match.
> > 
> > Then I tried the following cases:
> > 
> > 1. This works  
> > "query\_string": {  
> > "analyze\_wildcard": true,  
> > "query":"24/a\*"  
> > }
> > 
> > 2. This works  
> > "query\_string": {  
> > "query":"24/account"  
> > }
> > 
> > 3. This works  
> > "query\_string": {  
> > "query":"24 / a\*"  
> > }
> > 
> > 4. This works  
> > "query\_string": {  
> > "query":"a\*"  
> > }
> > 
> > 5. This doesn't work  
> > "query\_string": {  
> > "query":"24/a\*"  
> > }
> > 
> > I can't explain why 5 doesn't work. Perhaps without setting  
> > analyze\_wildcard  
> > to true, elasticsearch simply removes the slash and searches for "24a\*"?
> > 
> > What exactly does analyze\_wildcard do when set to true?  
> > As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
> > when do we need to set it to true?
> > 
> > Thanks,  
> > jimmy
> > 
> > --  
> > View this message in context:  
> > [http://elasticsearch-users.​115913.n3.nabble.com/wildcard-​and-slashes-tp3847057p3847057.​html](http://elasticsearch-users.xn--115913-9e0c.n3.nabble.com/wildcard-%E2%80%8Band-slashes-tp3847057p3847057.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html)  
> > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> * * *
> 
> If you reply to this email, your message will be added to the discussion  
> below:
> 
> [http://elasticsearch-users.​115913.n3.nabble.com/wildcard-​and-slashes-tp3847057p3847358.​html](http://elasticsearch-users.xn--115913-9e0c.n3.nabble.com/wildcard-%E2%80%8Band-slashes-tp3847057p3847358.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html)  
> To unsubscribe from wildcard and slashes, click here.  
> NAML[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro\_viewer&id=instant\_html!nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify\_subscribers!nabble%3Aemail.naml-instant\_emails!nabble%3Aemail.naml-send\_instant\_email!nabble%3Aemail.naml](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro_viewer&id=instant_html%21nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify_subscribers%21nabble%3Aemail.naml-instant_emails%21nabble%3Aemail.naml-send_instant_email%21nabble%3Aemail.naml)
> 
> * * *
> 
> View this message in context: Re: wildcard and slashes[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3849083.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3849083.html)  
> Sent from the Elasticsearch Users mailing list archive[http://elasticsearch-users.115913.n3.nabble.com/](http://elasticsearch-users.115913.n3.nabble.com/)at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![Gregory\_Rice](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gregory_rice/32/2884_2.png) [@Gregory\_Rice](https://discuss.elastic.co/u/Gregory_Rice)\
**Post date:** [March 22, 2012, 9:11pm UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/7 "2012-03-22T21:11:48Z")

</div>

Igor and David,

Thanks a ton for the info. One question:

Is there any way to specify which special characters are used for  
tokenization? Like, is there an easy way to say "Break on slashes, but  
not dashes", or do I need to make my own tokenizer to do that?

Thanks,  
Greg Rice

On Mar 22, 12:52 pm, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:

> 3 gets broken into queries by query parser.
> 
> I am not aware of any simple way to get the translated queries. When I need  
> to figure out what's actually going on with my queries I just start  
> elasticsearch under debugger, place breakpoint  
> herehttps://github.com/elasticsearch/elasticsearch/blob/master/src/main/j...  
> and execute my search. The query variable there points to the actual  
> Lucene query.
> 
> On Thursday, March 22, 2012 2:02:31 PM UTC-4, chimingc wrote:
> 
> > Igor,
> 
> > Thanks for the response. Really helpful.  
> > I have a few more questions though.
> 
> > Neither query 3 nor 5 is being analyzed, so why does 3 get broken down  
> > into 2 tokens but 5 doesn't?
> 
> > Also, how did you get the translated queries? Anyway to query elastic to  
> > get them? I think it's very helpful to know what the query strings  
> > eventually become.
> 
> > Thanks again,  
> > jimmy
> 
> > From: "Igor Motov-3 [via Elasticsearch Users]" \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=0](http://user/SendEmail.jtp?type=node&node=3849083&i=0)
> 
> > Date: Wed, 21 Mar 2012 21:11:28 -0500  
> > To: Jimmy Chen \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=1](http://user/SendEmail.jtp?type=node&node=3849083&i=1)
> 
> > Subject: Re: wildcard and slashes
> 
> > Assuming that you are using standard analyzer, this is what these 5  
> > queries are translated into on Lucene level:
> 
> > 1: \_all:24\* - prefix query for terms that start with "24"  
> > 2: \_all:24 \_all:account - query for the term "24" or the term "account"  
> > 3: \_all:24 \_all:a\* - query for the term "24" or prefix query for terms  
> > that start with "a"  
> > 4: \_all:a\* - prefix query for terms that start with "a"  
> > 5: \_all:24/a\* - prefix query for terms that start with "24/a"
> 
> > Cases 2-4 are obvious, but cases 1 and 5, probably, require some  
> > explanation. By default, wildcard terms are not analyzed. This is why 5th  
> > case is getting translated into prefix query with the prefix "24/a". As you  
> > correctly noticed, "24/account" is indexed as two tokens "24" and  
> > "account". So, there are no tokens in the index that start with 24/a and  
> > therefore 5th case doesn't return any results. In the case 1, wildcard  
> > terms are analyzed and "24/a" is getting translated into two tokens "24"  
> > and "a". The token "a\*" is a stopword and it's getting dropped and the  
> > query is getting translated into prefix query for terms that start with 24.
> 
> > On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:
> 
> > > I've indexed this: "24/account".  
> > > I understand that it's been tokenized into "24" and "account", which is  
> > > not  
> > > a problem for me.
> 
> > > However, when I query "24/a\*", it finds no match.
> 
> > > Then I tried the following cases:
> 
> > > 1. This works  
> > > "query\_string": {  
> > > "analyze\_wildcard": true,  
> > > "query":"24/a\*"  
> > > }
> 
> > > 1. This works  
> > > "query\_string": {  
> > > "query":"24/account"  
> > > }
> 
> > > 1. This works  
> > > "query\_string": {  
> > > "query":"24 / a\*"  
> > > }
> 
> > > 1. This works  
> > > "query\_string": {  
> > > "query":"a\*"  
> > > }
> 
> > > 1. This doesn't work  
> > > "query\_string": {  
> > > "query":"24/a\*"  
> > > }
> 
> > > I can't explain why 5 doesn't work. Perhaps without setting  
> > > analyze\_wildcard  
> > > to true, elasticsearch simply removes the slash and searches for "24a\*"?
> 
> > > What exactly does analyze\_wildcard do when set to true?  
> > > As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
> > > when do we need to set it to true?
> 
> > > Thanks,  
> > > jimmy
> 
> > > --  
> > > View this message in context:  
> > > [http://elasticsearch-users.​115913.n3.nabble.com/wildcard-​and-slashes-tp3847057p3847057.​html](http://elasticsearch-users.xn--115913-9e0c.n3.nabble.com/wildcard-%E2%80%8Band-slashes-tp3847057p3847057.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...)  
> > > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> > * * *
> > 
> > If you reply to this email, your message will be added to the discussion  
> > below:
> 
> > [http://elasticsearch-users.​115913.n3.nabble.com/wildcard-​and-slashes-tp3847057p3847358.​html](http://elasticsearch-users.xn--115913-9e0c.n3.nabble.com/wildcard-%E2%80%8Band-slashes-tp3847057p3847358.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...)  
> > To unsubscribe from wildcard and slashes, click here.  
> > NAML[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet....](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet....)
> 
> > * * *
> > 
> > View this message in context: Re: wildcard and slashes[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...)  
> > Sent from the Elasticsearch Users mailing list archive[http://elasticsearch-users.115913.n3.nabble.com/](http://elasticsearch-users.115913.n3.nabble.com/)at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 22, 2012, 9:50pm UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/8 "2012-03-22T21:50:18Z")

</div>

Perhaps this one : [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/pattern-tokenizer.html)

HTH  
David 😉  
Twitter : @dadoonet / @elasticsearchfr

Le 22 mars 2012 à 22:11, Gregory Rice [gregrice@gmail.com](mailto:gregrice@gmail.com) a écrit :

> Igor and David,
> 
> Thanks a ton for the info. One question:
> 
> Is there any way to specify which special characters are used for  
> tokenization? Like, is there an easy way to say "Break on slashes, but  
> not dashes", or do I need to make my own tokenizer to do that?
> 
> Thanks,  
> Greg Rice
> 
> On Mar 22, 12:52 pm, Igor Motov [imo...@gmail.com](mailto:imo...@gmail.com) wrote:
> 
> > 3 gets broken into queries by query parser.
> > 
> > I am not aware of any simple way to get the translated queries. When I need  
> > to figure out what's actually going on with my queries I just start  
> > elasticsearch under debugger, place breakpoint  
> > herehttps://github.com/elasticsearch/elasticsearch/blob/master/src/main/j...  
> > and execute my search. The query variable there points to the actual  
> > Lucene query.
> > 
> > On Thursday, March 22, 2012 2:02:31 PM UTC-4, chimingc wrote:
> > 
> > > Igor,
> > 
> > > Thanks for the response. Really helpful.  
> > > I have a few more questions though.
> > 
> > > Neither query 3 nor 5 is being analyzed, so why does 3 get broken down  
> > > into 2 tokens but 5 doesn't?
> > 
> > > Also, how did you get the translated queries? Anyway to query elastic to  
> > > get them? I think it's very helpful to know what the query strings  
> > > eventually become.
> > 
> > > Thanks again,  
> > > jimmy
> > 
> > > From: "Igor Motov-3 [via Elasticsearch Users]" \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=0](http://user/SendEmail.jtp?type=node&node=3849083&i=0)
> > 
> > > Date: Wed, 21 Mar 2012 21:11:28 -0500  
> > > To: Jimmy Chen \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=1](http://user/SendEmail.jtp?type=node&node=3849083&i=1)
> > 
> > > Subject: Re: wildcard and slashes
> > 
> > > Assuming that you are using standard analyzer, this is what these 5  
> > > queries are translated into on Lucene level:
> > 
> > > 1: \_all:24\* - prefix query for terms that start with "24"  
> > > 2: \_all:24 \_all:account - query for the term "24" or the term "account"  
> > > 3: \_all:24 \_all:a\* - query for the term "24" or prefix query for terms  
> > > that start with "a"  
> > > 4: \_all:a\* - prefix query for terms that start with "a"  
> > > 5: \_all:24/a\* - prefix query for terms that start with "24/a"
> > 
> > > Cases 2-4 are obvious, but cases 1 and 5, probably, require some  
> > > explanation. By default, wildcard terms are not analyzed. This is why 5th  
> > > case is getting translated into prefix query with the prefix "24/a". As you  
> > > correctly noticed, "24/account" is indexed as two tokens "24" and  
> > > "account". So, there are no tokens in the index that start with 24/a and  
> > > therefore 5th case doesn't return any results. In the case 1, wildcard  
> > > terms are analyzed and "24/a" is getting translated into two tokens "24"  
> > > and "a". The token "a\*" is a stopword and it's getting dropped and the  
> > > query is getting translated into prefix query for terms that start with 24.
> > 
> > > On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:
> > 
> > > > I've indexed this: "24/account".  
> > > > I understand that it's been tokenized into "24" and "account", which is  
> > > > not  
> > > > a problem for me.
> > 
> > > > However, when I query "24/a\*", it finds no match.
> > 
> > > > Then I tried the following cases:
> > 
> > > > 1. This works  
> > > > "query\_string": {  
> > > > "analyze\_wildcard": true,  
> > > > "query":"24/a\*"  
> > > > }
> > 
> > > > 1. This works  
> > > > "query\_string": {  
> > > > "query":"24/account"  
> > > > }
> > 
> > > > 1. This works  
> > > > "query\_string": {  
> > > > "query":"24 / a\*"  
> > > > }
> > 
> > > > 1. This works  
> > > > "query\_string": {  
> > > > "query":"a\*"  
> > > > }
> > 
> > > > 1. This doesn't work  
> > > > "query\_string": {  
> > > > "query":"24/a\*"  
> > > > }
> > 
> > > > I can't explain why 5 doesn't work. Perhaps without setting  
> > > > analyze\_wildcard  
> > > > to true, elasticsearch simply removes the slash and searches for "24a\*"?
> > 
> > > > What exactly does analyze\_wildcard do when set to true?  
> > > > As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
> > > > when do we need to set it to true?
> > 
> > > > Thanks,  
> > > > jimmy
> > 
> > > > --  
> > > > View this message in context:  
> > > > [http://elasticsearch-users.​115913.n3.nabble.com/wildcard-​and-slashes-tp3847057p3847057.​html](http://elasticsearch-users.xn--115913-9e0c.n3.nabble.com/wildcard-%E2%80%8Band-slashes-tp3847057p3847057.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...)  
> > > > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> > 
> > > * * *
> > > 
> > > If you reply to this email, your message will be added to the discussion  
> > > below:
> > 
> > > [http://elasticsearch-users.​115913.n3.nabble.com/wildcard-​and-slashes-tp3847057p3847358.​html](http://elasticsearch-users.xn--115913-9e0c.n3.nabble.com/wildcard-%E2%80%8Band-slashes-tp3847057p3847358.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...)  
> > > To unsubscribe from wildcard and slashes, click here.  
> > > NAML[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet....](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet....)
> > 
> > > * * *
> > > 
> > > View this message in context: Re: wildcard and slashes[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-...)  
> > > Sent from the Elasticsearch Users mailing list archive[http://elasticsearch-users.115913.n3.nabble.com/](http://elasticsearch-users.115913.n3.nabble.com/)at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [March 24, 2012, 1:54pm UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/9 "2012-03-24T13:54:04Z")

</div>

A better way to get translated queries is coming in 0.19.2 and 0.20.0. See [add extended validation information by imotov · Pull Request #1811 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/pull/1811)  
for details.

On Thursday, March 22, 2012 3:52:32 PM UTC-4, Igor Motov wrote:

> 3 gets broken into queries by query parser.
> 
> I am not aware of any simple way to get the translated queries. When I  
> need to figure out what's actually going on with my queries I just start  
> elasticsearch under debugger, place breakpoint here  
> [https://github.com/​elasticsearch/elasticsearch/​blob/master/src/main/java/org/​elasticsearch/search/query/​QueryPhase.java#L176](https://github.com/%E2%80%8Belasticsearch/elasticsearch/%E2%80%8Bblob/master/src/main/java/org/%E2%80%8Belasticsearch/search/query/%E2%80%8BQueryPhase.java#L176)[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/query/QueryPhase.java#L176](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/query/QueryPhase.java#L176) and execute my search. The query variable there points to the actual  
> Lucene query.
> 
> On Thursday, March 22, 2012 2:02:31 PM UTC-4, chimingc wrote:
> 
> > Igor,
> > 
> > Thanks for the response. Really helpful.  
> > I have a few more questions though.
> > 
> > Neither query 3 nor 5 is being analyzed, so why does 3 get broken down  
> > into 2 tokens but 5 doesn't?
> > 
> > Also, how did you get the translated queries? Anyway to query elastic to  
> > get them? I think it's very helpful to know what the query strings  
> > eventually become.
> > 
> > Thanks again,  
> > jimmy
> > 
> > From: "Igor Motov-3 [via Elasticsearch Users]" \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=0](http://user/SendEmail.jtp?type=node&node=3849083&i=0)
> > 
> > > 
> > 
> > Date: Wed, 21 Mar 2012 21:11:28 -0500  
> > To: Jimmy Chen \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=1](http://user/SendEmail.jtp?type=node&node=3849083&i=1)
> > 
> > > 
> > 
> > Subject: Re: wildcard and slashes
> > 
> > Assuming that you are using standard analyzer, this is what these 5  
> > queries are translated into on Lucene level:
> > 
> > 1: \_all:24\* - prefix query for terms that start with "24"  
> > 2: \_all:24 \_all:account - query for the term "24" or the term  
> > "account"  
> > 3: \_all:24 \_all:a\* - query for the term "24" or prefix query for  
> > terms that start with "a"  
> > 4: \_all:a\* - prefix query for terms that start with "a"  
> > 5: \_all:24/a\* - prefix query for terms that start with "24/a"
> > 
> > Cases 2-4 are obvious, but cases 1 and 5, probably, require some  
> > explanation. By default, wildcard terms are not analyzed. This is why 5th  
> > case is getting translated into prefix query with the prefix "24/a". As you  
> > correctly noticed, "24/account" is indexed as two tokens "24" and  
> > "account". So, there are no tokens in the index that start with 24/a and  
> > therefore 5th case doesn't return any results. In the case 1, wildcard  
> > terms are analyzed and "24/a" is getting translated into two tokens "24"  
> > and "a". The token "a\*" is a stopword and it's getting dropped and the  
> > query is getting translated into prefix query for terms that start with 24.
> > 
> > On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:
> > 
> > > I've indexed this: "24/account".  
> > > I understand that it's been tokenized into "24" and "account", which is  
> > > not  
> > > a problem for me.
> > > 
> > > However, when I query "24/a\*", it finds no match.
> > > 
> > > Then I tried the following cases:
> > > 
> > > 1. This works  
> > > "query\_string": {  
> > > "analyze\_wildcard": true,  
> > > "query":"24/a\*"  
> > > }
> > > 
> > > 2. This works  
> > > "query\_string": {  
> > > "query":"24/account"  
> > > }
> > > 
> > > 3. This works  
> > > "query\_string": {  
> > > "query":"24 / a\*"  
> > > }
> > > 
> > > 4. This works  
> > > "query\_string": {  
> > > "query":"a\*"  
> > > }
> > > 
> > > 5. This doesn't work  
> > > "query\_string": {  
> > > "query":"24/a\*"  
> > > }
> > > 
> > > I can't explain why 5 doesn't work. Perhaps without setting  
> > > analyze\_wildcard  
> > > to true, elasticsearch simply removes the slash and searches for "24a\*"?
> > > 
> > > What exactly does analyze\_wildcard do when set to true?  
> > > As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
> > > when do we need to set it to true?
> > > 
> > > Thanks,  
> > > jimmy
> > > 
> > > --  
> > > View this message in context:  
> > > [http://elasticsearch-users.​​115913.n3.nabble.com/wildcard-​​and-slashes-​tp3847057p3847057.​html](http://elasticsearch-users.xn--115913-9e0ca.n3.nabble.com/wildcard-%E2%80%8B%E2%80%8Band-slashes-%E2%80%8Btp3847057p3847057.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html)  
> > > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> > 
> > * * *
> > 
> > If you reply to this email, your message will be added to the  
> > discussion below:
> > 
> > [http://elasticsearch-users.​​115913.n3.nabble.com/wildcard-​​and-slashes-​tp3847057p3847358.​html](http://elasticsearch-users.xn--115913-9e0ca.n3.nabble.com/wildcard-%E2%80%8B%E2%80%8Band-slashes-%E2%80%8Btp3847057p3847358.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html)  
> > To unsubscribe from wildcard and slashes, click here.  
> > NAML[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro\_viewer&id=instant\_html!nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify\_subscribers!nabble%3Aemail.naml-instant\_emails!nabble%3Aemail.naml-send\_instant\_email!nabble%3Aemail.naml](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro_viewer&id=instant_html%21nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify_subscribers%21nabble%3Aemail.naml-instant_emails%21nabble%3Aemail.naml-send_instant_email%21nabble%3Aemail.naml)
> > 
> > * * *
> > 
> > View this message in context: Re: wildcard and slashes[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3849083.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3849083.html)  
> > Sent from the Elasticsearch Users mailing list archive[http://elasticsearch-users.115913.n3.nabble.com/](http://elasticsearch-users.115913.n3.nabble.com/)at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![chimingc](https://avatars.discourse-cdn.com/v4/letter/c/ba9def/32.png) [@chimingc](https://discuss.elastic.co/u/chimingc)\
**Post date:** [March 26, 2012, 4:18am UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/10 "2012-03-26T04:18:27Z")

</div>

Thanks, good to know.

From: "Igor Motov-3 [via ElasticSearch Users]" \<[ml-node+s115913n3853948h92@n3.nabble.com](mailto:ml-node+s115913n3853948h92@n3.nabble.com)[mailto:ml-node+s115913n3853948h92@n3.nabble.com](mailto:ml-node+s115913n3853948h92@n3.nabble.com)\>  
Date: Sat, 24 Mar 2012 10:24:42 -0500  
To: Jimmy Chen \<[jchen@sugarcrm.com](mailto:jchen@sugarcrm.com)[mailto:jchen@sugarcrm.com](mailto:jchen@sugarcrm.com)\>  
Subject: Re: wildcard and slashes

A better way to get translated queries is coming in 0.19.2 and 0.20.0. See [https://github.com/elasticsearch/elasticsearch/pull/1811](https://github.com/elasticsearch/elasticsearch/pull/1811) for details.

On Thursday, March 22, 2012 3:52:32 PM UTC-4, Igor Motov wrote:  
3 gets broken into queries by query parser.

I am not aware of any simple way to get the translated queries. When I need to figure out what's actually going on with my queries I just start elasticsearch under debugger, place breakpoint here [https://github.com/​elasticsearch/elasticsearch/​blob/master/src/main/java/org/​elasticsearch/search/query/​QueryPhase.java#L176](https://github.com/%E2%80%8Belasticsearch/elasticsearch/%E2%80%8Bblob/master/src/main/java/org/%E2%80%8Belasticsearch/search/query/%E2%80%8BQueryPhase.java#L176)[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/query/QueryPhase.java#L176](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/search/query/QueryPhase.java#L176) and execute my search. The query variable there points to the actual Lucene query.

On Thursday, March 22, 2012 2:02:31 PM UTC-4, chimingc wrote:  
Igor,

Thanks for the response. Really helpful.  
I have a few more questions though.

Neither query 3 nor 5 is being analyzed, so why does 3 get broken down into 2 tokens but 5 doesn't?

Also, how did you get the translated queries? Anyway to query elastic to get them? I think it's very helpful to know what the query strings eventually become.

Thanks again,  
jimmy

From: "Igor Motov-3 [via ElasticSearch Users]" \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=0](http://user/SendEmail.jtp?type=node&node=3849083&i=0)\>  
Date: Wed, 21 Mar 2012 21:11:28 -0500  
To: Jimmy Chen \<[hidden email][http://user/SendEmail.jtp?type=node&node=3849083&i=1](http://user/SendEmail.jtp?type=node&node=3849083&i=1)\>  
Subject: Re: wildcard and slashes

Assuming that you are using standard analyzer, this is what these 5 queries are translated into on Lucene level:

1: \_all:24\* - prefix query for terms that start with "24"  
2: \_all:24 \_all:account - query for the term "24" or the term "account"  
3: \_all:24 \_all:a\* - query for the term "24" or prefix query for terms that start with "a"  
4: \_all:a\* - prefix query for terms that start with "a"  
5: \_all:24/a\* - prefix query for terms that start with "24/a"

Cases 2-4 are obvious, but cases 1 and 5, probably, require some explanation. By default, wildcard terms are not analyzed. This is why 5th case is getting translated into prefix query with the prefix "24/a". As you correctly noticed, "24/account" is indexed as two tokens "24" and "account". So, there are no tokens in the index that start with 24/a and therefore 5th case doesn't return any results. In the case 1, wildcard terms are analyzed and "24/a" is getting translated into two tokens "24" and "a". The token "a\*" is a stopword and it's getting dropped and the query is getting translated into prefix query for terms that start with 24.

On Wednesday, March 21, 2012 6:29:49 PM UTC-4, chimingc wrote:  
I've indexed this: "24/account".  
I understand that it's been tokenized into "24" and "account", which is not  
a problem for me.

However, when I query "24/a\*", it finds no match.

Then I tried the following cases:

1. This works  
"query\_string": {  
"analyze\_wildcard": true,  
"query":"24/a\*"  
}

2. This works  
"query\_string": {  
"query":"24/account"  
}

3. This works  
"query\_string": {  
"query":"24 / a\*"  
}

4. This works  
"query\_string": {  
"query":"a\*"  
}

5. This doesn't work  
"query\_string": {  
"query":"24/a\*"  
}

I can't explain why 5 doesn't work. Perhaps without setting analyze\_wildcard  
to true, elasticsearch simply removes the slash and searches for "24a\*"?

What exactly does analyze\_wildcard do when set to true?  
As you can see 3 and 4 work without setting analyze\_wildcard to true. So  
when do we need to set it to true?

Thanks,  
jimmy

--  
View this message in context: [http://elasticsearch-users.​​115913.n3.nabble.com/wildcard-​​and-slashes-​tp3847057p3847057.​html](http://elasticsearch-users.xn--115913-9e0ca.n3.nabble.com/wildcard-%E2%80%8B%E2%80%8Band-slashes-%E2%80%8Btp3847057p3847057.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847057.html)  
Sent from the ElasticSearch Users mailing list archive at [Nabble.com](http://Nabble.com).

* * *

If you reply to this email, your message will be added to the discussion below:  
[http://elasticsearch-users.​​115913.n3.nabble.com/wildcard-​​and-slashes-​tp3847057p3847358.​html](http://elasticsearch-users.xn--115913-9e0ca.n3.nabble.com/wildcard-%E2%80%8B%E2%80%8Band-slashes-%E2%80%8Btp3847057p3847358.%E2%80%8Bhtml)[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3847358.html)  
To unsubscribe from wildcard and slashes, click here.  
NAML[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro\_viewer&id=instant\_html!nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify\_subscribers!nabble%3Aemail.naml-instant\_emails!nabble%3Aemail.naml-send\_instant\_email!nabble%3Aemail.naml](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro_viewer&id=instant_html%21nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify_subscribers%21nabble%3Aemail.naml-instant_emails%21nabble%3Aemail.naml-send_instant_email%21nabble%3Aemail.naml)

* * *

View this message in context: Re: wildcard and slashes[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3849083.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3849083.html)  
Sent from the ElasticSearch Users mailing list archive[http://elasticsearch-users.115913.n3.nabble.com/](http://elasticsearch-users.115913.n3.nabble.com/) at [Nabble.com](http://Nabble.com).

* * *

If you reply to this email, your message will be added to the discussion below:  
[http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3853948.html](http://elasticsearch-users.115913.n3.nabble.com/wildcard-and-slashes-tp3847057p3853948.html)  
To unsubscribe from wildcard and slashes, click here[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=unsubscribe\_by\_code&node=3847057&code=amNoZW5Ac3VnYXJjcm0uY29tfDM4NDcwNTd8LTcxNDM5MzQ0Nw==](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=unsubscribe_by_code&node=3847057&code=amNoZW5Ac3VnYXJjcm0uY29tfDM4NDcwNTd8LTcxNDM5MzQ0Nw==).  
NAML[http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro\_viewer&id=instant\_html!nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify\_subscribers!nabble%3Aemail.naml-instant\_emails!nabble%3Aemail.naml-send\_instant\_email!nabble%3Aemail.naml](http://elasticsearch-users.115913.n3.nabble.com/template/NamlServlet.jtp?macro=macro_viewer&id=instant_html%21nabble%3Aemail.naml&base=nabble.naml.namespaces.BasicNamespace-nabble.view.web.template.NabbleNamespace-nabble.view.web.template.NodeNamespace&breadcrumbs=notify_subscribers%21nabble%3Aemail.naml-instant_emails%21nabble%3Aemail.naml-send_instant_email%21nabble%3Aemail.naml)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:34am UTC](https://discuss.elastic.co/t/wildcard-and-slashes/7088/11 "2017-07-06T03:34:52Z")

</div>


