# EL setup for fulltext search

**URL:** <https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483>\
**Category:** Elasticsearch\
**Created:** [August 27, 2014, 6:57am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483 "2014-08-27T06:57:18Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Marc\_2](https://avatars.discourse-cdn.com/v4/letter/m/c4cdca/32.png) [@Marc\_2](https://discuss.elastic.co/u/Marc_2)\
**Post date:** [August 27, 2014, 6:57am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/1 "2014-08-27T06:57:18Z")

</div>

Hi,

I have quiet a simple scenario that already gives me a headache for quiet a  
while.  
I have one Field which is quiet big and full of special characters like  
(,),=,:,",' digits and text.  
Example:  
"msg" : "Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )"  
I essentially want to be able to search this things using text, wildcards  
etc.  
So far I have tried not analyzing the content and using the wildcard search  
and it doesn't work very well.  
Using different tokenizers and the query\_string query also only works to a  
certain degree.  
For example I want to be able to serach for following expressions:  
Service  
MyMDB  
onMessage  
MyMDB.onMessage  
appId=cs AND Times=Me:22

and other possible permutations.  
What is a correct setup?! I simply can't find a solution...

ps.: the data is imported to elasticsearch using logstash. We do acces the  
data using the java api (all software latest versions).

Cheeers,  
Marc

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 27, 2014, 7:20am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/2 "2014-08-27T07:20:41Z")

</div>

Off the top of my head, I would use a custom analyzer with a whitespace  
tokenizer and a word delimiter filter (preserving the original tokens as  
well). Perhaps a shingle filter to create bigrams. Or better yet a pattern  
tokenizer with spaces and parenthesis.

Cheers,

Ivan

On Tue, Aug 26, 2014 at 11:57 PM, Marc [mn.offman@googlemail.com](mailto:mn.offman@googlemail.com) wrote:

> Hi,
> 
> I have quiet a simple scenario that already gives me a headache for quiet  
> a while.  
> I have one Field which is quiet big and full of special characters like  
> (,),=,:,",' digits and text.  
> Example:  
> "msg" : "Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )"  
> I essentially want to be able to search this things using text, wildcards  
> etc.  
> So far I have tried not analyzing the content and using the wildcard  
> search and it doesn't work very well.  
> Using different tokenizers and the query\_string query also only works to a  
> certain degree.  
> For example I want to be able to serach for following expressions:  
> Service  
> MyMDB  
> onMessage  
> MyMDB.onMessage  
> appId=cs AND Times=Me:22
> 
> and other possible permutations.  
> What is a correct setup?! I simply can't find a solution...
> 
> ps.: the data is imported to elasticsearch using logstash. We do acces the  
> data using the java api (all software latest versions).
> 
> Cheeers,  
> Marc
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQBP26Q-H1Am3LBDXn6uLhg20tLregFdNLUan1Z8J2yTKg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQBP26Q-H1Am3LBDXn6uLhg20tLregFdNLUan1Z8J2yTKg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Marc\_2](https://avatars.discourse-cdn.com/v4/letter/m/c4cdca/32.png) [@Marc\_2](https://discuss.elastic.co/u/Marc_2)\
**Post date:** [August 28, 2014, 9:05am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/3 "2014-08-28T09:05:14Z")

</div>

Hi Ivan,

thanks for the help. Now it works almost... 😉  
I have used the following:  
"analysis": {  
"analyzer": {  
"msg\_excp\_analyzer": {  
"type": "custom",  
"tokenizer": "whitespace",  
"filters": ["split-up",  
"lowercase",  
"shingle",  
"ascii-folding"]  
}  
},  
"filter": {  
"split-up": {  
"type": "word\_delimiter",  
"preserve\_original": "true",  
"catenate\_all": "true",  
"type\_table": {  
"$": "DIGIT",  
"%": "DIGIT",  
".": "DIGIT",  
",": "DIGIT",  
":": "DIGIT",  
"/": "DIGIT",  
"\": "DIGIT",  
"=": "DIGIT",  
"&": "DIGIT",  
"(": "DIGIT",  
")": "DIGIT",  
"\<": "DIGIT",  
"\>": "DIGIT",  
"\U+000A": "DIGIT"  
}  
},  
"ascii-folding": {  
"type": "asciifolding",  
"preserve\_original": true  
}  
}  
If the above is wrong or not reasonable, please feel free to criticize!

Now the only thing that does not work is searching for subwords of  
concatenations with".".  
Having log Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ ) I cannot search for  
MyMDB or onMessage; only MyMDB.onMessage will work.

Anymore Ideas?

Cheers,  
Marc

On Wednesday, August 27, 2014 9:20:49 AM UTC+2, Ivan Brusic wrote:

> Off the top of my head, I would use a custom analyzer with a whitespace  
> tokenizer and a word delimiter filter (preserving the original tokens as  
> well). Perhaps a shingle filter to create bigrams. Or better yet a pattern  
> tokenizer with spaces and parenthesis.
> 
> Cheers,
> 
> Ivan
> 
> On Tue, Aug 26, 2014 at 11:57 PM, Marc \<[mn.o...@googlemail.com](mailto:mn.o...@googlemail.com)  
> \<javascript:\>\> wrote:
> 
> > Hi,
> > 
> > I have quiet a simple scenario that already gives me a headache for quiet  
> > a while.  
> > I have one Field which is quiet big and full of special characters like  
> > (,),=,:,",' digits and text.  
> > Example:  
> > "msg" : "Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> > attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )"  
> > I essentially want to be able to search this things using text, wildcards  
> > etc.  
> > So far I have tried not analyzing the content and using the wildcard  
> > search and it doesn't work very well.  
> > Using different tokenizers and the query\_string query also only works to  
> > a certain degree.  
> > For example I want to be able to serach for following expressions:  
> > Service  
> > MyMDB  
> > onMessage  
> > MyMDB.onMessage  
> > appId=cs AND Times=Me:22
> > 
> > and other possible permutations.  
> > What is a correct setup?! I simply can't find a solution...
> > 
> > ps.: the data is imported to elasticsearch using logstash. We do acces  
> > the data using the java api (all software latest versions).
> > 
> > Cheeers,  
> > Marc
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 28, 2014, 4:16pm UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/4 "2014-08-28T16:16:56Z")

</div>

Use the Analyze API to view what tokens are being generated? Keep it simple  
at first (maybe remove shingles) and build up as you encounter more  
edge-cases. What kind of query are you using?

--  
Ivan

On Thu, Aug 28, 2014 at 2:05 AM, Marc [mn.offman@googlemail.com](mailto:mn.offman@googlemail.com) wrote:

> Hi Ivan,
> 
> thanks for the help. Now it works almost... 😉  
> I have used the following:  
> "analysis": {  
> "analyzer": {  
> "msg\_excp\_analyzer": {  
> "type": "custom",  
> "tokenizer": "whitespace",  
> "filters": ["split-up",  
> "lowercase",  
> "shingle",  
> "ascii-folding"]  
> }  
> },  
> "filter": {  
> "split-up": {  
> "type": "word\_delimiter",  
> "preserve\_original": "true",  
> "catenate\_all": "true",  
> "type\_table": {  
> "$": "DIGIT",  
> "%": "DIGIT",  
> ".": "DIGIT",  
> ",": "DIGIT",  
> ":": "DIGIT",  
> "/": "DIGIT",  
> "\": "DIGIT",  
> "=": "DIGIT",  
> "&": "DIGIT",  
> "(": "DIGIT",  
> ")": "DIGIT",  
> "\<": "DIGIT",  
> "\>": "DIGIT",  
> "\U+000A": "DIGIT"  
> }  
> },  
> "ascii-folding": {  
> "type": "asciifolding",  
> "preserve\_original": true  
> }  
> }  
> If the above is wrong or not reasonable, please feel free to criticize!
> 
> Now the only thing that does not work is searching for subwords of  
> concatenations with".".  
> Having log Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ ) I cannot search for  
> MyMDB or onMessage; only MyMDB.onMessage will work.
> 
> Anymore Ideas?
> 
> Cheers,  
> Marc
> 
> On Wednesday, August 27, 2014 9:20:49 AM UTC+2, Ivan Brusic wrote:
> 
> > Off the top of my head, I would use a custom analyzer with a whitespace  
> > tokenizer and a word delimiter filter (preserving the original tokens as  
> > well). Perhaps a shingle filter to create bigrams. Or better yet a pattern  
> > tokenizer with spaces and parenthesis.
> > 
> > Cheers,
> > 
> > Ivan
> > 
> > On Tue, Aug 26, 2014 at 11:57 PM, Marc [mn.o...@googlemail.com](mailto:mn.o...@googlemail.com) wrote:
> > 
> > > Hi,
> > > 
> > > I have quiet a simple scenario that already gives me a headache for  
> > > quiet a while.  
> > > I have one Field which is quiet big and full of special characters like  
> > > (,),=,:,",' digits and text.  
> > > Example:  
> > > "msg" : "Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22  
> > > (updated attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )"  
> > > I essentially want to be able to search this things using text,  
> > > wildcards etc.  
> > > So far I have tried not analyzing the content and using the wildcard  
> > > search and it doesn't work very well.  
> > > Using different tokenizers and the query\_string query also only works to  
> > > a certain degree.  
> > > For example I want to be able to serach for following expressions:  
> > > Service  
> > > MyMDB  
> > > onMessage  
> > > MyMDB.onMessage  
> > > appId=cs AND Times=Me:22
> > > 
> > > and other possible permutations.  
> > > What is a correct setup?! I simply can't find a solution...
> > > 
> > > ps.: the data is imported to elasticsearch using logstash. We do acces  
> > > the data using the java api (all software latest versions).
> > > 
> > > Cheeers,  
> > > Marc
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).
> > > 
> > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%  
> > > [40googlegroups.com](http://40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .  
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQBbGG0fZ%2BGgwpcRqdmtnEeFoOFUu53P%3DZ34sLBrq39Lbw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQBbGG0fZ%2BGgwpcRqdmtnEeFoOFUu53P%3DZ34sLBrq39Lbw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Marc\_2](https://avatars.discourse-cdn.com/v4/letter/m/c4cdca/32.png) [@Marc\_2](https://discuss.elastic.co/u/Marc_2)\
**Post date:** [August 29, 2014, 8:48am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/5 "2014-08-29T08:48:31Z")

</div>

Hi Ivan,

thanks again. I have tried so and found a reasonable combination.  
Nevertheless, when I now try to use the analyze api with an index that has  
the said analyzer defined via template it doesn't seem to apply:

This is the complete template:  
{  
"template": "bogstash-\*",  
"settings": {  
"index.number\_of\_replicas": 0,  
"analysis": {  
"analyzer": {  
"msg\_excp\_analyzer": {  
"type": "custom",  
"tokenizer": "whitespace",  
"filters": ["word\_delimiter",  
"lowercase",  
"asciifolding",  
"shingle",  
"standard"]  
}  
},  
"filters": {  
"my\_word\_delimiter": {  
"type": "word\_delimiter",  
"preserve\_original": "true"  
},  
"my\_asciifolding": {  
"type": "asciifolding",  
"preserve\_original": true  
}  
}  
}  
},  
"mappings": {  
"_default_": {  
"properties": {  
"@excp": {  
"type": "string",  
"index": "analyzed",  
"analyzer": "msg\_excp\_analyzer"  
},  
"@msg": {  
"type": "string",  
"index": "analyzed",  
"analyzer": "msg\_excp\_analyzer"  
}  
}  
}  
}  
}  
I create the index bogstash-1.  
Now I test the following:  
curl -XGET  
'localhost:9200/bogstash-1/\_analyze?analyzer=msg\_excp\_analyzer&pretty=1' -d 'Service=MyMDB.onMessage  
appId=cs Times=Me:22/Total:22 (updated attributes=gps\_lng: 183731222/  
gps\_lat: 289309222/ )'  
and it returns:  
{  
"tokens" : [ {  
"token" : "Service=MyMDB.onMessage",  
"start\_offset" : 0,  
"end\_offset" : 23,  
"type" : "word",  
"position" : 1  
}, {  
"token" : "appId=cs",  
"start\_offset" : 24,  
"end\_offset" : 32,  
"type" : "word",  
"position" : 2  
}, {  
"token" : "Times=Me:22/Total:22",  
"start\_offset" : 33,  
"end\_offset" : 53,  
"type" : "word",  
"position" : 3  
}, {  
"token" : "(updated",  
"start\_offset" : 54,  
"end\_offset" : 62,  
"type" : "word",  
"position" : 4  
}, {  
"token" : "attributes=gps\_lng:",  
"start\_offset" : 63,  
"end\_offset" : 82,  
"type" : "word",  
"position" : 5  
}, {  
"token" : "183731222/",  
"start\_offset" : 83,  
"end\_offset" : 93,  
"type" : "word",  
"position" : 6  
}, {  
"token" : "gps\_lat:",  
"start\_offset" : 94,  
"end\_offset" : 102,  
"type" : "word",  
"position" : 7  
}, {  
"token" : "289309222/",  
"start\_offset" : 103,  
"end\_offset" : 113,  
"type" : "word",  
"position" : 8  
}, {  
"token" : ")",  
"start\_offset" : 114,  
"end\_offset" : 115,  
"type" : "word",  
"position" : 9  
} ]  
}  
Which is the output of a standard analyzer.  
Giving the tokenizer and filters in the analyze API directly works fine:  
curl -XGET  
'localhost:9200/\_analyze?tokenizer=whitespace&filters=lowercase,word\_delimiter,shingle,asciifolding,standard&pretty=1'  
-d 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
This results in:  
{  
"tokens" : [ {  
"token" : "service",  
"start\_offset" : 0,  
"end\_offset" : 7,  
"type" : "word",  
"position" : 1  
}, {  
"token" : "service mymdb",  
"start\_offset" : 0,  
"end\_offset" : 13,  
"type" : "shingle",  
"position" : 1  
}, {  
"token" : "mymdb",  
"start\_offset" : 8,  
"end\_offset" : 13,  
"type" : "word",  
"position" : 2  
}, {  
"token" : "mymdb onmessage",  
"start\_offset" : 8,  
"end\_offset" : 23,  
"type" : "shingle",  
"position" : 2  
}, {  
"token" : "onmessage",  
"start\_offset" : 14,  
"end\_offset" : 23,  
"type" : "word",  
"position" : 3  
}, {  
"token" : "onmessage appid",  
"start\_offset" : 14,  
"end\_offset" : 29,  
"type" : "shingle",  
"position" : 3  
}, {  
"token" : "appid",  
"start\_offset" : 24,  
"end\_offset" : 29,  
"type" : "word",  
"position" : 4  
}, {  
"token" : "appid cs",  
"start\_offset" : 24,  
"end\_offset" : 32,  
"type" : "shingle",  
"position" : 4  
}, {  
"token" : "cs",  
"start\_offset" : 30,  
"end\_offset" : 32,  
"type" : "word",  
"position" : 5  
}, {  
"token" : "cs times",  
"start\_offset" : 30,  
"end\_offset" : 38,  
"type" : "shingle",  
"position" : 5  
}, {  
"token" : "times",  
"start\_offset" : 33,  
"end\_offset" : 38,  
"type" : "word",  
"position" : 6  
}, {  
"token" : "times me",  
"start\_offset" : 33,  
"end\_offset" : 41,  
"type" : "shingle",  
"position" : 6  
}, {  
"token" : "me",  
"start\_offset" : 39,  
"end\_offset" : 41,  
"type" : "word",  
"position" : 7  
}, {  
"token" : "me 22",  
"start\_offset" : 39,  
"end\_offset" : 44,  
"type" : "shingle",  
"position" : 7  
}, {  
"token" : "22",  
"start\_offset" : 42,  
"end\_offset" : 44,  
"type" : "word",  
"position" : 8  
}, {  
"token" : "22 total",  
"start\_offset" : 42,  
"end\_offset" : 50,  
"type" : "shingle",  
"position" : 8  
}, {  
"token" : "total",  
"start\_offset" : 45,  
"end\_offset" : 50,  
"type" : "word",  
"position" : 9  
}, {  
"token" : "total 22",  
"start\_offset" : 45,  
"end\_offset" : 53,  
"type" : "shingle",  
"position" : 9  
}, {  
"token" : "22",  
"start\_offset" : 51,  
"end\_offset" : 53,  
"type" : "word",  
"position" : 10  
}, {  
"token" : "22 updated",  
"start\_offset" : 51,  
"end\_offset" : 62,  
"type" : "shingle",  
"position" : 10  
}, {  
"token" : "updated",  
"start\_offset" : 55,  
"end\_offset" : 62,  
"type" : "word",  
"position" : 11  
}, {  
"token" : "updated attributes",  
"start\_offset" : 55,  
"end\_offset" : 73,  
"type" : "shingle",  
"position" : 11  
}, {  
"token" : "attributes",  
"start\_offset" : 63,  
"end\_offset" : 73,  
"type" : "word",  
"position" : 12  
}, {  
"token" : "attributes gps",  
"start\_offset" : 63,  
"end\_offset" : 77,  
"type" : "shingle",  
"position" : 12  
}, {  
"token" : "gps",  
"start\_offset" : 74,  
"end\_offset" : 77,  
"type" : "word",  
"position" : 13  
}, {  
"token" : "gps lng",  
"start\_offset" : 74,  
"end\_offset" : 81,  
"type" : "shingle",  
"position" : 13  
}, {  
"token" : "lng",  
"start\_offset" : 78,  
"end\_offset" : 81,  
"type" : "word",  
"position" : 14  
}, {  
"token" : "lng 183731222",  
"start\_offset" : 78,  
"end\_offset" : 92,  
"type" : "shingle",  
"position" : 14  
}, {  
"token" : "183731222",  
"start\_offset" : 83,  
"end\_offset" : 92,  
"type" : "word",  
"position" : 15  
}, {  
"token" : "183731222 gps",  
"start\_offset" : 83,  
"end\_offset" : 97,  
"type" : "shingle",  
"position" : 15  
}, {  
"token" : "gps",  
"start\_offset" : 94,  
"end\_offset" : 97,  
"type" : "word",  
"position" : 16  
}, {  
"token" : "gps lat",  
"start\_offset" : 94,  
"end\_offset" : 101,  
"type" : "shingle",  
"position" : 16  
}, {  
"token" : "lat",  
"start\_offset" : 98,  
"end\_offset" : 101,  
"type" : "word",  
"position" : 17  
}, {  
"token" : "lat 289309222",  
"start\_offset" : 98,  
"end\_offset" : 112,  
"type" : "shingle",  
"position" : 17  
}, {  
"token" : "289309222",  
"start\_offset" : 103,  
"end\_offset" : 112,  
"type" : "word",  
"position" : 18  
} ]  
}

So it seems the template is not used?! Any obvious reason/mistakes?

Thx,  
Marc

On Thursday, August 28, 2014 6:17:08 PM UTC+2, Ivan Brusic wrote:

> Use the Analyze API to view what tokens are being generated? Keep it  
> simple at first (maybe remove shingles) and build up as you encounter more  
> edge-cases. What kind of query are you using?
> 
> --  
> Ivan
> 
> On Thu, Aug 28, 2014 at 2:05 AM, Marc \<[mn.o...@googlemail.com](mailto:mn.o...@googlemail.com)  
> \<javascript:\>\> wrote:
> 
> > Hi Ivan,
> > 
> > thanks for the help. Now it works almost... 😉  
> > I have used the following:  
> > "analysis": {  
> > "analyzer": {  
> > "msg\_excp\_analyzer": {  
> > "type": "custom",  
> > "tokenizer": "whitespace",  
> > "filters": ["split-up",  
> > "lowercase",  
> > "shingle",  
> > "ascii-folding"]  
> > }  
> > },  
> > "filter": {  
> > "split-up": {  
> > "type": "word\_delimiter",  
> > "preserve\_original": "true",  
> > "catenate\_all": "true",  
> > "type\_table": {  
> > "$": "DIGIT",  
> > "%": "DIGIT",  
> > ".": "DIGIT",  
> > ",": "DIGIT",  
> > ":": "DIGIT",  
> > "/": "DIGIT",  
> > "\": "DIGIT",  
> > "=": "DIGIT",  
> > "&": "DIGIT",  
> > "(": "DIGIT",  
> > ")": "DIGIT",  
> > "\<": "DIGIT",  
> > "\>": "DIGIT",  
> > "\U+000A": "DIGIT"  
> > }  
> > },  
> > "ascii-folding": {  
> > "type": "asciifolding",  
> > "preserve\_original": true  
> > }  
> > }  
> > If the above is wrong or not reasonable, please feel free to criticize!
> > 
> > Now the only thing that does not work is searching for subwords of  
> > concatenations with".".  
> > Having log Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22  
> > (updated attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ ) I cannot  
> > search for MyMDB or onMessage; only MyMDB.onMessage will work.
> > 
> > Anymore Ideas?
> > 
> > Cheers,  
> > Marc
> > 
> > On Wednesday, August 27, 2014 9:20:49 AM UTC+2, Ivan Brusic wrote:
> > 
> > > Off the top of my head, I would use a custom analyzer with a whitespace  
> > > tokenizer and a word delimiter filter (preserving the original tokens as  
> > > well). Perhaps a shingle filter to create bigrams. Or better yet a pattern  
> > > tokenizer with spaces and parenthesis.
> > > 
> > > Cheers,
> > > 
> > > Ivan
> > > 
> > > On Tue, Aug 26, 2014 at 11:57 PM, Marc [mn.o...@googlemail.com](mailto:mn.o...@googlemail.com) wrote:
> > > 
> > > > Hi,
> > > > 
> > > > I have quiet a simple scenario that already gives me a headache for  
> > > > quiet a while.  
> > > > I have one Field which is quiet big and full of special characters like  
> > > > (,),=,:,",' digits and text.  
> > > > Example:  
> > > > "msg" : "Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22  
> > > > (updated attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )"  
> > > > I essentially want to be able to search this things using text,  
> > > > wildcards etc.  
> > > > So far I have tried not analyzing the content and using the wildcard  
> > > > search and it doesn't work very well.  
> > > > Using different tokenizers and the query\_string query also only works  
> > > > to a certain degree.  
> > > > For example I want to be able to serach for following expressions:  
> > > > Service  
> > > > MyMDB  
> > > > onMessage  
> > > > MyMDB.onMessage  
> > > > appId=cs AND Times=Me:22
> > > > 
> > > > and other possible permutations.  
> > > > What is a correct setup?! I simply can't find a solution...
> > > > 
> > > > ps.: the data is imported to elasticsearch using logstash. We do acces  
> > > > the data using the java api (all software latest versions).
> > > > 
> > > > Cheeers,  
> > > > Marc
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google  
> > > > Groups "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).
> > > > 
> > > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > > msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%  
> > > > [40googlegroups.com](http://40googlegroups.com)  
> > > > [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > > .  
> > > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google Groups  
> > > "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send an  
> > > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > To view this discussion on the web visit  
> > > [https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/cb70139a-da96-41ab-9d6f-a5a2c19bfc0c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/cb70139a-da96-41ab-9d6f-a5a2c19bfc0c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 29, 2014, 4:49pm UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/6 "2014-08-29T16:49:31Z")

</div>

That output does not look like the something generated from the standard  
analyzer since it contains uppercase letters and various non-word  
characters such as '='.

Your two analysis requests will differ since the second one contains the  
default word\_delimiter filter instead of your custom my\_word\_delimiter.  
What you are trying to achieve is somewhat difficult, but you can get there  
if you keep on tweaking. 🙂 Try using a pattern tokenizer instead of the  
whitespace tokenizer if you want more control over word boundaries.

--  
Ivan

On Fri, Aug 29, 2014 at 1:48 AM, Marc [mn.offman@googlemail.com](mailto:mn.offman@googlemail.com) wrote:

> Hi Ivan,
> 
> thanks again. I have tried so and found a reasonable combination.  
> Nevertheless, when I now try to use the analyze api with an index that has  
> the said analyzer defined via template it doesn't seem to apply:
> 
> This is the complete template:  
> {  
> "template": "bogstash-\*",  
> "settings": {  
> "index.number\_of\_replicas": 0,  
> "analysis": {  
> "analyzer": {  
> "msg\_excp\_analyzer": {  
> "type": "custom",  
> "tokenizer": "whitespace",  
> "filters": ["word\_delimiter",  
> "lowercase",  
> "asciifolding",  
> "shingle",  
> "standard"]  
> }  
> },  
> "filters": {  
> "my\_word\_delimiter": {  
> "type": "word\_delimiter",  
> "preserve\_original": "true"  
> },  
> "my\_asciifolding": {  
> "type": "asciifolding",  
> "preserve\_original": true  
> }  
> }  
> }  
> },  
> "mappings": {  
> "_default_": {  
> "properties": {  
> "@excp": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> },  
> "@msg": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> }  
> }  
> }  
> }  
> }  
> I create the index bogstash-1.  
> Now I test the following:  
> curl -XGET  
> 'localhost:9200/bogstash-1/\_analyze?analyzer=msg\_excp\_analyzer&pretty=1' -d  
> 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> and it returns:  
> {  
> "tokens" : [ {  
> "token" : "Service=MyMDB.onMessage",  
> "start\_offset" : 0,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "appId=cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "Times=Me:22/Total:22",  
> "start\_offset" : 33,  
> "end\_offset" : 53,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "(updated",  
> "start\_offset" : 54,  
> "end\_offset" : 62,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "attributes=gps\_lng:",  
> "start\_offset" : 63,  
> "end\_offset" : 82,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "183731222/",  
> "start\_offset" : 83,  
> "end\_offset" : 93,  
> "type" : "word",  
> "position" : 6  
> }, {  
> "token" : "gps\_lat:",  
> "start\_offset" : 94,  
> "end\_offset" : 102,  
> "type" : "word",  
> "position" : 7  
> }, {  
> "token" : "289309222/",  
> "start\_offset" : 103,  
> "end\_offset" : 113,  
> "type" : "word",  
> "position" : 8  
> }, {  
> "token" : ")",  
> "start\_offset" : 114,  
> "end\_offset" : 115,  
> "type" : "word",  
> "position" : 9  
> } ]  
> }  
> Which is the output of a standard analyzer.  
> Giving the tokenizer and filters in the analyze API directly works fine:  
> curl -XGET  
> 'localhost:9200/\_analyze?tokenizer=whitespace&filters=lowercase,word\_delimiter,shingle,asciifolding,standard&pretty=1'  
> -d 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> This results in:  
> {  
> "tokens" : [ {  
> "token" : "service",  
> "start\_offset" : 0,  
> "end\_offset" : 7,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "service mymdb",  
> "start\_offset" : 0,  
> "end\_offset" : 13,  
> "type" : "shingle",  
> "position" : 1  
> }, {  
> "token" : "mymdb",  
> "start\_offset" : 8,  
> "end\_offset" : 13,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "mymdb onmessage",  
> "start\_offset" : 8,  
> "end\_offset" : 23,  
> "type" : "shingle",  
> "position" : 2  
> }, {  
> "token" : "onmessage",  
> "start\_offset" : 14,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "onmessage appid",  
> "start\_offset" : 14,  
> "end\_offset" : 29,  
> "type" : "shingle",  
> "position" : 3  
> }, {  
> "token" : "appid",  
> "start\_offset" : 24,  
> "end\_offset" : 29,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "appid cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "shingle",  
> "position" : 4  
> }, {  
> "token" : "cs",  
> "start\_offset" : 30,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "cs times",  
> "start\_offset" : 30,  
> "end\_offset" : 38,  
> "type" : "shingle",  
> "position" : 5  
> }, {  
> "token" : "times",  
> "start\_offset" : 33,  
> "end\_offset" : 38,  
> "type" : "word",  
> "position" : 6  
> }, {  
> "token" : "times me",  
> "start\_offset" : 33,  
> "end\_offset" : 41,  
> "type" : "shingle",  
> "position" : 6  
> }, {  
> "token" : "me",  
> "start\_offset" : 39,  
> "end\_offset" : 41,  
> "type" : "word",  
> "position" : 7  
> }, {  
> "token" : "me 22",  
> "start\_offset" : 39,  
> "end\_offset" : 44,  
> "type" : "shingle",  
> "position" : 7  
> }, {  
> "token" : "22",  
> "start\_offset" : 42,  
> "end\_offset" : 44,  
> "type" : "word",  
> "position" : 8  
> }, {  
> "token" : "22 total",  
> "start\_offset" : 42,  
> "end\_offset" : 50,  
> "type" : "shingle",  
> "position" : 8  
> }, {  
> "token" : "total",  
> "start\_offset" : 45,  
> "end\_offset" : 50,  
> "type" : "word",  
> "position" : 9  
> }, {  
> "token" : "total 22",  
> "start\_offset" : 45,  
> "end\_offset" : 53,  
> "type" : "shingle",  
> "position" : 9  
> }, {  
> "token" : "22",  
> "start\_offset" : 51,  
> "end\_offset" : 53,  
> "type" : "word",  
> "position" : 10  
> }, {  
> "token" : "22 updated",  
> "start\_offset" : 51,  
> "end\_offset" : 62,  
> "type" : "shingle",  
> "position" : 10  
> }, {  
> "token" : "updated",  
> "start\_offset" : 55,  
> "end\_offset" : 62,  
> "type" : "word",  
> "position" : 11  
> }, {  
> "token" : "updated attributes",  
> "start\_offset" : 55,  
> "end\_offset" : 73,  
> "type" : "shingle",  
> "position" : 11  
> }, {  
> "token" : "attributes",  
> "start\_offset" : 63,  
> "end\_offset" : 73,  
> "type" : "word",  
> "position" : 12  
> }, {  
> "token" : "attributes gps",  
> "start\_offset" : 63,  
> "end\_offset" : 77,  
> "type" : "shingle",  
> "position" : 12  
> }, {  
> "token" : "gps",  
> "start\_offset" : 74,  
> "end\_offset" : 77,  
> "type" : "word",  
> "position" : 13  
> }, {  
> "token" : "gps lng",  
> "start\_offset" : 74,  
> "end\_offset" : 81,  
> "type" : "shingle",  
> "position" : 13  
> }, {  
> "token" : "lng",  
> "start\_offset" : 78,  
> "end\_offset" : 81,  
> "type" : "word",  
> "position" : 14  
> }, {  
> "token" : "lng 183731222",  
> "start\_offset" : 78,  
> "end\_offset" : 92,  
> "type" : "shingle",  
> "position" : 14  
> }, {  
> "token" : "183731222",  
> "start\_offset" : 83,  
> "end\_offset" : 92,  
> "type" : "word",  
> "position" : 15  
> }, {  
> "token" : "183731222 gps",  
> "start\_offset" : 83,  
> "end\_offset" : 97,  
> "type" : "shingle",  
> "position" : 15  
> }, {  
> "token" : "gps",  
> "start\_offset" : 94,  
> "end\_offset" : 97,  
> "type" : "word",  
> "position" : 16  
> }, {  
> "token" : "gps lat",  
> "start\_offset" : 94,  
> "end\_offset" : 101,  
> "type" : "shingle",  
> "position" : 16  
> }, {  
> "token" : "lat",  
> "start\_offset" : 98,  
> "end\_offset" : 101,  
> "type" : "word",  
> "position" : 17  
> }, {  
> "token" : "lat 289309222",  
> "start\_offset" : 98,  
> "end\_offset" : 112,  
> "type" : "shingle",  
> "position" : 17  
> }, {  
> "token" : "289309222",  
> "start\_offset" : 103,  
> "end\_offset" : 112,  
> "type" : "word",  
> "position" : 18  
> } ]  
> }
> 
> So it seems the template is not used?! Any obvious reason/mistakes?
> 
> Thx,  
> Marc
> 
> On Thursday, August 28, 2014 6:17:08 PM UTC+2, Ivan Brusic wrote:
> 
> > Use the Analyze API to view what tokens are being generated? Keep it  
> > simple at first (maybe remove shingles) and build up as you encounter more  
> > edge-cases. What kind of query are you using?
> > 
> > --  
> > Ivan
> > 
> > On Thu, Aug 28, 2014 at 2:05 AM, Marc [mn.o...@googlemail.com](mailto:mn.o...@googlemail.com) wrote:
> > 
> > > Hi Ivan,
> > > 
> > > thanks for the help. Now it works almost... 😉  
> > > I have used the following:  
> > > "analysis": {  
> > > "analyzer": {  
> > > "msg\_excp\_analyzer": {  
> > > "type": "custom",  
> > > "tokenizer": "whitespace",  
> > > "filters": ["split-up",  
> > > "lowercase",  
> > > "shingle",  
> > > "ascii-folding"]  
> > > }  
> > > },  
> > > "filter": {  
> > > "split-up": {  
> > > "type": "word\_delimiter",  
> > > "preserve\_original": "true",  
> > > "catenate\_all": "true",  
> > > "type\_table": {  
> > > "$": "DIGIT",  
> > > "%": "DIGIT",  
> > > ".": "DIGIT",  
> > > ",": "DIGIT",  
> > > ":": "DIGIT",  
> > > "/": "DIGIT",  
> > > "\": "DIGIT",  
> > > "=": "DIGIT",  
> > > "&": "DIGIT",  
> > > "(": "DIGIT",  
> > > ")": "DIGIT",  
> > > "\<": "DIGIT",  
> > > "\>": "DIGIT",  
> > > "\U+000A": "DIGIT"  
> > > }  
> > > },  
> > > "ascii-folding": {  
> > > "type": "asciifolding",  
> > > "preserve\_original": true  
> > > }  
> > > }  
> > > If the above is wrong or not reasonable, please feel free to criticize!
> > > 
> > > Now the only thing that does not work is searching for subwords of  
> > > concatenations with".".  
> > > Having log Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22  
> > > (updated attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ ) I cannot  
> > > search for MyMDB or onMessage; only MyMDB.onMessage will work.
> > > 
> > > Anymore Ideas?
> > > 
> > > Cheers,  
> > > Marc
> > > 
> > > On Wednesday, August 27, 2014 9:20:49 AM UTC+2, Ivan Brusic wrote:
> > > 
> > > > Off the top of my head, I would use a custom analyzer with a whitespace  
> > > > tokenizer and a word delimiter filter (preserving the original tokens as  
> > > > well). Perhaps a shingle filter to create bigrams. Or better yet a pattern  
> > > > tokenizer with spaces and parenthesis.
> > > > 
> > > > Cheers,
> > > > 
> > > > Ivan
> > > > 
> > > > On Tue, Aug 26, 2014 at 11:57 PM, Marc [mn.o...@googlemail.com](mailto:mn.o...@googlemail.com) wrote:
> > > > 
> > > > > Hi,
> > > > > 
> > > > > I have quiet a simple scenario that already gives me a headache for  
> > > > > quiet a while.  
> > > > > I have one Field which is quiet big and full of special characters  
> > > > > like (,),=,:,",' digits and text.  
> > > > > Example:  
> > > > > "msg" : "Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22  
> > > > > (updated attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )"  
> > > > > I essentially want to be able to search this things using text,  
> > > > > wildcards etc.  
> > > > > So far I have tried not analyzing the content and using the wildcard  
> > > > > search and it doesn't work very well.  
> > > > > Using different tokenizers and the query\_string query also only works  
> > > > > to a certain degree.  
> > > > > For example I want to be able to serach for following expressions:  
> > > > > Service  
> > > > > MyMDB  
> > > > > onMessage  
> > > > > MyMDB.onMessage  
> > > > > appId=cs AND Times=Me:22
> > > > > 
> > > > > and other possible permutations.  
> > > > > What is a correct setup?! I simply can't find a solution...
> > > > > 
> > > > > ps.: the data is imported to elasticsearch using logstash. We do acces  
> > > > > the data using the java api (all software latest versions).
> > > > > 
> > > > > Cheeers,  
> > > > > Marc
> > > > > 
> > > > > --  
> > > > > You received this message because you are subscribed to the Google  
> > > > > Groups "elasticsearch" group.  
> > > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).
> > > > > 
> > > > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > > > msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40goo  
> > > > > [glegroups.com](http://glegroups.com)  
> > > > > [https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ada9c759-41e0-46ad-9941-3a0f2fb7c122%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > > > .  
> > > > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google  
> > > > Groups "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > > msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%  
> > > > [40googlegroups.com](http://40googlegroups.com)  
> > > > [https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/a4350999-f089-4b52-bccd-d10821630066%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > > .
> > > 
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/cb70139a-da96-41ab-9d6f-a5a2c19bfc0c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/cb70139a-da96-41ab-9d6f-a5a2c19bfc0c%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/cb70139a-da96-41ab-9d6f-a5a2c19bfc0c%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/cb70139a-da96-41ab-9d6f-a5a2c19bfc0c%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQAU6AchiAU6F%2BbTf0LyOecL2YjpLY%2B\_e\_KF2SjuQWKL0g%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQAU6AchiAU6F%2BbTf0LyOecL2YjpLY%2B_e_KF2SjuQWKL0g%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Marc\_2](https://avatars.discourse-cdn.com/v4/letter/m/c4cdca/32.png) [@Marc\_2](https://discuss.elastic.co/u/Marc_2)\
**Post date:** [September 1, 2014, 11:15am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/7 "2014-09-01T11:15:27Z")

</div>

Hi Ivan,

Using a test index and the analyze API, I was no able to create a config,  
which is fine for me... theoretically.  
{  
"template": "logstash-\*",  
"settings": {  
"analysis": {  
"filter": {  
"my\_word\_delimiter": {  
"type": "word\_delimiter",  
"preserve\_original": "true"  
}  
},  
"analyzer": {  
"b2v\_analyzer": {  
"type": "custom",  
"tokenizer": "standard",  
"filter": ["standard",  
"lowercase",  
"stop",  
"my\_word\_delimiter",  
"asciifolding"]  
}  
}  
}  
},  
"mappings": {  
"_default_": {  
"properties": {  
"excp": {  
"type": "string",  
"index": "analyzed",  
"analyzer": "b2v\_analyzer"  
},  
"msg": {  
"type": "string",  
"index": "not\_analyzed",  
"analyzer": "b2v\_analyzer"  
}  
}  
}  
}  
}  
The problem now is, as soon as I activate this for the two fields and have  
a new logstash index created I cannot use a simpleQueryString query to  
retrieve any results.  
It won't find anything via the REST api. Using the standard logstash  
template and mapping it works fine.  
Have you observed anything simililar?

Thx  
Marc

On Friday, August 29, 2014 6:49:41 PM UTC+2, Ivan Brusic wrote:

> That output does not look like the something generated from the standard  
> analyzer since it contains uppercase letters and various non-word  
> characters such as '='.
> 
> Your two analysis requests will differ since the second one contains the  
> default word\_delimiter filter instead of your custom my\_word\_delimiter.  
> What you are trying to achieve is somewhat difficult, but you can get there  
> if you keep on tweaking. 🙂 Try using a pattern tokenizer instead of the  
> whitespace tokenizer if you want more control over word boundaries.
> 
> --  
> Ivan
> 
> On Fri, Aug 29, 2014 at 1:48 AM, Marc \<[mn.o...@googlemail.com](mailto:mn.o...@googlemail.com)  
> \<javascript:\>\> wrote:
> 
> Hi Ivan,
> 
> thanks again. I have tried so and found a reasonable combination.  
> Nevertheless, when I now try to use the analyze api with an index that has  
> the said analyzer defined via template it doesn't seem to apply:
> 
> This is the complete template:  
> {  
> "template": "bogstash-\*",  
> "settings": {  
> "index.number\_of\_replicas": 0,  
> "analysis": {  
> "analyzer": {  
> "msg\_excp\_analyzer": {  
> "type": "custom",  
> "tokenizer": "whitespace",  
> "filters": ["word\_delimiter",  
> "lowercase",  
> "asciifolding",  
> "shingle",  
> "standard"]  
> }  
> },  
> "filters": {  
> "my\_word\_delimiter": {  
> "type": "word\_delimiter",  
> "preserve\_original": "true"  
> },  
> "my\_asciifolding": {  
> "type": "asciifolding",  
> "preserve\_original": true  
> }  
> }  
> }  
> },  
> "mappings": {  
> "_default_": {  
> "properties": {  
> "@excp": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> },  
> "@msg": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> }  
> }  
> }  
> }  
> }  
> I create the index bogstash-1.  
> Now I test the following:  
> curl -XGET  
> 'localhost:9200/bogstash-1/\_analyze?analyzer=msg\_excp\_analyzer&pretty=1' -d  
> 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> and it returns:  
> {  
> "tokens" : [ {  
> "token" : "Service=MyMDB.onMessage",  
> "start\_offset" : 0,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "appId=cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "Times=Me:22/Total:22",  
> "start\_offset" : 33,  
> "end\_offset" : 53,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "(updated",  
> "start\_offset" : 54,  
> "end\_offset" : 62,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "attributes=gps\_lng:",  
> "start\_offset" : 63,  
> "end\_offset" : 82,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "183731222/",  
> "start\_offset" : 83,  
> "end\_offset" : 93,  
> "type" : "word",  
> "position" : 6  
> }, {  
> "token" : "gps\_lat:",  
> "start\_offset" : 94,  
> "end\_offset" : 102,  
> "type" : "word",  
> "position" : 7  
> }, {  
> "token" : "289309222/",  
> "start\_offset" : 103,  
> "end\_offset" : 113,  
> "type" : "word",  
> "position" : 8  
> }, {  
> "token" : ")",  
> "start\_offset" : 114,  
> "end\_offset" : 115,  
> "type" : "word",  
> "position" : 9  
> } ]  
> }  
> Which is the output of a standard analyzer.  
> Giving the tokenizer and filters in the analyze API directly works fine:  
> curl -XGET  
> 'localhost:9200/\_analyze?tokenizer=whitespace&filters=lowercase,word\_delimiter,shingle,asciifolding,standard&pretty=1'  
> -d 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> This results in:  
> {  
> "tokens" : [ {  
> "token" : "service",  
> "start\_offset" : 0,  
> "end\_offset" : 7,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "service mymdb",  
> "start\_offset" : 0,  
> "end\_offset" : 13,  
> "type" : "shingle",  
> "position" : 1  
> }, {  
> "token" : "mymdb",  
> "start\_offset" : 8,  
> "end\_offset" : 13,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "mymdb onmessage",  
> "start\_offset" : 8,  
> "end\_offset" : 23,  
> "type" : "shingle",  
> "position" : 2  
> }, {  
> "token" : "onmessage",  
> "start\_offset" : 14,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "onmessage appid",  
> "start\_offset" : 14,  
> "end\_offset" : 29,  
> "type" : "shingle",  
> "position" : 3  
> }, {  
> "token" : "appid",  
> "start\_offset" : 24,  
> "end\_offset" : 29,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "appid cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "shingle",  
> "position" : 4  
> }, {  
> "token" : "cs",  
> "start\_offset" : 30,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "cs times",  
> "start\_offset" : 30,  
> "end\_offset" : 38,  
> "type" : "shingle",  
> "position" : 5  
> }, {  
> "token" : "times",  
> "start\_offset" : 33,  
> "end\_offset" : \<span style="color:#06
> 
> ...

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ddd5fc06-e4e1-41e7-9810-cdc60a2c9aea%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ddd5fc06-e4e1-41e7-9810-cdc60a2c9aea%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Marc\_2](https://avatars.discourse-cdn.com/v4/letter/m/c4cdca/32.png) [@Marc\_2](https://discuss.elastic.co/u/Marc_2)\
**Post date:** [September 1, 2014, 11:16am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/8 "2014-09-01T11:16:43Z")

</div>

Hi Ivan,

Using a test index and the analyze API, I was no able to create a config,  
which is fine for me... theoretically.  
{  
"template": "logstash-\*",  
"settings": {  
"analysis": {  
"filter": {

```
            "my_word_delimiter": {
                "type": "word_delimiter",
                "preserve_original": "true"
            }
        },

        "analyzer": {
            "my_analyzer": {
                "type": "custom",
                "tokenizer": "standard",
                "filter": ["standard",
                "lowercase",
                "stop",
                "my_word_delimiter",
                "asciifolding"]
            }
        }
    }
},
"mappings": {
    "_default_": {
        "properties": {

            "excp": {
                "type": "string",
                "index": "analyzed",

                "analyzer": "my_analyzer"
            },
            "msg": {
                "type": "string",
                "index": "not_analyzed",
                "analyzer": "my_analyzer"
            }
        }
    }
}

```

}  
The problem now is, as soon as I activate this for the two fields and have  
a new logstash index created I cannot use a simpleQueryString query to  
retrieve any results.  
It won't find anything via the REST api. Using the standard logstash  
template and mapping it works fine.  
Have you observed anything simililar?

Thx  
Marc

On Friday, August 29, 2014 6:49:41 PM UTC+2, Ivan Brusic wrote:

> That output does not look like the something generated from the standard  
> analyzer since it contains uppercase letters and various non-word  
> characters such as '='.
> 
> Your two analysis requests will differ since the second one contains the  
> default word\_delimiter filter instead of your custom my\_word\_delimiter.  
> What you are trying to achieve is somewhat difficult, but you can get there  
> if you keep on tweaking. 🙂 Try using a pattern tokenizer instead of the  
> whitespace tokenizer if you want more control over word boundaries.
> 
> --  
> Ivan
> 
> On Fri, Aug 29, 2014 at 1:48 AM, Marc \<[mn.o...@googlemail.com](mailto:mn.o...@googlemail.com)  
> \<javascript:\>\> wrote:
> 
> Hi Ivan,
> 
> thanks again. I have tried so and found a reasonable combination.  
> Nevertheless, when I now try to use the analyze api with an index that has  
> the said analyzer defined via template it doesn't seem to apply:
> 
> This is the complete template:  
> {  
> "template": "bogstash-\*",  
> "settings": {  
> "index.number\_of\_replicas": 0,  
> "analysis": {  
> "analyzer": {  
> "msg\_excp\_analyzer": {  
> "type": "custom",  
> "tokenizer": "whitespace",  
> "filters": ["word\_delimiter",  
> "lowercase",  
> "asciifolding",  
> "shingle",  
> "standard"]  
> }  
> },  
> "filters": {  
> "my\_word\_delimiter": {  
> "type": "word\_delimiter",  
> "preserve\_original": "true"  
> },  
> "my\_asciifolding": {  
> "type": "asciifolding",  
> "preserve\_original": true  
> }  
> }  
> }  
> },  
> "mappings": {  
> "_default_": {  
> "properties": {  
> "@excp": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> },  
> "@msg": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> }  
> }  
> }  
> }  
> }  
> I create the index bogstash-1.  
> Now I test the following:  
> curl -XGET  
> 'localhost:9200/bogstash-1/\_analyze?analyzer=msg\_excp\_analyzer&pretty=1' -d  
> 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> and it returns:  
> {  
> "tokens" : [ {  
> "token" : "Service=MyMDB.onMessage",  
> "start\_offset" : 0,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "appId=cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "Times=Me:22/Total:22",  
> "start\_offset" : 33,  
> "end\_offset" : 53,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "(updated",  
> "start\_offset" : 54,  
> "end\_offset" : 62,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "attributes=gps\_lng:",  
> "start\_offset" : 63,  
> "end\_offset" : 82,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "183731222/",  
> "start\_offset" : 83,  
> "end\_offset" : 93,  
> "type" : "word",  
> "position" : 6  
> }, {  
> "token" : "gps\_lat:",  
> "start\_offset" : 94,  
> "end\_offset" : 102,  
> "type" : "word",  
> "position" : 7  
> }, {  
> "token" : "289309222/",  
> "start\_offset" : 103,  
> "end\_offset" : 113,  
> "type" : "word",  
> "position" : 8  
> }, {  
> "token" : ")",  
> "start\_offset" : 114,  
> "end\_offset" : 115,  
> "type" : "word",  
> "position" : 9  
> } ]  
> }  
> Which is the output of a standard analyzer.  
> Giving the tokenizer and filters in the analyze API directly works fine:  
> curl -XGET  
> 'localhost:9200/\_analyze?tokenizer=whitespace&filters=lowercase,word\_delimiter,shingle,asciifolding,standard&pretty=1'  
> -d 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> This results in:  
> {  
> "tokens" : [ {  
> "token" : "service",  
> "start\_offset" : 0,  
> "end\_offset" : 7,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "service mymdb",  
> "start\_offset" : 0,  
> "end\_offset" : 13,  
> "type" : "shingle",  
> "position" : 1  
> }, {  
> "token" : "mymdb",  
> "start\_offset" : 8,  
> "end\_offset" : 13,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "mymdb onmessage",  
> "start\_offset" : 8,  
> "end\_offset" : 23,  
> "type" : "shingle",  
> "position" : 2  
> }, {  
> "token" : "onmessage",  
> "start\_offset" : 14,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "onmessage appid",  
> "start\_offset" : 14,  
> "end\_offset" : 29,  
> "type" : "shingle",  
> "position" : 3  
> }, {  
> "token" : "appid",  
> "start\_offset" : 24,  
> "end\_offset" : 29,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "appid cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "shingle",  
> "position" : 4  
> }, {  
> "token" : "cs",  
> "start\_offset" : 30,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "cs times",  
> "start\_offset" : 30,  
> "end\_offset" : 38,  
> "type" : "shingle",  
> "position" : 5  
> }, {  
> "token" : "times",  
> "start\_offset" : 33,  
> "end\_offset" : \<span style="color:#06
> 
> ...

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/9bf32a56-a490-44e9-8efd-676587c22621%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9bf32a56-a490-44e9-8efd-676587c22621%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Marc\_2](https://avatars.discourse-cdn.com/v4/letter/m/c4cdca/32.png) [@Marc\_2](https://discuss.elastic.co/u/Marc_2)\
**Post date:** [September 1, 2014, 11:17am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/9 "2014-09-01T11:17:52Z")

</div>

Hi Ivan,

Using a test index and the analyze API, I was no able to create a config,  
which is fine for me... theoretically.  
{  
"template": "logstash-\*",  
"settings": {  
"analysis": {  
"filter": {

```
            "my_word_delimiter": {
                "type": "word_delimiter",
                "preserve_original": "true"
            }
        },

        "analyzer": {
            "my_analyzer": {
                "type": "custom",
                "tokenizer": "standard",
                "filter": ["standard",
                "lowercase",
                "stop",
                "my_word_delimiter",
                "asciifolding"]
            }
        }
    }
},
"mappings": {
    "_default_": {
        "properties": {

            "excp": {
                "type": "string",
                "index": "analyzed",
                "analyzer": "my_analyzer"
            },
            "msg": {
                "type": "string",
                "index": "analyzed",
                "analyzer": "my_analyzer"
            }
        }
    }
}

```

}  
The problem now is, as soon as I activate this for the two fields and have  
a new logstash index created I cannot use a simpleQueryString query to  
retrieve any results.  
It won't find anything via the REST api. Using the standard logstash  
template and mapping it works fine.  
Have you observed anything simililar?

Thx  
Marc

On Friday, August 29, 2014 6:49:41 PM UTC+2, Ivan Brusic wrote:

> That output does not look like the something generated from the standard  
> analyzer since it contains uppercase letters and various non-word  
> characters such as '='.
> 
> Your two analysis requests will differ since the second one contains the  
> default word\_delimiter filter instead of your custom my\_word\_delimiter.  
> What you are trying to achieve is somewhat difficult, but you can get there  
> if you keep on tweaking. 🙂 Try using a pattern tokenizer instead of the  
> whitespace tokenizer if you want more control over word boundaries.
> 
> --  
> Ivan
> 
> On Fri, Aug 29, 2014 at 1:48 AM, Marc \<[mn.o...@googlemail.com](mailto:mn.o...@googlemail.com)  
> \<javascript:\>\> wrote:
> 
> Hi Ivan,
> 
> thanks again. I have tried so and found a reasonable combination.  
> Nevertheless, when I now try to use the analyze api with an index that has  
> the said analyzer defined via template it doesn't seem to apply:
> 
> This is the complete template:  
> {  
> "template": "bogstash-\*",  
> "settings": {  
> "index.number\_of\_replicas": 0,  
> "analysis": {  
> "analyzer": {  
> "msg\_excp\_analyzer": {  
> "type": "custom",  
> "tokenizer": "whitespace",  
> "filters": ["word\_delimiter",  
> "lowercase",  
> "asciifolding",  
> "shingle",  
> "standard"]  
> }  
> },  
> "filters": {  
> "my\_word\_delimiter": {  
> "type": "word\_delimiter",  
> "preserve\_original": "true"  
> },  
> "my\_asciifolding": {  
> "type": "asciifolding",  
> "preserve\_original": true  
> }  
> }  
> }  
> },  
> "mappings": {  
> "_default_": {  
> "properties": {  
> "@excp": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> },  
> "@msg": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> }  
> }  
> }  
> }  
> }  
> I create the index bogstash-1.  
> Now I test the following:  
> curl -XGET  
> 'localhost:9200/bogstash-1/\_analyze?analyzer=msg\_excp\_analyzer&pretty=1' -d  
> 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> and it returns:  
> {  
> "tokens" : [ {  
> "token" : "Service=MyMDB.onMessage",  
> "start\_offset" : 0,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "appId=cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "Times=Me:22/Total:22",  
> "start\_offset" : 33,  
> "end\_offset" : 53,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "(updated",  
> "start\_offset" : 54,  
> "end\_offset" : 62,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "attributes=gps\_lng:",  
> "start\_offset" : 63,  
> "end\_offset" : 82,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "183731222/",  
> "start\_offset" : 83,  
> "end\_offset" : 93,  
> "type" : "word",  
> "position" : 6  
> }, {  
> "token" : "gps\_lat:",  
> "start\_offset" : 94,  
> "end\_offset" : 102,  
> "type" : "word",  
> "position" : 7  
> }, {  
> "token" : "289309222/",  
> "start\_offset" : 103,  
> "end\_offset" : 113,  
> "type" : "word",  
> "position" : 8  
> }, {  
> "token" : ")",  
> "start\_offset" : 114,  
> "end\_offset" : 115,  
> "type" : "word",  
> "position" : 9  
> } ]  
> }  
> Which is the output of a standard analyzer.  
> Giving the tokenizer and filters in the analyze API directly works fine:  
> curl -XGET  
> 'localhost:9200/\_analyze?tokenizer=whitespace&filters=lowercase,word\_delimiter,shingle,asciifolding,standard&pretty=1'  
> -d 'Service=MyMDB.onMessage appId=cs Times=Me:22/Total:22 (updated  
> attributes=gps\_lng: 183731222/ gps\_lat: 289309222/ )'  
> This results in:  
> {  
> "tokens" : [ {  
> "token" : "service",  
> "start\_offset" : 0,  
> "end\_offset" : 7,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "service mymdb",  
> "start\_offset" : 0,  
> "end\_offset" : 13,  
> "type" : "shingle",  
> "position" : 1  
> }, {  
> "token" : "mymdb",  
> "start\_offset" : 8,  
> "end\_offset" : 13,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "mymdb onmessage",  
> "start\_offset" : 8,  
> "end\_offset" : 23,  
> "type" : "shingle",  
> "position" : 2  
> }, {  
> "token" : "onmessage",  
> "start\_offset" : 14,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "onmessage appid",  
> "start\_offset" : 14,  
> "end\_offset" : 29,  
> "type" : "shingle",  
> "position" : 3  
> }, {  
> "token" : "appid",  
> "start\_offset" : 24,  
> "end\_offset" : 29,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "appid cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "shingle",  
> "position" : 4  
> }, {  
> "token" : "cs",  
> "start\_offset" : 30,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "cs times",  
> "start\_offset" : 30,  
> "end\_offset" : 38,  
> "type" : "shingle",  
> "position" : 5  
> }, {  
> "token" : "times",  
> "start\_offset" : 33,  
> "end\_offset" : \<span style="color:#06
> 
> ...

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/8ef1a8eb-3e4b-413a-b751-6e0d84bfca6a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8ef1a8eb-3e4b-413a-b751-6e0d84bfca6a%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 2, 2014, 5:54pm UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/10 "2014-09-02T17:54:18Z")

</div>

Hard to say without looking at your query, but perhaps you are experiencing  
query parser issues. The query string query uses the standard query parser,  
which might does not tokenize terms in the way your custom tokenizer might.  
Try using match queries, which does not use the query parser to see if it  
"fixes" the problem. Of course, you will have have the query syntax at your  
disposal, but you can find workarounds.

--  
Ivan

On Mon, Sep 1, 2014 at 4:17 AM, Marc [mn.offman@googlemail.com](mailto:mn.offman@googlemail.com) wrote:

> Hi Ivan,
> 
> Using a test index and the analyze API, I was no able to create a config,  
> which is fine for me... theoretically.  
> {  
> "template": "logstash-\*",  
> "settings": {  
> "analysis": {  
> "filter": {
> 
> ```
> "my_word_delimiter": {
> 
> "type": "word_delimiter",
> "preserve_original": "true"
> }
> },
> 
> "analyzer": {
> 
> "my_analyzer": {
> "type": "custom",
> "tokenizer": "standard",
> "filter": ["standard",
> "lowercase",
> "stop",
> "my_word_delimiter",
> "asciifolding"]
> }
> }
> }
> },
> "mappings": {
> "_default_": {
> "properties": {
> 
> "excp": {
> 
> "type": "string",
> "index": "analyzed",
> "analyzer": "my_analyzer"
> },
> "msg": {
> 
> "type": "string",
> "index": "analyzed",
> 
> "analyzer": "my_analyzer"
> }
> }
> }
> }
> 
> ```
> 
> }  
> The problem now is, as soon as I activate this for the two fields and have  
> a new logstash index created I cannot use a simpleQueryString query to  
> retrieve any results.  
> It won't find anything via the REST api. Using the standard logstash  
> template and mapping it works fine.  
> Have you observed anything simililar?
> 
> Thx  
> Marc
> 
> On Friday, August 29, 2014 6:49:41 PM UTC+2, Ivan Brusic wrote:
> 
> > That output does not look like the something generated from the standard  
> > analyzer since it contains uppercase letters and various non-word  
> > characters such as '='.
> > 
> > Your two analysis requests will differ since the second one contains the  
> > default word\_delimiter filter instead of your custom my\_word\_delimiter.  
> > What you are trying to achieve is somewhat difficult, but you can get there  
> > if you keep on tweaking. 🙂 Try using a pattern tokenizer instead of the  
> > whitespace tokenizer if you want more control over word boundaries.
> > 
> > --  
> > Ivan
> > 
> > On Fri, Aug 29, 2014 at 1:48 AM, Marc [mn.o...@googlemail.com](mailto:mn.o...@googlemail.com) wrote:
> > 
> > Hi Ivan,
> > 
> > thanks again. I have tried so and found a reasonable combination.  
> > Nevertheless, when I now try to use the analyze api with an index that  
> > has the said analyzer defined via template it doesn't seem to apply:
> > 
> > This is the complete template:  
> > {  
> > "template": "bogstash-\*",  
> > "settings": {  
> > "index.number\_of\_replicas": 0,  
> > "analysis": {  
> > "analyzer": {  
> > "msg\_excp\_analyzer": {  
> > "type": "custom",  
> > "tokenizer": "whitespace",  
> > "filters": ["word\_delimiter",  
> > "lowercase",  
> > "asciifolding",  
> > "shingle",  
> > "standard"]  
> > }  
> > },  
> > "filters": {  
> > "my\_word\_delimiter": {  
> > "type": "word\_delimiter",  
> > "preserve\_original": "true"  
> > },  
> > "my\_asciifolding": {  
> > "type": "asciifolding",  
> > "preserve\_original": true  
> > }  
> > }  
> > }  
> > },  
> > "mappings": {  
> > "_default_": {  
> > "properties": {  
> > "@excp": {  
> > "type": "string",  
> > "index": "analyzed",  
> > "analyzer": "msg\_excp\_analyzer"  
> > },  
> > "@msg": {  
> > "type": "string",  
> > "index": "analyzed",  
> > "analyzer": "msg\_excp\_analyzer"  
> > }  
> > }  
> > }  
> > }  
> > }  
> > I create the index bogstash-1.  
> > Now I test the following:  
> > curl -XGET 'localhost:9200/bogstash-1/_analyze?analyzer=msg\_excp_  
> > analyzer&pretty=1' -d 'Service=MyMDB.onMessage appId=cs  
> > Times=Me:22/Total:22 (updated attributes=gps\_lng: 183731222/ gps\_lat:  
> > 289309222/ )'  
> > and it returns:  
> > {  
> > "tokens" : [ {  
> > "token" : "Service=MyMDB.onMessage",  
> > "start\_offset" : 0,  
> > "end\_offset" : 23,  
> > "type" : "word",  
> > "position" : 1  
> > }, {  
> > "token" : "appId=cs",  
> > "start\_offset" : 24,  
> > "end\_offset" : 32,  
> > "type" : "word",  
> > "position" : 2  
> > }, {  
> > "token" : "Times=Me:22/Total:22",  
> > "start\_offset" : 33,  
> > "end\_offset" : 53,  
> > "type" : "word",  
> > "position" : 3  
> > }, {  
> > "token" : "(updated",  
> > "start\_offset" : 54,  
> > "end\_offset" : 62,  
> > "type" : "word",  
> > "position" : 4  
> > }, {  
> > "token" : "attributes=gps\_lng:",  
> > "start\_offset" : 63,  
> > "end\_offset" : 82,  
> > "type" : "word",  
> > "position" : 5  
> > }, {  
> > "token" : "183731222/",  
> > "start\_offset" : 83,  
> > "end\_offset" : 93,  
> > "type" : "word",  
> > "position" : 6  
> > }, {  
> > "token" : "gps\_lat:",  
> > "start\_offset" : 94,  
> > "end\_offset" : 102,  
> > "type" : "word",  
> > "position" : 7  
> > }, {  
> > "token" : "289309222/",  
> > "start\_offset" : 103,  
> > "end\_offset" : 113,  
> > "type" : "word",  
> > "position" : 8  
> > }, {  
> > "token" : ")",  
> > "start\_offset" : 114,  
> > "end\_offset" : 115,  
> > "type" : "word",  
> > "position" : 9  
> > } ]  
> > }  
> > Which is the output of a standard analyzer.  
> > Giving the tokenizer and filters in the analyze API directly works fine:  
> > curl -XGET 'localhost:9200/\_analyze?tokenizer=whitespace&filters=  
> > lowercase,word\_delimiter,shingle,asciifolding,standard&pretty=1' -d 'Service=MyMDB.onMessage  
> > appId=cs Times=Me:22/Total:22 (updated attributes=gps\_lng: 183731222/  
> > gps\_lat: 289309222/ )'  
> > This results in:  
> > {  
> > "tokens" : [ {  
> > "token" : "service",  
> > "start\_offset" : 0,  
> > "end\_offset" : 7,  
> > "type" : "word",  
> > "position" : 1  
> > }, {  
> > "token" : "service mymdb",  
> > "start\_offset" : 0,  
> > "end\_offset" : 13,  
> > "type" : "shingle",  
> > "position" : 1  
> > }, {  
> > "token" : "mymdb",  
> > "start\_offset" : 8,  
> > "end\_offset" : 13,  
> > "type" : "word",  
> > "position" : 2  
> > }, {  
> > "token" : "mymdb onmessage",  
> > "start\_offset" : 8,  
> > "end\_offset" : 23,  
> > "type" : "shingle",  
> > "position" : 2  
> > }, {  
> > "token" : "onmessage",  
> > "start\_offset" : 14,  
> > "end\_offset" : 23,  
> > "type" : "word",  
> > "position" : 3  
> > }, {  
> > "token" : "onmessage appid",  
> > "start\_offset" : 14,  
> > "end\_offset" : 29,  
> > "type" : "shingle",  
> > "position" : 3  
> > }, {  
> > "token" : "appid",  
> > "start\_offset" : 24,  
> > "end\_offset" : 29,  
> > "type" : "word",  
> > "position" : 4  
> > }, {  
> > "token" : "appid cs",  
> > "start\_offset" : 24,  
> > "end\_offset" : 32,  
> > "type" : "shingle",  
> > "position" : 4  
> > }, {  
> > "token" : "cs",  
> > "start\_offset" : 30,  
> > "end\_offset" : 32,  
> > "type" : "word",  
> > "position" : 5  
> > }, {  
> > "token" : "cs times",  
> > "start\_offset" : 30,  
> > "end\_offset" : 38,  
> > "type" : "shingle",  
> > "position" : 5  
> > }, {  
> > "token" : "times",  
> > "start\_offset" : 33,
> > 
> > ```
> > "end_offset" : <span style="color:#06
> > 
> > ```
> > 
> > ...
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/8ef1a8eb-3e4b-413a-b751-6e0d84bfca6a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8ef1a8eb-3e4b-413a-b751-6e0d84bfca6a%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/8ef1a8eb-3e4b-413a-b751-6e0d84bfca6a%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/8ef1a8eb-3e4b-413a-b751-6e0d84bfca6a%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQDctYOw\_3uqc0zgQAfaQpjTGUhYWPB03v0HYy1PsLT9\_w%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQDctYOw_3uqc0zgQAfaQpjTGUhYWPB03v0HYy1PsLT9_w%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Marc\_2](https://avatars.discourse-cdn.com/v4/letter/m/c4cdca/32.png) [@Marc\_2](https://discuss.elastic.co/u/Marc_2)\
**Post date:** [September 3, 2014, 9:33am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/11 "2014-09-03T09:33:29Z")

</div>

Hi Ivan,

I have resolved the problem. It works fine now. The template was wrong. The  
simpleQueryString works fine now too.

Cheers,  
Marc

On Tuesday, September 2, 2014 7:54:27 PM UTC+2, Ivan Brusic wrote:

> Hard to say without looking at your query, but perhaps you are  
> experiencing query parser issues. The query string query uses the standard  
> query parser, which might does not tokenize terms in the way your custom  
> tokenizer might. Try using match queries, which does not use the query  
> parser to see if it "fixes" the problem. Of course, you will have have the  
> query syntax at your disposal, but you can find workarounds.
> 
> --  
> Ivan
> 
> On Mon, Sep 1, 2014 at 4:17 AM, Marc \<[mn.o...@googlemail.com](mailto:mn.o...@googlemail.com) \<javascript:\>
> 
> > wrote:
> 
> Hi Ivan,
> 
> Using a test index and the analyze API, I was no able to create a config,  
> which is fine for me... theoretically.  
> {  
> "template": "logstash-\*",  
> "settings": {  
> "analysis": {  
> "filter": {
> 
> ```
> "my_word_delimiter": {
> 
> "type": "word_delimiter",
> "preserve_original": "true"
> }
> },
> 
> "analyzer": {
> 
> "my_analyzer": {
> "type": "custom",
> "tokenizer": "standard",
> "filter": ["standard",
> "lowercase",
> "stop",
> "my_word_delimiter",
> "asciifolding"]
> }
> }
> }
> },
> "mappings": {
> "_default_": {
> "properties": {
> 
> "excp": {
> 
> "type": "string",
> "index": "analyzed",
> "analyzer": "my_analyzer"
> },
> "msg": {
> 
> "type": "string",
> "index": "analyzed",
> 
> "analyzer": "my_analyzer"
> }
> }
> }
> }
> 
> ```
> 
> }  
> The problem now is, as soon as I activate this for the two fields and have  
> a new logstash index created I cannot use a simpleQueryString query to  
> retrieve any results.  
> It won't find anything via the REST api. Using the standard logstash  
> template and mapping it works fine.  
> Have you observed anything simililar?
> 
> Thx  
> Marc
> 
> On Friday, August 29, 2014 6:49:41 PM UTC+2, Ivan Brusic wrote:
> 
> That output does not look like the something generated from the standard  
> analyzer since it contains uppercase letters and various non-word  
> characters such as '='.
> 
> Your two analysis requests will differ since the second one contains the  
> default word\_delimiter filter instead of your custom my\_word\_delimiter.  
> What you are trying to achieve is somewhat difficult, but you can get there  
> if you keep on tweaking. 🙂 Try using a pattern tokenizer instead of the  
> whitespace tokenizer if you want more control over word boundaries.
> 
> --  
> Ivan
> 
> On Fri, Aug 29, 2014 at 1:48 AM, Marc [mn.o...@googlemail.com](mailto:mn.o...@googlemail.com) wrote:
> 
> Hi Ivan,
> 
> thanks again. I have tried so and found a reasonable combination.  
> Nevertheless, when I now try to use the analyze api with an index that has  
> the said analyzer defined via template it doesn't seem to apply:
> 
> This is the complete template:  
> {  
> "template": "bogstash-\*",  
> "settings": {  
> "index.number\_of\_replicas": 0,  
> "analysis": {  
> "analyzer": {  
> "msg\_excp\_analyzer": {  
> "type": "custom",  
> "tokenizer": "whitespace",  
> "filters": ["word\_delimiter",  
> "lowercase",  
> "asciifolding",  
> "shingle",  
> "standard"]  
> }  
> },  
> "filters": {  
> "my\_word\_delimiter": {  
> "type": "word\_delimiter",  
> "preserve\_original": "true"  
> },  
> "my\_asciifolding": {  
> "type": "asciifolding",  
> "preserve\_original": true  
> }  
> }  
> }  
> },  
> "mappings": {  
> "_default_": {  
> "properties": {  
> "@excp": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> },  
> "@msg": {  
> "type": "string",  
> "index": "analyzed",  
> "analyzer": "msg\_excp\_analyzer"  
> }  
> }  
> }  
> }  
> }  
> I create the index bogstash-1.  
> Now I test the following:  
> curl -XGET 'localhost:9200/bogstash-1/_analyze?analyzer=msg\_excp_  
> analyzer&pretty=1' -d 'Service=MyMDB.onMessage appId=cs  
> Times=Me:22/Total:22 (updated attributes=gps\_lng: 183731222/ gps\_lat:  
> 289309222/ )'  
> and it returns:  
> {  
> "tokens" : [ {  
> "token" : "Service=MyMDB.onMessage",  
> "start\_offset" : 0,  
> "end\_offset" : 23,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "appId=cs",  
> "start\_offset" : 24,  
> "end\_offset" : 32,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "Times=Me:22/Total:22",  
> "start\_offset" : 33,  
> "end\_offset" : 53,  
> "type" : "word",  
> "position" : 3  
> }, {  
> "token" : "(updated",  
> "start\_offset" : 54,  
> "end\_offset" : 62,  
> "type" : "word",  
> "position" : 4  
> }, {  
> "token" : "attributes=gps\_lng:",  
> "start\_offset" : 63,  
> "end\_offset" : 82,  
> "type" : "word",  
> "position" : 5  
> }, {  
> "token" : "183731222/",  
> "start\_offset" : 83,  
> "end\_offset" : 93,  
> "type" : "word",  
> "position" : 6  
> }, {  
> "token" : "gps\_lat:",  
> "start\_offset" : 94,  
> "end\_offset" : 102,  
> "type" : "word",  
> "position" : 7  
> }, {  
> "token" : "289309222/",  
> "start\_offset" : 103,  
> "end\_offset" : 113,  
> "type" : "word",  
> "position" : 8  
> }, {  
> "token" : ")",  
> "start\_offset" : 114,  
> "end\_offset" : 115,  
> "type" : "word",  
> "position" : 9  
> } ]  
> }  
> Which is the output of a standard analyzer.  
> Giving the tokenizer and filters in the analyze API directly works fine:  
> curl -XGET 'localhost:9200/\_analyze?tokenizer=whitespace&filters=  
> lowercase,word\_delimiter,shingle,asciifolding,standard&pretty=1' -d 'Service=MyMDB.onMessage  
> appId=cs Times=Me:22/Total:22 (updated attributes=gps\_lng: 183731222/  
> gps\_lat: 289309222/ )'  
> This results in:  
> {  
> "tokens" : [ {  
> "token" : "service",  
> "start\_offset" : 0,  
> "end\_offset" : 7,  
> "type" : "word",  
> "position" : 1  
> }, {  
> "token" : "service mymdb",  
> "start\_offset" : 0,  
> "end\_offset" : 13,  
> "type" : "shingle",  
> "position" : 1  
> }, {  
> "token" : "mymdb",  
> "start\_offset" : 8,  
> "end\_offset" : 13,  
> "type" : "word",  
> "position" : 2  
> }, {  
> "token" : "mymdb onmessage",  
> "start\_offset" : 8,  
> "end\_offset" : 23,  
> "type" : "shingle",  
> "position" : 2  
> }, {  
> "token" : "onmessage",  
> "start\_offset" : \<span style="colo
> 
> ...

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/8b242461-9899-4de7-8ad3-da6645be9947%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8b242461-9899-4de7-8ad3-da6645be9947%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:04am UTC](https://discuss.elastic.co/t/el-setup-for-fulltext-search/19483/12 "2017-07-06T01:04:42Z")

</div>


