# Word delimiter

**URL:** <https://discuss.elastic.co/t/word-delimiter/20315>\
**Category:** Elasticsearch\
**Created:** [October 17, 2014, 11:57pm UTC](https://discuss.elastic.co/t/word-delimiter/20315 "2014-10-17T23:57:52Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![nicktackes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicktackes/32/1174_2.png) [@nicktackes](https://discuss.elastic.co/u/nicktackes)\
**Post date:** [October 17, 2014, 11:57pm UTC](https://discuss.elastic.co/t/word-delimiter/20315/1 "2014-10-17T23:57:52Z")

</div>

Hello, I am experimenting with word\_delimiter and have an example with a  
special character that is indexed. The character is in the type table for  
the word delimiter. analysis of the tokenization looks good, but when i  
attempt to do a match query it doesnt seem to respect tokenization as  
expected.  
The example indexes 'HER2+ Breast Cancer'. Tokenization is 'her2+',  
'breast', 'cancer', which is good. searching for 'HER2\+' results in a  
hit, as well as 'HER2\-'

#!/bin/sh  
curl -XPUT '[http://localhost:9200/specialchars](http://localhost:9200/specialchars)' -d '{  
"settings" : {  
"index" : {  
"number\_of\_shards" : 1,  
"number\_of\_replicas" : 1  
},  
"analysis" : {  
"filter" : {  
"special\_character\_spliter" : {  
"type" : "word\_delimiter",  
"split\_on\_numerics":false,  
"type\_table": ["+ =\> ALPHA", "- =\> ALPHA"]  
}  
},  
"analyzer" : {  
"schar\_analyzer" : {  
"type" : "custom",  
"tokenizer" : "whitespace",  
"filter" : ["lowercase", "special\_character\_spliter"]  
}  
}  
}  
},  
"mappings" : {  
"specialchars" : {  
"properties" : {  
"msg" : {  
"type" : "string",  
"analyzer" : "schar\_analyzer"  
}  
}  
}  
}  
}'

curl -XPOST localhost:9200/specialchars/1 -d '{"msg" : "HER2+ Breast  
Cancer"}'  
curl -XPOST localhost:9200/specialchars/2 -d '{"msg" : "Non-Small Cell Lung  
Cancer"}'  
curl -XPOST localhost:9200/specialchars/3 -d '{"msg" : "c.2573T\>G NSCLC"}'

curl -XPOST localhost:9200/specialchars/\_refresh

curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
"HER2+ Breast Cancer"  
#curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
"Non-Small Cell Lung Cancer"  
#curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
"c.2573T\>G NSCLC"

printf "HER2+\n"  
curl -XGET localhost:9200/specialchars/\_search?pretty -d '{  
"query" : {  
"match" : {  
"msg" : {  
"query" : "HER2\+"  
}  
}  
}  
}'

printf "HER2-\n"  
curl -XGET localhost:9200/specialchars/\_search?pretty -d '{  
"query" : {  
"match" : {  
"msg" : {  
"query" : "HER2\-"  
}  
}  
}  
}'

curl -X DELETE localhost:9200/specialchars

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/becb02b7-72f0-42dd-b347-5f031fa154d3%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/becb02b7-72f0-42dd-b347-5f031fa154d3%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nicktackes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicktackes/32/1174_2.png) [@nicktackes](https://discuss.elastic.co/u/nicktackes)\
**Post date:** [October 20, 2014, 4:49pm UTC](https://discuss.elastic.co/t/word-delimiter/20315/2 "2014-10-20T16:49:15Z")

</div>

any thoughts on how I am constructing my search query? I have tried  
escaping the special characters as well as passing the unescaped special  
chars. I was hoping to stick with a match query although i tried query  
string query, and match phrase and term query and had found no solution  
there.

very appreciative of your thoughts.

Nick

On Friday, October 17, 2014 4:57:52 PM UTC-7, Nick Tackes wrote:

> Hello, I am experimenting with word\_delimiter and have an example with a  
> special character that is indexed. The character is in the type table for  
> the word delimiter. analysis of the tokenization looks good, but when i  
> attempt to do a match query it doesnt seem to respect tokenization as  
> expected.  
> The example indexes 'HER2+ Breast Cancer'. Tokenization is 'her2+',  
> 'breast', 'cancer', which is good. searching for 'HER2\+' results in a  
> hit, as well as 'HER2\-'
> 
> #!/bin/sh  
> curl -XPUT '[http://localhost:9200/specialchars](http://localhost:9200/specialchars)' -d '{  
> "settings" : {  
> "index" : {  
> "number\_of\_shards" : 1,  
> "number\_of\_replicas" : 1  
> },  
> "analysis" : {  
> "filter" : {  
> "special\_character\_spliter" : {  
> "type" : "word\_delimiter",  
> "split\_on\_numerics":false,  
> "type\_table": ["+ =\> ALPHA", "- =\> ALPHA"]  
> }  
> },  
> "analyzer" : {  
> "schar\_analyzer" : {  
> "type" : "custom",  
> "tokenizer" : "whitespace",  
> "filter" : ["lowercase", "special\_character\_spliter"]  
> }  
> }  
> }  
> },  
> "mappings" : {  
> "specialchars" : {  
> "properties" : {  
> "msg" : {  
> "type" : "string",  
> "analyzer" : "schar\_analyzer"  
> }  
> }  
> }  
> }  
> }'
> 
> curl -XPOST localhost:9200/specialchars/1 -d '{"msg" : "HER2+ Breast  
> Cancer"}'  
> curl -XPOST localhost:9200/specialchars/2 -d '{"msg" : "Non-Small Cell  
> Lung Cancer"}'  
> curl -XPOST localhost:9200/specialchars/3 -d '{"msg" : "c.2573T\>G NSCLC"}'
> 
> curl -XPOST localhost:9200/specialchars/\_refresh
> 
> curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
> "HER2+ Breast Cancer"  
> #curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
> "Non-Small Cell Lung Cancer"  
> #curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
> "c.2573T\>G NSCLC"
> 
> printf "HER2+\n"  
> curl -XGET localhost:9200/specialchars/\_search?pretty -d '{  
> "query" : {  
> "match" : {  
> "msg" : {  
> "query" : "HER2\+"  
> }  
> }  
> }  
> }'
> 
> printf "HER2-\n"  
> curl -XGET localhost:9200/specialchars/\_search?pretty -d '{  
> "query" : {  
> "match" : {  
> "msg" : {  
> "query" : "HER2\-"  
> }  
> }  
> }  
> }'
> 
> curl -X DELETE localhost:9200/specialchars

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/62d1f188-92e1-4ea4-b94b-b47c696db78f%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/62d1f188-92e1-4ea4-b94b-b47c696db78f%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nicktackes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nicktackes/32/1174_2.png) [@nicktackes](https://discuss.elastic.co/u/nicktackes)\
**Post date:** [October 20, 2014, 11:44pm UTC](https://discuss.elastic.co/t/word-delimiter/20315/3 "2014-10-20T23:44:18Z")

</div>

I resolved my own issue.

#!/bin/sh  
curl -XPUT '[http://localhost:9200/specialchars](http://localhost:9200/specialchars)' -d '{  
"settings" : {  
"index" : {  
"number\_of\_shards" : 1,  
"number\_of\_replicas" : 1  
},  
"analysis" : {  
"filter" : {  
"special\_character\_splitter" : {  
"type" : "word\_delimiter",  
"split\_on\_numerics":false,  
"type\_table": ["+ =\> ALPHANUM", "- =\> ALPHANUM", "@ =\>  
ALPHANUM"]  
}  
},  
"analyzer" : {  
"schar\_analyzer" : {  
"type" : "custom",  
"tokenizer" : "whitespace",  
"filter" : ["lowercase", "special\_character\_splitter"]  
}  
}  
}  
},  
"mappings" : {  
"specialchars" : {  
"properties" : {  
"msg" : {  
"type" : "string",  
"analyzer" : "schar\_analyzer"  
}  
}  
}  
}  
}'

curl -XPOST localhost:9200/specialchars/specialchars/1 -d '{"msg" : "HER2+  
Breast Cancer"}'  
curl -XPOST localhost:9200/specialchars/specialchars/2 -d '{"msg" :  
"Non-Small Cell Lung Cancer"}'  
curl -XPOST localhost:9200/specialchars/specialchars/3 -d '{"msg" :  
"c.2573T\>G NSCLC"}'

curl -XPOST localhost:9200/specialchars/\_refresh

curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
"HER2+ Breast Cancer"  
#curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
"Non-Small Cell Lung Cancer"  
#curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
"c.2573T\>G NSCLC"

printf "HER2+\n"  
curl -XGET localhost:9200/specialchars/specialchars/\_search?pretty -d '{  
"query" : {  
"match" : {  
"msg" : {  
"query" : "HER2+"  
}  
}  
}  
}'

printf "HER2-\n"  
curl -XGET localhost:9200/specialchars/specialchars/\_search?pretty -d '{  
"query" : {  
"match" : {  
"msg" : {  
"query" : "HER2-"  
}  
}  
}  
}'

printf "HER2@\n"  
curl -XGET localhost:9200/specialchars/specialchars/\_search?pretty -d '{  
"query" : {  
"match" : {  
"msg" : {  
"query" : "HER2@"  
}  
}  
}  
}'

curl -X DELETE localhost:9200/specialchars

On Friday, October 17, 2014 4:57:52 PM UTC-7, Nick Tackes wrote:

> Hello, I am experimenting with word\_delimiter and have an example with a  
> special character that is indexed. The character is in the type table for  
> the word delimiter. analysis of the tokenization looks good, but when i  
> attempt to do a match query it doesnt seem to respect tokenization as  
> expected.  
> The example indexes 'HER2+ Breast Cancer'. Tokenization is 'her2+',  
> 'breast', 'cancer', which is good. searching for 'HER2\+' results in a  
> hit, as well as 'HER2\-'
> 
> #!/bin/sh  
> curl -XPUT '[http://localhost:9200/specialchars](http://localhost:9200/specialchars)' -d '{  
> "settings" : {  
> "index" : {  
> "number\_of\_shards" : 1,  
> "number\_of\_replicas" : 1  
> },  
> "analysis" : {  
> "filter" : {  
> "special\_character\_spliter" : {  
> "type" : "word\_delimiter",  
> "split\_on\_numerics":false,  
> "type\_table": ["+ =\> ALPHA", "- =\> ALPHA"]  
> }  
> },  
> "analyzer" : {  
> "schar\_analyzer" : {  
> "type" : "custom",  
> "tokenizer" : "whitespace",  
> "filter" : ["lowercase", "special\_character\_spliter"]  
> }  
> }  
> }  
> },  
> "mappings" : {  
> "specialchars" : {  
> "properties" : {  
> "msg" : {  
> "type" : "string",  
> "analyzer" : "schar\_analyzer"  
> }  
> }  
> }  
> }  
> }'
> 
> curl -XPOST localhost:9200/specialchars/1 -d '{"msg" : "HER2+ Breast  
> Cancer"}'  
> curl -XPOST localhost:9200/specialchars/2 -d '{"msg" : "Non-Small Cell  
> Lung Cancer"}'  
> curl -XPOST localhost:9200/specialchars/3 -d '{"msg" : "c.2573T\>G NSCLC"}'
> 
> curl -XPOST localhost:9200/specialchars/\_refresh
> 
> curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
> "HER2+ Breast Cancer"  
> #curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
> "Non-Small Cell Lung Cancer"  
> #curl -XGET 'localhost:9200/specialchars/\_analyze?field=msg&pretty=1' -d  
> "c.2573T\>G NSCLC"
> 
> printf "HER2+\n"  
> curl -XGET localhost:9200/specialchars/\_search?pretty -d '{  
> "query" : {  
> "match" : {  
> "msg" : {  
> "query" : "HER2\+"  
> }  
> }  
> }  
> }'
> 
> printf "HER2-\n"  
> curl -XGET localhost:9200/specialchars/\_search?pretty -d '{  
> "query" : {  
> "match" : {  
> "msg" : {  
> "query" : "HER2\-"  
> }  
> }  
> }  
> }'
> 
> curl -X DELETE localhost:9200/specialchars

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ad8ebeac-a75d-461d-920d-cba1a25a3226%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ad8ebeac-a75d-461d-920d-cba1a25a3226%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:54am UTC](https://discuss.elastic.co/t/word-delimiter/20315/4 "2017-07-06T00:54:52Z")

</div>


