# Do entries in a synonym list always get whitespace tokenized?

**URL:** <https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782>\
**Category:** Elasticsearch\
**Created:** [September 27, 2013, 1:14am UTC](https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782 "2013-09-27T01:14:24Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Glen\_Smith](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/glen_smith/32/111656_2.png) [@Glen\_Smith](https://discuss.elastic.co/u/Glen_Smith)\
**Post date:** [September 27, 2013, 1:14am UTC](https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782/1 "2013-09-27T01:14:24Z")

</div>

It appears to me they do.

(Apologies if this is a repost. Posted over an hour ago and it hasn't shown  
up here.)

#!/bin/sh  
echo "\nattempt to delete the index"  
curl -XDELETE "[http://localhost:9200/syndex/?pretty=false](http://localhost:9200/syndex/?pretty=false)"  
echo "\ncreate the index"  
curl -XPUT "[http://localhost:9200/syndex/?pretty=true](http://localhost:9200/syndex/?pretty=true)" -d '{  
"settings": {  
"number\_of\_shards": 1,  
"number\_of\_replicas": 0,  
"analysis": {  
"analyzer": {  
"syn": {  
"tokenizer": "keyword",  
"filter": ["color\_synonym"]  
}  
},  
"filter": {  
"color\_synonym": {  
"type" : "synonym",  
"synonyms" : ["red, another shade"]  
}  
}  
}  
}  
}'  
echo "\n analyze red: I want a single token _another shade_"  
echo "\n instead I get two tokens _another_ & _shade_"  
curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "red"  
echo "\n sanity check how keyword tokenizer handes another shade"  
curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "another shade"  
echo "\n and of course it does not split them"

_Generates the output_

attempt to delete the index  
{"ok":true,"acknowledged":true}  
create the index  
{  
"ok" : true,  
"acknowledged" : true  
}  
analyze red: I want a single token _another shade_

instead I get two tokens _another_ & _shade_  
{  
"tokens" : [ {  
"token" : "red",  
"start\_offset" : 0,  
"end\_offset" : 3,  
"type" : "SYNONYM",  
"position" : 1  
}, {  
"token" : "another",  
"start\_offset" : 0,  
"end\_offset" : 3,  
"type" : "SYNONYM",  
"position" : 1  
}, {  
"token" : "shade",  
"start\_offset" : 0,  
"end\_offset" : 3,  
"type" : "SYNONYM",  
"position" : 2  
} ]  
}  
sanity check how keyword tokenizer handes another shade  
{  
"tokens" : [ {  
"token" : "another shade",  
"start\_offset" : 0,  
"end\_offset" : 13,  
"type" : "word",  
"position" : 1  
} ]  
}  
and of course it does not split them

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 27, 2013, 4:22pm UTC](https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782/2 "2013-09-27T16:22:29Z")

</div>

Correct, the terms are being tokenized with a WhitespaceTokenizer. You can  
see the code in SynonymTokenFilterFactory:

[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85)

Cheers,

Ivan

On Thu, Sep 26, 2013 at 6:14 PM, Glen Smith [glen@smithsrock.com](mailto:glen@smithsrock.com) wrote:

> It appears to me they do.
> 
> (Apologies if this is a repost. Posted over an hour ago and it hasn't  
> shown up here.)
> 
> #!/bin/sh  
> echo "\nattempt to delete the index"  
> curl -XDELETE "[http://localhost:9200/syndex/?pretty=false](http://localhost:9200/syndex/?pretty=false)"  
> echo "\ncreate the index"  
> curl -XPUT "[http://localhost:9200/syndex/?pretty=true](http://localhost:9200/syndex/?pretty=true)" -d '{  
> "settings": {  
> "number\_of\_shards": 1,  
> "number\_of\_replicas": 0,  
> "analysis": {  
> "analyzer": {  
> "syn": {  
> "tokenizer": "keyword",  
> "filter": ["color\_synonym"]  
> }  
> },  
> "filter": {  
> "color\_synonym": {  
> "type" : "synonym",  
> "synonyms" : ["red, another shade"]  
> }  
> }  
> }  
> }  
> }'  
> echo "\n analyze red: I want a single token _another shade_"  
> echo "\n instead I get two tokens _another_ & _shade_"  
> curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "red"  
> echo "\n sanity check how keyword tokenizer handes another shade"  
> curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "another shade"  
> echo "\n and of course it does not split them"
> 
> _Generates the output_
> 
> attempt to delete the index  
> {"ok":true,"acknowledged":true}  
> create the index  
> {  
> "ok" : true,  
> "acknowledged" : true  
> }  
> analyze red: I want a single token _another shade_
> 
> instead I get two tokens _another_ & _shade_  
> {  
> "tokens" : [ {  
> "token" : "red",  
> "start\_offset" : 0,  
> "end\_offset" : 3,  
> "type" : "SYNONYM",  
> "position" : 1  
> }, {  
> "token" : "another",  
> "start\_offset" : 0,  
> "end\_offset" : 3,  
> "type" : "SYNONYM",  
> "position" : 1  
> }, {  
> "token" : "shade",  
> "start\_offset" : 0,  
> "end\_offset" : 3,  
> "type" : "SYNONYM",  
> "position" : 2  
> } ]  
> }  
> sanity check how keyword tokenizer handes another shade  
> {  
> "tokens" : [ {  
> "token" : "another shade",  
> "start\_offset" : 0,  
> "end\_offset" : 13,  
> "type" : "word",  
> "position" : 1  
> } ]  
> }  
> and of course it does not split them
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 27, 2013, 4:23pm UTC](https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782/3 "2013-09-27T16:23:34Z")

</div>

Hit reply too soon. You can change the tokenizer used by passing setting  
the tokenizer setting.

--  
Ivan

On Fri, Sep 27, 2013 at 9:22 AM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> Correct, the terms are being tokenized with a WhitespaceTokenizer. You can  
> see the code in SynonymTokenFilterFactory:
> 
> [https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85)
> 
> Cheers,
> 
> Ivan
> 
> On Thu, Sep 26, 2013 at 6:14 PM, Glen Smith [glen@smithsrock.com](mailto:glen@smithsrock.com) wrote:
> 
> > It appears to me they do.
> > 
> > (Apologies if this is a repost. Posted over an hour ago and it hasn't  
> > shown up here.)
> > 
> > #!/bin/sh  
> > echo "\nattempt to delete the index"  
> > curl -XDELETE "[http://localhost:9200/syndex/?pretty=false](http://localhost:9200/syndex/?pretty=false)"  
> > echo "\ncreate the index"  
> > curl -XPUT "[http://localhost:9200/syndex/?pretty=true](http://localhost:9200/syndex/?pretty=true)" -d '{  
> > "settings": {  
> > "number\_of\_shards": 1,  
> > "number\_of\_replicas": 0,  
> > "analysis": {  
> > "analyzer": {  
> > "syn": {  
> > "tokenizer": "keyword",  
> > "filter": ["color\_synonym"]  
> > }  
> > },  
> > "filter": {  
> > "color\_synonym": {  
> > "type" : "synonym",  
> > "synonyms" : ["red, another shade"]  
> > }  
> > }  
> > }  
> > }  
> > }'  
> > echo "\n analyze red: I want a single token _another shade_"  
> > echo "\n instead I get two tokens _another_ & _shade_"  
> > curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "red"  
> > echo "\n sanity check how keyword tokenizer handes another shade"  
> > curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "another shade"  
> > echo "\n and of course it does not split them"
> > 
> > _Generates the output_
> > 
> > attempt to delete the index  
> > {"ok":true,"acknowledged":true}  
> > create the index  
> > {  
> > "ok" : true,  
> > "acknowledged" : true  
> > }  
> > analyze red: I want a single token _another shade_
> > 
> > instead I get two tokens _another_ & _shade_  
> > {  
> > "tokens" : [ {  
> > "token" : "red",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "SYNONYM",  
> > "position" : 1  
> > }, {  
> > "token" : "another",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "SYNONYM",  
> > "position" : 1  
> > }, {  
> > "token" : "shade",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "SYNONYM",  
> > "position" : 2  
> > } ]  
> > }  
> > sanity check how keyword tokenizer handes another shade  
> > {  
> > "tokens" : [ {  
> > "token" : "another shade",  
> > "start\_offset" : 0,  
> > "end\_offset" : 13,  
> > "type" : "word",  
> > "position" : 1  
> > } ]  
> > }  
> > and of course it does not split them
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Glen\_Smith](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/glen_smith/32/111656_2.png) [@Glen\_Smith](https://discuss.elastic.co/u/Glen_Smith)\
**Post date:** [September 27, 2013, 5:00pm UTC](https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782/4 "2013-09-27T17:00:56Z")

</div>

Thanks, Ivan. Yeah, I finally did figure out the obvious - the filter gets  
its _own_ tokenizer.

Hopefully this reply doesn't take 15 hours to show up...

On Friday, September 27, 2013 12:22:29 PM UTC-4, Ivan Brusic wrote:

> Correct, the terms are being tokenized with a WhitespaceTokenizer. You can  
> see the code in SynonymTokenFilterFactory:
> 
> [https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85)
> 
> Cheers,
> 
> Ivan
> 
> On Thu, Sep 26, 2013 at 6:14 PM, Glen Smith \<[gl...@smithsrock.com](mailto:gl...@smithsrock.com)\<javascript:\>
> 
> > wrote:
> 
> > It appears to me they do.
> > 
> > (Apologies if this is a repost. Posted over an hour ago and it hasn't  
> > shown up here.)
> > 
> > #!/bin/sh  
> > echo "\nattempt to delete the index"  
> > curl -XDELETE "[http://localhost:9200/syndex/?pretty=false](http://localhost:9200/syndex/?pretty=false)"  
> > echo "\ncreate the index"  
> > curl -XPUT "[http://localhost:9200/syndex/?pretty=true](http://localhost:9200/syndex/?pretty=true)" -d '{  
> > "settings": {  
> > "number\_of\_shards": 1,  
> > "number\_of\_replicas": 0,  
> > "analysis": {  
> > "analyzer": {  
> > "syn": {  
> > "tokenizer": "keyword",  
> > "filter": ["color\_synonym"]  
> > }  
> > },  
> > "filter": {  
> > "color\_synonym": {  
> > "type" : "synonym",  
> > "synonyms" : ["red, another shade"]  
> > }  
> > }  
> > }  
> > }  
> > }'  
> > echo "\n analyze red: I want a single token _another shade_"  
> > echo "\n instead I get two tokens _another_ & _shade_"  
> > curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "red"  
> > echo "\n sanity check how keyword tokenizer handes another shade"  
> > curl -XGET "localhost:9200/syndex/\_analyze?analyzer=syn&pretty=true" -d "another shade"  
> > echo "\n and of course it does not split them"
> > 
> > _Generates the output_
> > 
> > attempt to delete the index  
> > {"ok":true,"acknowledged":true}  
> > create the index  
> > {  
> > "ok" : true,  
> > "acknowledged" : true  
> > }  
> > analyze red: I want a single token _another shade_
> > 
> > instead I get two tokens _another_ & _shade_  
> > {  
> > "tokens" : [ {  
> > "token" : "red",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "SYNONYM",  
> > "position" : 1  
> > }, {  
> > "token" : "another",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "SYNONYM",  
> > "position" : 1  
> > }, {  
> > "token" : "shade",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "SYNONYM",  
> > "position" : 2  
> > } ]  
> > }  
> > sanity check how keyword tokenizer handes another shade  
> > {  
> > "tokens" : [ {  
> > "token" : "another shade",  
> > "start\_offset" : 0,  
> > "end\_offset" : 13,  
> > "type" : "word",  
> > "position" : 1  
> > } ]  
> > }  
> > and of course it does not split them
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 27, 2013, 6:13pm UTC](https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782/5 "2013-09-27T18:13:23Z")

</div>

Nope, it showed up right away. 🙂 BTW, I love your avatar.

--  
Ivan

On Fri, Sep 27, 2013 at 10:00 AM, Glen Smith [glen@smithsrock.com](mailto:glen@smithsrock.com) wrote:

> Thanks, Ivan. Yeah, I finally did figure out the obvious - the filter gets  
> its _own_ tokenizer.
> 
> Hopefully this reply doesn't take 15 hours to show up...
> 
> On Friday, September 27, 2013 12:22:29 PM UTC-4, Ivan Brusic wrote:
> 
> > Correct, the terms are being tokenized with a WhitespaceTokenizer. You  
> > can see the code in SynonymTokenFilterFactory:
> > 
> > [https://github.com/\*\*elasticsearch/elasticsearch/](https://github.com/**elasticsearch/elasticsearch/)\*\*  
> > blob/master/src/main/java/org/ **elasticsearch/index/analysis/**  
> > SynonymTokenFilterFactory.\*\*java#L85[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/SynonymTokenFilterFactory.java#L85)
> > 
> > Cheers,
> > 
> > Ivan
> > 
> > On Thu, Sep 26, 2013 at 6:14 PM, Glen Smith [gl...@smithsrock.com](mailto:gl...@smithsrock.com) wrote:
> > 
> > > It appears to me they do.
> > > 
> > > (Apologies if this is a repost. Posted over an hour ago and it hasn't  
> > > shown up here.)
> > > 
> > > #!/bin/sh  
> > > echo "\nattempt to delete the index"  
> > > curl -XDELETE "[http://localhost:9200/syndex/\*\*?pretty=false](http://localhost:9200/syndex/**?pretty=false) [http://localhost:9200/syndex/?pretty=false](http://localhost:9200/syndex/?pretty=false)"  
> > > echo "\ncreate the index"  
> > > curl -XPUT "[http://localhost:9200/syndex/\*\*?pretty=true](http://localhost:9200/syndex/**?pretty=true) [http://localhost:9200/syndex/?pretty=true](http://localhost:9200/syndex/?pretty=true)" -d '{  
> > > "settings": {  
> > > "number\_of\_shards": 1,  
> > > "number\_of\_replicas": 0,  
> > > "analysis": {  
> > > "analyzer": {  
> > > "syn": {  
> > > "tokenizer": "keyword",  
> > > "filter": ["color\_synonym"]  
> > > }  
> > > },  
> > > "filter": {  
> > > "color\_synonym": {  
> > > "type" : "synonym",  
> > > "synonyms" : ["red, another shade"]  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }'  
> > > echo "\n analyze red: I want a single token _another shade_"  
> > > echo "\n instead I get two tokens _another_ & _shade_"  
> > > curl -XGET "localhost:9200/syndex/_\*\*analyze?analyzer=syn&pretty=\*\*true" -d "red"  
> > > echo "\n sanity check how keyword tokenizer handes another shade"  
> > > curl -XGET "localhost:9200/syndex/_\*\*analyze?analyzer=syn&pretty=\*\*true" -d "another shade"  
> > > echo "\n and of course it does not split them"
> > > 
> > > _Generates the output_
> > > 
> > > attempt to delete the index  
> > > {"ok":true,"acknowledged":\*\*true}  
> > > create the index  
> > > {  
> > > "ok" : true,  
> > > "acknowledged" : true  
> > > }  
> > > analyze red: I want a single token _another shade_
> > > 
> > > instead I get two tokens _another_ & _shade_  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "red",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 3,  
> > > "type" : "SYNONYM",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "another",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 3,  
> > > "type" : "SYNONYM",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "shade",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 3,  
> > > "type" : "SYNONYM",  
> > > "position" : 2  
> > > } ]  
> > > }  
> > > sanity check how keyword tokenizer handes another shade  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "another shade",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 13,  
> > > "type" : "word",  
> > > "position" : 1  
> > > } ]  
> > > }  
> > > and of course it does not split them
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > 
> > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > .
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:14am UTC](https://discuss.elastic.co/t/do-entries-in-a-synonym-list-always-get-whitespace-tokenized/13782/6 "2017-07-06T02:14:23Z")

</div>


