# Issue with using word delimiter

**URL:** <https://discuss.elastic.co/t/issue-with-using-word-delimiter/17142>\
**Category:** Elasticsearch\
**Created:** [April 23, 2014, 2:22am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter/17142 "2014-04-23T02:22:30Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Amit\_Soni](https://avatars.discourse-cdn.com/v4/letter/a/c0e974/32.png) [@Amit\_Soni](https://discuss.elastic.co/u/Amit_Soni)\
**Post date:** [April 23, 2014, 2:22am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter/17142/1 "2014-04-23T02:22:30Z")

</div>

Hi team - I just wanted to share complete config file wherein I am able to  
see this problem with word delimiter (unless I got the config wrong). My  
config is below and if I analyze the string "650-454-2343", I get the  
following tokens:

1. 650-454-2343 [expected since we have "preserve\_original": true]
2. 650 [unexpected]
3. 454 [unexpected]
4. 2343 [unexpected]
5. 6504542343 [expected since we have "catenate\_all": true]

thoughts?

{  
"settings": {  
"number\_of\_shards": 5,  
"number\_of\_replicas": 0,  
"analysis": {  
"analyzer": {  
"phoneAnalyzer": {  
"type": "custom",  
"tokenizer": "whitespace",  
"filter": [  
"word\_delimiter\_for\_phone"  
]  
}  
},  
"filter": {  
"word\_delimiter\_for\_phone": {  
"type": "word\_delimiter",  
"catenate\_all": true,  
"generate\_number\_parts ": false,  
"split\_on\_case\_change": false,  
"generate\_word\_parts": false,  
"split\_on\_numerics": false,  
"preserve\_original": true  
}  
}  
}  
},  
"mappings": {  
"my\_type": {  
"properties": {  
"phone": {  
"type": "string",  
"index\_analyzer": "phoneAnalyzer",  
"include\_in\_all": false  
}  
}  
}  
}  
}

-Amit.

On Mon, Apr 21, 2014 at 8:46 PM, Amit Soni [amitsoni29@gmail.com](mailto:amitsoni29@gmail.com) wrote:

> hi everyone - I have changed the mapping so that it now looks like below.  
> However for a given input say 123-456-8989, the generated tokens are:
> 
> a) 123-456-8989 b) 123 c) 456 d) 8989 e) 1234568989
> 
> I was expecting just two tokens: a) 123-456-8989 b) 1234568989
> 
> Would you know what might be going wrong here?
> 
> "default\_index": {  
> "tokenizer": "keyword",  
> "filter": [  
> "lowercase"  
> ]
> 
> },
> 
> "phoneAnalyzer": {  
> "type": "custom",  
> "tokenizer": "keyword",  
> "filter": [  
> "word\_delimiter\_for\_phone"  
> ]  
> },
> 
> "word\_delimiter\_for\_phone": {  
> "type": "word\_delimiter",  
> "catenate\_all": true,  
> "generate\_number\_parts ": false,  
> "split\_on\_case\_change": false,  
> "generate\_word\_parts": false,  
> "split\_on\_numerics": false,  
> "preserve\_original": true  
> },
> 
> -Amit.
> 
> On Fri, Nov 1, 2013 at 1:07 AM, David Pilato [david@pilato.fr](mailto:david@pilato.fr) wrote:
> 
> > Sorry. Forget my answer. Useless here.
> > 
> > --  
> > David 😉  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > 
> > Le 1 nov. 2013 à 08:05, David Pilato [david@pilato.fr](mailto:david@pilato.fr) a écrit :
> > 
> > Or disable analysis for this field.
> > 
> > HTH
> > 
> > --  
> > David 😉  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > 
> > Le 1 nov. 2013 à 07:42, [sina.tamanna@gmail.com](mailto:sina.tamanna@gmail.com) a écrit :
> > 
> > Analysis starts by using tokenizer, which in your case is "standard".  
> > Therefore the input "345 678-1234" will be tokenized to "345", "678", and  
> > "1234", and only then the filters will be applied. A solution to get the  
> > original and the concatenated input would be to use the "keyword" tokenizer.
> > 
> > On Thursday, October 31, 2013 8:10:55 PM UTC+1, amit.soni wrote:
> > 
> > > Hi all - I have a phone number field and I am trying to use  
> > > word\_delimiter filter in order break it up into tokens, preserve the  
> > > original entry and concatenate all the numbers in the entry. I have the  
> > > following entry:
> > > 
> > > "phoneAnalyzer" : {  
> > > "type": "custom",  
> > > "tokenizer": "standard",  
> > > "filter": [  
> > > "word\_delimiter\_for\_phone"  
> > > ]  
> > > }
> > > 
> > > "filter": {  
> > > "word\_delimiter\_for\_phone": {  
> > > "type": "word\_delimiter",
> > > 
> > > - 
> > > 
> > > ```
> > > "catenate_numbers" : true,*
> > > "preserve_original" : true
> > > },
> > > 
> > > ```
> > > 
> > > }
> > > 
> > > Using this, when I run it on input "345 678-1234" I get the following:
> > > 
> > > {  
> > > "tokens" : [ {  
> > > "token" : "_345_",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 3,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "_678_",  
> > > "start\_offset" : 4,  
> > > "end\_offset" : 7,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "_1234_",  
> > > "start\_offset" : 8,  
> > > "end\_offset" : 12,  
> > > "type" : "",  
> > > "position" : 3  
> > > } ]  
> > > }
> > > 
> > > Question: Should this also not have generated a concatenated string of  
> > > the form: 3456781234.
> > > 
> > > Anything I am missing here?
> > > 
> > > -Amit.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAAOGaQKGSvbjGHFvRKtqhW0zPitD-DK%2B2%3DMVBrjS4THJUN4Duw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAAOGaQKGSvbjGHFvRKtqhW0zPitD-DK%2B2%3DMVBrjS4THJUN4Duw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:34am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter/17142/2 "2017-07-06T01:34:10Z")

</div>


