# Issue with using word delimiter filter

**URL:** <https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201>\
**Category:** Elasticsearch\
**Created:** [October 31, 2013, 7:10pm UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201 "2013-10-31T19:10:55Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Amit\_Soni](https://avatars.discourse-cdn.com/v4/letter/a/c0e974/32.png) [@Amit\_Soni](https://discuss.elastic.co/u/Amit_Soni)\
**Post date:** [October 31, 2013, 7:10pm UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201/1 "2013-10-31T19:10:55Z")

</div>

Hi all - I have a phone number field and I am trying to use word\_delimiter  
filter in order break it up into tokens, preserve the original entry and  
concatenate all the numbers in the entry. I have the following entry:

"phoneAnalyzer" : {  
"type": "custom",  
"tokenizer": "standard",  
"filter": [  
"word\_delimiter\_for\_phone"  
]  
}

"filter": {  
"word\_delimiter\_for\_phone": {  
"type": "word\_delimiter",

- 

```
                "catenate_numbers" : true,*
               "preserve_original" : true
          },

```

}

Using this, when I run it on input "345 678-1234" I get the following:

{  
"tokens" : [ {  
"token" : "_345_",  
"start\_offset" : 0,  
"end\_offset" : 3,  
"type" : "",  
"position" : 1  
}, {  
"token" : "_678_",  
"start\_offset" : 4,  
"end\_offset" : 7,  
"type" : "",  
"position" : 2  
}, {  
"token" : "_1234_",  
"start\_offset" : 8,  
"end\_offset" : 12,  
"type" : "",  
"position" : 3  
} ]  
}

Question: Should this also not have generated a concatenated string of the  
form: 3456781234.

Anything I am missing here?

-Amit.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![sina\_tamanna](https://avatars.discourse-cdn.com/v4/letter/s/a6a055/32.png) [@sina\_tamanna](https://discuss.elastic.co/u/sina_tamanna)\
**Post date:** [November 1, 2013, 7:42am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201/2 "2013-11-01T07:42:51Z")

</div>

Analysis starts by using tokenizer, which in your case is "standard".  
Therefore the input "345 678-1234" will be tokenized to "345", "678", and  
"1234", and only then the filters will be applied. A solution to get the  
original and the concatenated input would be to use the "keyword" tokenizer.

On Thursday, October 31, 2013 8:10:55 PM UTC+1, amit.soni wrote:

> Hi all - I have a phone number field and I am trying to use word\_delimiter  
> filter in order break it up into tokens, preserve the original entry and  
> concatenate all the numbers in the entry. I have the following entry:
> 
> "phoneAnalyzer" : {  
> "type": "custom",  
> "tokenizer": "standard",  
> "filter": [  
> "word\_delimiter\_for\_phone"  
> ]  
> }
> 
> "filter": {  
> "word\_delimiter\_for\_phone": {  
> "type": "word\_delimiter",
> 
> - 
> 
> ```
> "catenate_numbers" : true,*
> "preserve_original" : true 
> },
> 
> ```
> 
> }
> 
> Using this, when I run it on input "345 678-1234" I get the following:
> 
> {  
> "tokens" : [ {  
> "token" : "_345_",  
> "start\_offset" : 0,  
> "end\_offset" : 3,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "_678_",  
> "start\_offset" : 4,  
> "end\_offset" : 7,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "_1234_",  
> "start\_offset" : 8,  
> "end\_offset" : 12,  
> "type" : "",  
> "position" : 3  
> } ]  
> }
> 
> Question: Should this also not have generated a concatenated string of the  
> form: 3456781234.
> 
> Anything I am missing here?
> 
> -Amit.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 1, 2013, 8:05am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201/3 "2013-11-01T08:05:56Z")

</div>

Or disable analysis for this field.

HTH

--  
David 😉  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

Le 1 nov. 2013 à 07:42, [sina.tamanna@gmail.com](mailto:sina.tamanna@gmail.com) a écrit :

> Analysis starts by using tokenizer, which in your case is "standard". Therefore the input "345 678-1234" will be tokenized to "345", "678", and "1234", and only then the filters will be applied. A solution to get the original and the concatenated input would be to use the "keyword" tokenizer.
> 
> On Thursday, October 31, 2013 8:10:55 PM UTC+1, amit.soni wrote:
> 
> > Hi all - I have a phone number field and I am trying to use word\_delimiter filter in order break it up into tokens, preserve the original entry and concatenate all the numbers in the entry. I have the following entry:
> > 
> > "phoneAnalyzer" : {  
> > "type": "custom",  
> > "tokenizer": "standard",  
> > "filter": [  
> > "word\_delimiter\_for\_phone"  
> > ]  
> > }
> > 
> > "filter": {  
> > "word\_delimiter\_for\_phone": {  
> > "type": "word\_delimiter",  
> > "catenate\_numbers" : true,  
> > "preserve\_original" : true  
> > },  
> > }
> > 
> > Using this, when I run it on input "345 678-1234" I get the following:
> > 
> > {  
> > "tokens" : [ {  
> > "token" : "345",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "",  
> > "position" : 1  
> > }, {  
> > "token" : "678",  
> > "start\_offset" : 4,  
> > "end\_offset" : 7,  
> > "type" : "",  
> > "position" : 2  
> > }, {  
> > "token" : "1234",  
> > "start\_offset" : 8,  
> > "end\_offset" : 12,  
> > "type" : "",  
> > "position" : 3  
> > } ]  
> > }
> > 
> > Question: Should this also not have generated a concatenated string of the form: 3456781234.
> > 
> > Anything I am missing here?
> > 
> > -Amit.
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 1, 2013, 8:07am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201/4 "2013-11-01T08:07:03Z")

</div>

Sorry. Forget my answer. Useless here.

--  
David 😉  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

Le 1 nov. 2013 à 08:05, David Pilato [david@pilato.fr](mailto:david@pilato.fr) a écrit :

> Or disable analysis for this field.
> 
> HTH
> 
> --  
> David 😉  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> Le 1 nov. 2013 à 07:42, [sina.tamanna@gmail.com](mailto:sina.tamanna@gmail.com) a écrit :
> 
> > Analysis starts by using tokenizer, which in your case is "standard". Therefore the input "345 678-1234" will be tokenized to "345", "678", and "1234", and only then the filters will be applied. A solution to get the original and the concatenated input would be to use the "keyword" tokenizer.
> > 
> > On Thursday, October 31, 2013 8:10:55 PM UTC+1, amit.soni wrote:
> > 
> > > Hi all - I have a phone number field and I am trying to use word\_delimiter filter in order break it up into tokens, preserve the original entry and concatenate all the numbers in the entry. I have the following entry:
> > > 
> > > "phoneAnalyzer" : {  
> > > "type": "custom",  
> > > "tokenizer": "standard",  
> > > "filter": [  
> > > "word\_delimiter\_for\_phone"  
> > > ]  
> > > }
> > > 
> > > "filter": {  
> > > "word\_delimiter\_for\_phone": {  
> > > "type": "word\_delimiter",  
> > > "catenate\_numbers" : true,  
> > > "preserve\_original" : true  
> > > },  
> > > }
> > > 
> > > Using this, when I run it on input "345 678-1234" I get the following:
> > > 
> > > {  
> > > "tokens" : [ {  
> > > "token" : "345",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 3,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "678",  
> > > "start\_offset" : 4,  
> > > "end\_offset" : 7,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "1234",  
> > > "start\_offset" : 8,  
> > > "end\_offset" : 12,  
> > > "type" : "",  
> > > "position" : 3  
> > > } ]  
> > > }
> > > 
> > > Question: Should this also not have generated a concatenated string of the form: 3456781234.
> > > 
> > > Anything I am missing here?
> > > 
> > > -Amit.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).  
> > --  
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Amit\_Soni](https://avatars.discourse-cdn.com/v4/letter/a/c0e974/32.png) [@Amit\_Soni](https://discuss.elastic.co/u/Amit_Soni)\
**Post date:** [April 22, 2014, 3:46am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201/5 "2014-04-22T03:46:06Z")

</div>

hi everyone - I have changed the mapping so that it now looks like below.  
However for a given input say 123-456-8989, the generated tokens are:

a) 123-456-8989 b) 123 c) 456 d) 8989 e) 1234568989

I was expecting just two tokens: a) 123-456-8989 b) 1234568989

Would you know what might be going wrong here?

"default\_index": {  
"tokenizer": "keyword",  
"filter": [  
"lowercase"  
]  
},

"phoneAnalyzer": {  
"type": "custom",  
"tokenizer": "keyword",  
"filter": [  
"word\_delimiter\_for\_phone"  
]  
},

"word\_delimiter\_for\_phone": {  
"type": "word\_delimiter",  
"catenate\_all": true,  
"generate\_number\_parts ": false,  
"split\_on\_case\_change": false,  
"generate\_word\_parts": false,  
"split\_on\_numerics": false,  
"preserve\_original": true  
},

-Amit.

On Fri, Nov 1, 2013 at 1:07 AM, David Pilato [david@pilato.fr](mailto:david@pilato.fr) wrote:

> Sorry. Forget my answer. Useless here.
> 
> --  
> David 😉  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> Le 1 nov. 2013 à 08:05, David Pilato [david@pilato.fr](mailto:david@pilato.fr) a écrit :
> 
> Or disable analysis for this field.
> 
> HTH
> 
> --  
> David 😉  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> Le 1 nov. 2013 à 07:42, [sina.tamanna@gmail.com](mailto:sina.tamanna@gmail.com) a écrit :
> 
> Analysis starts by using tokenizer, which in your case is "standard".  
> Therefore the input "345 678-1234" will be tokenized to "345", "678", and  
> "1234", and only then the filters will be applied. A solution to get the  
> original and the concatenated input would be to use the "keyword" tokenizer.
> 
> On Thursday, October 31, 2013 8:10:55 PM UTC+1, amit.soni wrote:
> 
> > Hi all - I have a phone number field and I am trying to use  
> > word\_delimiter filter in order break it up into tokens, preserve the  
> > original entry and concatenate all the numbers in the entry. I have the  
> > following entry:
> > 
> > "phoneAnalyzer" : {  
> > "type": "custom",  
> > "tokenizer": "standard",  
> > "filter": [  
> > "word\_delimiter\_for\_phone"  
> > ]  
> > }
> > 
> > "filter": {  
> > "word\_delimiter\_for\_phone": {  
> > "type": "word\_delimiter",
> > 
> > - 
> > 
> > ```
> > "catenate_numbers" : true,*
> > "preserve_original" : true
> > },
> > 
> > ```
> > 
> > }
> > 
> > Using this, when I run it on input "345 678-1234" I get the following:
> > 
> > {  
> > "tokens" : [ {  
> > "token" : "_345_",  
> > "start\_offset" : 0,  
> > "end\_offset" : 3,  
> > "type" : "",  
> > "position" : 1  
> > }, {  
> > "token" : "_678_",  
> > "start\_offset" : 4,  
> > "end\_offset" : 7,  
> > "type" : "",  
> > "position" : 2  
> > }, {  
> > "token" : "_1234_",  
> > "start\_offset" : 8,  
> > "end\_offset" : 12,  
> > "type" : "",  
> > "position" : 3  
> > } ]  
> > }
> > 
> > Question: Should this also not have generated a concatenated string of  
> > the form: 3456781234.
> > 
> > Anything I am missing here?
> > 
> > -Amit.
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAAOGaQKiEQhwJFfVwTBEHkeF%2BCK%2B8zpw6WC%2BpmSDeUgjTFtN2Q%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAAOGaQKiEQhwJFfVwTBEHkeF%2BCK%2B8zpw6WC%2BpmSDeUgjTFtN2Q%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:34am UTC](https://discuss.elastic.co/t/issue-with-using-word-delimiter-filter/14201/6 "2017-07-06T01:34:25Z")

</div>


