# WordDelimiterTokenFilter used twice in same analyzer with different configurations causes issues

**URL:** <https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259>\
**Category:** Elasticsearch\
**Created:** [January 19, 2018, 3:00pm UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259 "2018-01-19T15:00:26Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Atul\_Bagga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atul_bagga/32/15047_2.png) [@Atul\_Bagga](https://discuss.elastic.co/u/Atul_Bagga)\
**Post date:** [January 19, 2018, 3:00pm UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/1 "2018-01-19T15:00:26Z")

</div>

ES 5.4.1

Config-1  
GenerateWordParts = true, // [wi-fi] ---\> [wi,fi]  
GenerateNumberParts = true, // [12-03] ---\> [12,03]  
CatenateWords = false, // [wi-fi] -/-\> [wifi]  
CatenateNumbers = false, // [12-03] -/-\> [1203]  
CatenateAll = false, // [wi-fi-12] -/-\> [wifi12]  
SplitOnCaseChange = false, // [WiFi] -/-\> [Wi,Fi]  
PreserveOriginal = false, // [wi-fi] -/-\> [wi-fi,wi,fi]  
SplitOnNumerics = false, // [j2ee] -/-\> [j,2,ee]  
StemEnglishPossessive = true // [Jack's] ---\> [Jack]

Config-2  
GenerateWordParts = true, // [wi-fi] ---\> [wi,fi]  
GenerateNumberParts = true, // [12-03] ---\> [12,03]  
CatenateWords = false, // [wi-fi] -/-\> [wifi]  
CatenateNumbers = false, // [12-03] -/-\> [1203]  
CatenateAll = false, // [wi-fi-12] -/-\> [wifi12]  
SplitOnCaseChange = true, // [WiFi] -/-\> [Wi,Fi]  
PreserveOriginal = true, // [wi-fi] -/-\> [wi-fi,wi,fi]  
SplitOnNumerics = true, // [j2ee] -/-\> [j,2,ee]  
StemEnglishPossessive = true // [Jack's] ---\> [Jack]

I am using these two configs of wordDelimiterFilter on same analyzer this starts giving error "startOffset must be non-negative, and endOffset must be \>= startOffset, and offsets must not go backwards" on indexing text like "AtulBagga.TestConfig". Is this an issue with wordDelimiterTokenFilter on lucene?

Requirement: I want the tokens in such a way that "AtulBagga24.TestConfig" is searchable with all of the following keywords-  
atul, bagga, atulbagga, atulbagga24, atulbagga24.textConfig, test, config, testconfig, 24

Is there a way to solve this if above approach has known issues?

---

<div class="post-metadata">

**Author:** ![AlanWoodward](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alanwoodward/32/41962_2.png) [@AlanWoodward](https://discuss.elastic.co/u/AlanWoodward)\
**Post date:** [January 31, 2018, 9:20am UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/2 "2018-01-31T09:20:32Z")

</div>

Rather than chaining WordDelimiterFilter (which will often break offsets, as you've found), can you instead use a char\_filter to remove the period? Something like:  
POST testindex/\_analyze  
{  
"char\_filter": [  
{ "type" : "pattern\_replace",  
"pattern" : "\\.",  
"replacement" : " "  
}  
],  
"tokenizer": "standard",  
"filter" : [  
{"type": "word\_delimiter",  
"generate\_word\_parts": "true",  
"generate\_number\_parts": "true",  
"catenate\_words": "true",  
"catenate\_numbers": "false",  
"catenate\_all": "false",  
"split\_on\_case\_change": "true",  
"preserve\_original": "true",  
"split\_on\_numerics": "true",  
"stem\_english\_possessive": "true"},  
"lowercase"  
],  
"text": "AtulBagga24.TestProject"  
}

---

<div class="post-metadata">

**Author:** ![Atul\_Bagga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atul_bagga/32/15047_2.png) [@Atul\_Bagga](https://discuss.elastic.co/u/Atul_Bagga)\
**Post date:** [February 1, 2018, 7:27am UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/3 "2018-02-01T07:27:25Z")

</div>

Thanks a lot for reply.  
This won't work (It will not match "atulbagga24 testproject" because of position offsets.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/0/5/05a9d76c9345e34b77bdd9c9473c282be749725a.png)

I also think there is also a genuine issue here in WordDelimiterTokenFilter WITHOUT chaining which I filed but got closed on github [[https://github.com/elastic/elasticsearch/issues/28439](https://github.com/elastic/elasticsearch/issues/28439)]

---

<div class="post-metadata">

**Author:** ![AlanWoodward](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alanwoodward/32/41962_2.png) [@AlanWoodward](https://discuss.elastic.co/u/AlanWoodward)\
**Post date:** [February 1, 2018, 10:49am UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/4 "2018-02-01T10:49:24Z")

</div>

I think it will work if you use word\_delimiter\_graph instead of word\_delimiter - the graph version also records positionLength, which is then used by query parsers to correctly construct phrase queries with gaps.

---

<div class="post-metadata">

**Author:** ![Atul\_Bagga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atul_bagga/32/15047_2.png) [@Atul\_Bagga](https://discuss.elastic.co/u/Atul_Bagga)\
**Post date:** [February 1, 2018, 2:54pm UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/5 "2018-02-01T14:54:47Z")

</div>

Thanks! It doesn't seem to work even with graph but i will try something with it.

Can you help me with confirming that this is a genuine issue [[https://github.com/elastic/elasticsearch/issues/28439](https://github.com/elastic/elasticsearch/issues/28439) ?

---

<div class="post-metadata">

**Author:** ![AlanWoodward](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alanwoodward/32/41962_2.png) [@AlanWoodward](https://discuss.elastic.co/u/AlanWoodward)\
**Post date:** [February 15, 2018, 2:36pm UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/6 "2018-02-15T14:36:41Z")

</div>

Hi @Atul_Bagga,

I think this is an issue in WordDelimiterFilter itself (so in lucene, rather than in ES). As currently implemented, if a token is broken on case change multiple times, catenate\_words will then string all of those subtokens together, but it won't produce the intermediate tokens. For example, the token 'OneTwoThree' would produce 'One', 'Two', 'Three' and 'OneTwoThree', but not 'OneTwo' or 'TwoThree'.

In your example, is 'AtulBagga24.TestProject' a standalone field, or part of a larger run of text? There might be ways to combine WDF with shingles if it's standalone.

Stringing together WordDelimiterFilters will always be broken, it looks like, because neither WDF nor WDGF can consume token graphs, and they both produce graphs (correctly in the case of WDGF, broken ones in the case of WDF). Again, that's a lucene issue.

---

<div class="post-metadata">

**Author:** ![Atul\_Bagga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atul_bagga/32/15047_2.png) [@Atul\_Bagga](https://discuss.elastic.co/u/Atul_Bagga)\
**Post date:** [February 21, 2018, 8:55am UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/7 "2018-02-21T08:55:16Z")

</div>

Thanks @AlanWoodward!

It is a standalone field. Yes, Shingle filter solves the problem partially. (AtulBagga24 and TestProject are generated but the positions are still not adjacent)

For fixing this I am keeping the same field analyzed using different analyzer and doing a search on both fields.  
Combining highlighting from two fields is a bit of pain in that case but I can live with that for now. I see there is already an open issue to combine highlights for unified highlighter for multi field approach.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 21, 2018, 8:55am UTC](https://discuss.elastic.co/t/worddelimitertokenfilter-used-twice-in-same-analyzer-with-different-configurations-causes-issues/116259/8 "2018-03-21T08:55:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
