# Elastic tokenizer customization question

**URL:** <https://discuss.elastic.co/t/elastic-tokenizer-customization-question/233378>\
**Category:** Elasticsearch\
**Created:** [May 19, 2020, 5:00pm UTC](https://discuss.elastic.co/t/elastic-tokenizer-customization-question/233378 "2020-05-19T17:00:53Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Attila816](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/attila816/32/68007_2.png) [@Attila816](https://discuss.elastic.co/u/Attila816)\
**Post date:** [May 19, 2020, 5:00pm UTC](https://discuss.elastic.co/t/elastic-tokenizer-customization-question/233378/1 "2020-05-19T17:00:53Z")

</div>

How can I use a tokenizer which is similar to the [Word Delimiter Graph Token tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-word-delimiter-graph-tokenfilter.html) but without using the following rules:  
• Split tokens at letter case transitions. For example: PowerShot → Power, Shot  
• Split tokens at letter-number transitions. For example: XL500 → XL, 500  
• Remove the English possessive ('s) from the end of each token. For example: Neil's → Neil

So from the example: "Neil's-Super-Duper-XL500--42+AutoCoder"  
instead of these tokens:  
[Neil, Super, Duper, XL, 500, 42, Auto, Coder]  
the analyzer need to produce these tokens:  
[Neil, s, Super, Duper, XL500, 42, AutoCoder]

Thanks, Attila

---

<div class="post-metadata">

**Author:** ![Attila816](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/attila816/32/68007_2.png) [@Attila816](https://discuss.elastic.co/u/Attila816)\
**Post date:** [May 19, 2020, 9:29pm UTC](https://discuss.elastic.co/t/elastic-tokenizer-customization-question/233378/2 "2020-05-19T21:29:52Z")

</div>

I've found the solution on [elasticsearch site](https://www.elastic.co/guide/en/elasticsearch/reference/7.x/analysis-word-delimiter-graph-tokenfilter.html#analysis-word-delimiter-graph-tokenfilter-customize).

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 19, 2020, 11:11pm UTC](https://discuss.elastic.co/t/elastic-tokenizer-customization-question/233378/3 "2020-05-19T23:11:28Z")

</div>

Thanks for sharing your solution 🙂

---

<div class="post-metadata">

**Author:** ![Attila816](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/attila816/32/68007_2.png) [@Attila816](https://discuss.elastic.co/u/Attila816)\
**Post date:** [May 20, 2020, 8:48am UTC](https://discuss.elastic.co/t/elastic-tokenizer-customization-question/233378/4 "2020-05-20T08:48:00Z")

</div>

Thanks @warkolm,  
my problem is when I try to search for a text field which contains concatenated text and numbers in this formula: "{text}{number}" like DOC0000000009 then I don't know how to search on them with queries like these:

- DOC0000000009 : For this I tried to use SpanTermQuery, MatchQuery without success
- DOC000000004? : For this I tried to use WildcardQuery without success

At indexing time I set WordDelimiterGraphTokenFilter and lowercase filter for this text field analyzer and search analyzer property.  
The queries work only with lowercase letters like doc0000000009, doc000000004 even if I try to use MatchQuery with setting the same analyzer.

I am only able to execute these queries with QuerystringQuery but if I use then I cannot use a ProximityQuery which contains QueryStringQuery. Therefore I can use proximity queries with lowercased queries.

Could You please help in that?

Thanks, Attila

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 17, 2020, 8:48am UTC](https://discuss.elastic.co/t/elastic-tokenizer-customization-question/233378/5 "2020-06-17T08:48:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
