# Custom tokenizer for letter and digits

**URL:** <https://discuss.elastic.co/t/custom-tokenizer-for-letter-and-digits/225317>\
**Category:** Elasticsearch\
**Created:** [March 27, 2020, 12:44am UTC](https://discuss.elastic.co/t/custom-tokenizer-for-letter-and-digits/225317 "2020-03-27T00:44:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![lfcnassif](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lfcnassif/32/49500_2.png) [@lfcnassif](https://discuss.elastic.co/u/lfcnassif)\
**Post date:** [March 27, 2020, 12:44am UTC](https://discuss.elastic.co/t/custom-tokenizer-for-letter-and-digits/225317/1 "2020-03-27T00:44:28Z")

</div>

Hi,

Searched the docs and I was not able to find a solution to create a custom tokenizer to break text at any char different from digit or unicode letter (like those returned by java Character.isLetterOrDigit()). In the past I coded one using pure Lucene...

In Elastic, tried the simple pattern tokenizer, creating a regex with all chars returned by java Character.isLetterOrDigit() (154,137 chars) but that caused a stack overflow and put Elastic down.

I would not like to use standard pattern tokenizer because it is slow. Had bad experience with java regex in the past...

Found char group tokenizer, but, if I understood correctly, I need exactly the opposite: be able to define valid token chars, not delimiter chars.

Thanks,  
Luis Nassif

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [March 30, 2020, 3:30pm UTC](https://discuss.elastic.co/t/custom-tokenizer-for-letter-and-digits/225317/2 "2020-03-30T15:30:27Z")

</div>

Hey,

I do not think that there is an out of the box tokenizer doing that for you. The `LetterTokenizer` checks for letters only, and you could probably use that one, and write a plugin based on that. You would need to implement `AnalysisPlugin` and write your own plugin. See [https://github.com/elastic/elasticsearch/blob/master/plugins/analysis-stempel/src/main/java/org/elasticsearch/plugin/analysis/stempel/AnalysisStempelPlugin.java](https://github.com/elastic/elasticsearch/blob/master/plugins/analysis-stempel/src/main/java/org/elasticsearch/plugin/analysis/stempel/AnalysisStempelPlugin.java) and [https://github.com/elastic/elasticsearch/tree/master/plugins/examples](https://github.com/elastic/elasticsearch/tree/master/plugins/examples) for some help on how to do that.

--Alex

---

<div class="post-metadata">

**Author:** ![lfcnassif](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lfcnassif/32/49500_2.png) [@lfcnassif](https://discuss.elastic.co/u/lfcnassif)\
**Post date:** [March 30, 2020, 3:49pm UTC](https://discuss.elastic.co/t/custom-tokenizer-for-letter-and-digits/225317/3 "2020-03-30T15:49:38Z")

</div>

Thank you for replying. I didn't know it is possible to write analysis plugins, will take a look.

Luis

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 27, 2020, 3:49pm UTC](https://discuss.elastic.co/t/custom-tokenizer-for-letter-and-digits/225317/4 "2020-04-27T15:49:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
