# Create an analyzer to tokenize non-alphanumeric characters

**URL:** <https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059>\
**Category:** Elasticsearch\
**Created:** [September 20, 2016, 10:23pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059 "2016-09-20T22:23:04Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Josh\_Harrison](https://avatars.discourse-cdn.com/v4/letter/j/0ea827/32.png) [@Josh\_Harrison](https://discuss.elastic.co/u/Josh_Harrison)\
**Post date:** [September 20, 2016, 10:23pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/1 "2016-09-20T22:23:04Z")

</div>

I want a custom analyzer that can take a string like "((hello world!))" and give me a token list of:  
["(", "(", "hello", "world", "!", ")", ")"]  
That is to say, I basically want the "letter" tokenizer, but I want to keep the non letter characters and tokenize them as a single character length token.

Is this feasible?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [September 21, 2016, 4:57pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/2 "2016-09-21T16:57:31Z")

</div>

Hey there,

have you considered using the [pattern tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/2.4/analysis-pattern-tokenizer.html) to define your own regex for tokenization?

--Alex

---

<div class="post-metadata">

**Author:** ![Josh\_Harrison](https://avatars.discourse-cdn.com/v4/letter/j/0ea827/32.png) [@Josh\_Harrison](https://discuss.elastic.co/u/Josh_Harrison)\
**Post date:** [September 23, 2016, 4:45pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/3 "2016-09-23T16:45:25Z")

</div>

Ah, ok - this isn't well documented but to be able to write a pattern for tokens (instead of delimiters, by default), you have to use the "group" value!  
I've put together the following:

> {  
> "settings": {  
> "analysis": {  
> "analyzer": {  
> "my\_analyzer": {  
> "tokenizer": "my\_tokenizer",  
> "filter":["lowercase"]  
> }  
> },  
> "tokenizer": {  
> "my\_tokenizer": {  
> "type": "pattern",  
> "pattern": "(\W|\w+)",  
> "group": 1,  
> "flags":"UNICODE\_CASE"  
> }  
> }  
> }  
> }  
> }

I'd ideally like to use the UNICODE\_CHARACTER\_CLASS Java regex flag, but I get an error when I include it. This means that values in Chinese, Japanese, etc, are being treated as non-letters and each letter is therefore a single token. Is there any way to do this without using CJK?

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [September 23, 2016, 5:01pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/4 "2016-09-23T17:01:50Z")

</div>

If you have an idea for improving the documentation for the tokenizer can you open an issue? I'd certainly be happy to review it.

Both `UNICODE_CHAR_CLASS` and `UNICODE_CHARACTER_CLASS` should work. What error are you seeing?

---

<div class="post-metadata">

**Author:** ![Josh\_Harrison](https://avatars.discourse-cdn.com/v4/letter/j/0ea827/32.png) [@Josh\_Harrison](https://discuss.elastic.co/u/Josh_Harrison)\
**Post date:** [September 23, 2016, 5:15pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/5 "2016-09-23T17:15:29Z")

</div>

When I attempt to PUT

> ```
> {
> "settings": {
> "analysis": {
> "analyzer": {
> "my_analyzer": {
> "tokenizer": "my_tokenizer",
> "filter":["lowercase"]
> }
> },
> "tokenizer": {
> "my_tokenizer": {
> "type": "pattern",
> "pattern": "(\\W|\\w+)",
> "group": 1, 
> "flags":"UNICODE_CASE|UNICODE_CHARACTER_CLASS"
> }
> }
> }
> }
> }
> 
> ```

I get:

> ```
> {
> "error": {
> "root_cause": [
> {
> "type": "index_creation_exception",
> "reason": "failed to create index"
> }
> ],
> "type": "illegal_argument_exception",
> "reason": "Unknown regex flag [UNICODE_CHARACTER_CLASS]"
> },
> "status": 400
> }
> 
> ```

This is on ES 2.3.5

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [September 23, 2016, 5:45pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/6 "2016-09-23T17:45:16Z")

</div>

Looks like only `UNICODE_CHAR_CLASS` is supported in 2.3: [https://github.com/elastic/elasticsearch/blob/2.3/core/src/main/java/org/elasticsearch/common/regex/Regex.java#L153](https://github.com/elastic/elasticsearch/blob/2.3/core/src/main/java/org/elasticsearch/common/regex/Regex.java#L153)

And that was fixed in 5.0:

> <https://github.com/elastic/elasticsearch/pull/11598>

Which has yet had a production release.

---

<div class="post-metadata">

**Author:** ![Josh\_Harrison](https://avatars.discourse-cdn.com/v4/letter/j/0ea827/32.png) [@Josh\_Harrison](https://discuss.elastic.co/u/Josh_Harrison)\
**Post date:** [September 23, 2016, 5:48pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/7 "2016-09-23T17:48:01Z")

</div>

Great, ok - we'll take another pass on this once 5.x is out, and we move to it!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:17pm UTC](https://discuss.elastic.co/t/create-an-analyzer-to-tokenize-non-alphanumeric-characters/61059/8 "2017-07-05T22:17:36Z")

</div>


