# Pattern analyzer regex help

**URL:** <https://discuss.elastic.co/t/pattern-analyzer-regex-help/310694>\
**Category:** Elasticsearch\
**Created:** [July 26, 2022, 10:38pm UTC](https://discuss.elastic.co/t/pattern-analyzer-regex-help/310694 "2022-07-26T22:38:46Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ansamHox](https://avatars.discourse-cdn.com/v4/letter/a/54ee81/32.png) [@ansamHox](https://discuss.elastic.co/u/ansamHox)\
**Post date:** [July 26, 2022, 10:38pm UTC](https://discuss.elastic.co/t/pattern-analyzer-regex-help/310694/1 "2022-07-26T22:38:46Z")

</div>

Got a question regarding the pattern analyzer. Example text:

```auto
S6UlZgYCJaSIQcy03OOA==Ieuwc7Ix/CQfwoDSOVJl== 2oZjflRSRkcj4/OHcp78==

```

It's encrypted (each letter is hashed to 20 characters ending with `==` sign). I would like to create tokenizer for each letter, tokens should be:

```auto
S6UlZgYCJaSIQcy03OOA==
Ieuwc7Ix/CQfwoDSOVJl==
2oZjflRSRkcj4/OHcp78==

```

If I put `"pattern": "=="` it will create token without `==` sign (e.g. `S6UlZgYCJaSIQcy03OOA`). Is there any way to also include a separator as part of the token, or maybe some another logic like "take 20 characters [skip whitespace] and create 1 token, then take another 20 characters [skip whitespace] and create 2nd token, etc?

The rules are quite simple

- 1 letter = 20 characters
- Every hashed letter ends with `==` sign
- Whitespace is also a separator, if the new word starts, it will be separated by whitespace.

It would be cool if I can combine just pattern to be included in the token, and the whitespace analyzer together.

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [July 27, 2022, 12:37am UTC](https://discuss.elastic.co/t/pattern-analyzer-regex-help/310694/2 "2022-07-27T00:37:44Z")

</div>

Hi @ansamHox

I hope that help you.

```auto
PUT my-index-000001
{
  "settings": {
    "analysis": {
      "char_filter": {
        "my_char_filter": {
          "type": "pattern_replace",
          "pattern": "==",
          "replacement": "==<>"
        }
      },
      "analyzer": {
        "my_analyzer": {
          "tokenizer": "my_tokenizer",
          "char_filter": [
            "my_char_filter"
          ]
        }
      },
      "tokenizer": {
        "my_tokenizer": {
          "type": "pattern",
          "pattern": ["<>"]
        }
      }
    }
  }
}

POST my-index-000001/_analyze
{
  "analyzer": "my_analyzer",
  "text": "S6UlZgYCJaSIQcy03OOA==Ieuwc7Ix/CQfwoDSOVJl== 2oZjflRSRkcj4/OHcp78=="
}

```

---

<div class="post-metadata">

**Author:** ![ansamHox](https://avatars.discourse-cdn.com/v4/letter/a/54ee81/32.png) [@ansamHox](https://discuss.elastic.co/u/ansamHox)\
**Post date:** [July 27, 2022, 12:31pm UTC](https://discuss.elastic.co/t/pattern-analyzer-regex-help/310694/3 "2022-07-27T12:31:10Z")

</div>

yep, that's it, I never thought about replacing 🙂

just had to add whitespace in the pattern list and it works as it should:

`"pattern": ["<>", " "]`

tnx

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 24, 2022, 12:32pm UTC](https://discuss.elastic.co/t/pattern-analyzer-regex-help/310694/4 "2022-08-24T12:32:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
