# Non-standart analizer/tokenizer

**URL:** <https://discuss.elastic.co/t/non-standart-analizer-tokenizer/41139>\
**Category:** Elasticsearch\
**Created:** [February 7, 2016, 1:28pm UTC](https://discuss.elastic.co/t/non-standart-analizer-tokenizer/41139 "2016-02-07T13:28:58Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![enp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/enp/32/3825_2.png) [@enp](https://discuss.elastic.co/u/enp)\
**Post date:** [February 7, 2016, 1:28pm UTC](https://discuss.elastic.co/t/non-standart-analizer-tokenizer/41139/1 "2016-02-07T13:28:58Z")

</div>

Hi,

What is the best way to configure non-standart analizer/tokenizer: split not only my whitespaces but even with underscores, dashes and slashes, exclude numbers - so only words with more than 3 letter without numbers and in lower case must stay?

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [February 7, 2016, 8:46pm UTC](https://discuss.elastic.co/t/non-standart-analizer-tokenizer/41139/2 "2016-02-07T20:46:41Z")

</div>

If any of the tokenizers [here](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-tokenizers.html) will do then you can just configure them. The Pattern tokenizer lets you define a regex and so its super flexible. It might be the best thing in your case.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:18pm UTC](https://discuss.elastic.co/t/non-standart-analizer-tokenizer/41139/3 "2017-07-05T23:18:18Z")

</div>


