# Combining multiple Tokenizer features on single \_all field

**URL:** <https://discuss.elastic.co/t/combining-multiple-tokenizer-features-on-single-all-field/148288>\
**Category:** Elasticsearch\
**Created:** [September 12, 2018, 9:46am UTC](https://discuss.elastic.co/t/combining-multiple-tokenizer-features-on-single-all-field/148288 "2018-09-12T09:46:24Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![cdekker](https://avatars.discourse-cdn.com/v4/letter/c/f05b48/32.png) [@cdekker](https://discuss.elastic.co/u/cdekker)\
**Post date:** [September 12, 2018, 9:46am UTC](https://discuss.elastic.co/t/combining-multiple-tokenizer-features-on-single-all-field/148288/1 "2018-09-12T09:46:24Z")

</div>

We have several customer defined indices on ES 6 with 100+ fields, where each field has a copy\_to mapping to an [\_all](https://www.elastic.co/guide/en/elasticsearch/reference/current/mapping-all-field.html) field. This allows us to perform full-text search over all user-defined fields in the index.

I have several specific tokenizer requirements for this (and any other) field in those indices:

1. Emails should be tokenized as-is and not broken up: [uax\_url\_email tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-uaxurlemail-tokenizer.html)
2. Support non-western languages: [icu\_tokenizer](https://www.elastic.co/guide/en/elasticsearch/plugins/current/analysis-icu-tokenizer.html)
3. (Company) domain names should be normalized without TLD ('[Amazon.com](http://Amazon.com)' \> 'Amazon'), so they will match queries without the '.com'.

I currently implemented 2. and 3. as follows in one Analyzer:

```
"analysis": {
  "filter": {
    "domain_name": {
      "type": "pattern_capture",
      "preserve_original": "true",
      "patterns": [
        "^(?:www\\.)?([^.]{3,})\\.[^.]+"
      ]
    }
  },
  "analyzer": {
    "icu": {
      "filter": [
        "icu_folding",
        "domain_name"
      ],
      "type": "custom",
      "tokenizer": "icu_tokenizer"
    }
  }
}

```

How can I also add requirement 1. to this to support email addresses? How can I somehow 'combine' the 2 different tokenizers?

Is there a better way to implement the (company) domain name tokenization?

---

<div class="post-metadata">

**Author:** ![cdekker](https://avatars.discourse-cdn.com/v4/letter/c/f05b48/32.png) [@cdekker](https://discuss.elastic.co/u/cdekker)\
**Post date:** [September 18, 2018, 9:39am UTC](https://discuss.elastic.co/t/combining-multiple-tokenizer-features-on-single-all-field/148288/2 "2018-09-18T09:39:55Z")

</div>

Does anyone have an idea on how to achieve the 3 different tokenizations of terms?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 16, 2018, 9:39am UTC](https://discuss.elastic.co/t/combining-multiple-tokenizer-features-on-single-all-field/148288/3 "2018-10-16T09:39:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
