# Choose Correct Text Analyzer/ Tokenizer

**URL:** <https://discuss.elastic.co/t/choose-correct-text-analyzer-tokenizer/185867>\
**Category:** Elasticsearch\
**Created:** [June 14, 2019, 1:58pm UTC](https://discuss.elastic.co/t/choose-correct-text-analyzer-tokenizer/185867 "2019-06-14T13:58:41Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![M.alsioufi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/m.alsioufi/32/52269_2.png) [@M.alsioufi](https://discuss.elastic.co/u/M.alsioufi)\
**Post date:** [June 14, 2019, 1:58pm UTC](https://discuss.elastic.co/t/choose-correct-text-analyzer-tokenizer/185867/1 "2019-06-14T13:58:42Z")

</div>

I have a text field that represents a file name, this field follow a specific format. It contains several parts separated by underscore (\_) one part could have letters, numbers and dashes (-) and it ends with a dot (.) and an extension. For example: X\_Y-B\_Z.ext.  
I want to be able to search this field by each of its parts however the "-" preventing this behavior. I think I should use some custom analyzer/ tokenizer to solve this but I am not sure how exactly I should fix it.

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [June 16, 2019, 4:28pm UTC](https://discuss.elastic.co/t/choose-correct-text-analyzer-tokenizer/185867/2 "2019-06-16T16:28:23Z")

</div>

Yes, you could use a custom analyzer with a tokenizer that just breaks on those characters that you want to break on (in your case `_` and `.` ). The [Char Group tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-chargroup-tokenizer.html) is perhaps the most straightforward choice.

Here's a little example:

```auto
PUT my_index
{
  "settings": {
    "analysis": {
      "tokenizer": {
        "my_tokenizer": {
          "type": "char_group",
          "tokenize_on_chars": [
            "_",
            "."
          ]
        }
      },
      "analyzer": {
        "my_analyzer": {
          "char_filter": [],
          "tokenizer": "my_tokenizer",
          "filter": [
            "lowercase"
          ]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "my_field": {
        "type": "text",
        "analyzer": "my_analyzer"
      }
    }
  }
}

GET my_index/_analyze
{
  "text": "X_Y-B_Z.ext",
  "analyzer": "my_analyzer"
}

PUT my_index/_doc/1
{
  "my_field": "X_Y-B_Z.ext"
}

GET my_index/_search
{
 "query": {
   "match": {
     "my_field": "y-b"
   }
 } 
}

```

If you don't need case-insensitive search you can remove the `lowercase` token filter from the analyzer.

---

<div class="post-metadata">

**Author:** ![M.alsioufi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/m.alsioufi/32/52269_2.png) [@M.alsioufi](https://discuss.elastic.co/u/M.alsioufi)\
**Post date:** [June 17, 2019, 11:29am UTC](https://discuss.elastic.co/t/choose-correct-text-analyzer-tokenizer/185867/3 "2019-06-17T11:29:37Z")

</div>

Thanks for the suggested solution. This solved the porblem mentioned in the post very well, however, I forgot to mention that X, Y, B, Z in the example I gave could be a set of alphanumeric characters like `ABC01_DG-102_102_203.ext` and`ABC01_DG-102_102_204.ext`.

With the provided solution if I search for: "`ABC01_DG-102_102_20*`" in Kibana searchbar I expected to get both `ABC01_DG-102_102_203.ext` and `ABC01_DG-102_102_204.ext` however, I got no results.  
How can I modify this solution to handle that case as well?

---

<div class="post-metadata">

**Author:** ![M.alsioufi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/m.alsioufi/32/52269_2.png) [@M.alsioufi](https://discuss.elastic.co/u/M.alsioufi)\
**Post date:** [June 19, 2019, 11:46am UTC](https://discuss.elastic.co/t/choose-correct-text-analyzer-tokenizer/185867/4 "2019-06-19T11:46:04Z")

</div>

I used abdon's solution with a small modification.  
I solved the issue by adding only "." in the "tokenize\_on\_chars" .

```
PUT my_index
{
  "settings": {
    "analysis": {
      "tokenizer": {
        "my_tokenizer": {
          "type": "char_group",
          "tokenize_on_chars": [
            "."
          ]
        }
      },
      "analyzer": {
        "my_analyzer": {
          "char_filter": [],
          "tokenizer": "my_tokenizer",
          "filter": [
            "lowercase"
          ]
        }
      }
    }
  },
  "mappings": {
    "doc": {
      "properties": {
        "my_field": {
          "type": "text",
          "analyzer": "my_analyzer"
        }
      }
    }
  }
}
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 17, 2019, 11:46am UTC](https://discuss.elastic.co/t/choose-correct-text-analyzer-tokenizer/185867/5 "2019-07-17T11:46:08Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
