# How do I extend the default analyzer?

**URL:** <https://discuss.elastic.co/t/how-do-i-extend-the-default-analyzer/272390>\
**Category:** Elasticsearch\
**Created:** [May 7, 2021, 8:29am UTC](https://discuss.elastic.co/t/how-do-i-extend-the-default-analyzer/272390 "2021-05-07T08:29:54Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jongpyo.lee](https://avatars.discourse-cdn.com/v4/letter/j/b38774/32.png) [@jongpyo.lee](https://discuss.elastic.co/u/jongpyo.lee)\
**Post date:** [May 7, 2021, 8:29am UTC](https://discuss.elastic.co/t/how-do-i-extend-the-default-analyzer/272390/1 "2021-05-07T08:29:54Z")

</div>

This is example data.

```auto
POST coding/_bulk
{"index":{"_id":"1"}}
{"language":"xyz_foo@abc"}

```

I confirmed that the defulat `analyzer` distinguishes `@` but not the `_` (underscore)symbol through \_termvectors.

```auto
GET coding/_termvectors/1?fields=language
{
  "_index" : "coding",
  ...
  "term_vectors" : {
    "language" : {
      ...
      "terms" : {
        "abc" : {
           ...
        },
        "xyz_foo" : {
          ...
        }
      }
    }
  }
}

```

The default `analyzer` didn't distinguish between the `_` (underscore)symbols, so I couldn't search with `xyz` or `foo`.  
How do I create an analyzer that can search up to `xyz` or `foo` and `abc` by separating the `_` (underscore)symbol?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [May 10, 2021, 9:10am UTC](https://discuss.elastic.co/t/how-do-i-extend-the-default-analyzer/272390/2 "2021-05-10T09:10:08Z")

</div>

Hey,

you need to find the proper tokenizer in order to split tokens. See this example

```auto
GET _analyze
{
  "text": ["xyz_foo@abc"],
  "tokenizer": "letter"
}

```

So the analyze API allows you to figure out how the tokens are tokenized and modified before saved in the inverted index. Take a look at [Tokenizer reference | Elasticsearch Guide [7.12] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.12/analysis-tokenizers.html) and check out which tokenizer might be for you. The char group tokenizer might be something for you as well, to come up with your own set of characaters to tokenize on [Character group tokenizer | Elasticsearch Guide [7.12] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.12/analysis-chargroup-tokenizer.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 7, 2021, 9:10am UTC](https://discuss.elastic.co/t/how-do-i-extend-the-default-analyzer/272390/3 "2021-06-07T09:10:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
