# Character group tokenizer in ElasticSearch

**URL:** <https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212>\
**Category:** Elasticsearch\
**Created:** [June 16, 2023, 11:27am UTC](https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212 "2023-06-16T11:27:27Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Rakhshunda\_Noorein\_J](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rakhshunda_noorein_j/32/99407_2.png) [@Rakhshunda\_Noorein\_J](https://discuss.elastic.co/u/Rakhshunda_Noorein_J)\
**Post date:** [June 16, 2023, 11:27am UTC](https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212/1 "2023-06-16T11:27:27Z")

</div>

Hello, I want to implement Character group tokenizer in elasticsearch. How Do I implement an index with `char_group` tokenizer.  
I am putting this setting in my index:

```auto
{
  "index": {
    "analysis": {
      "number_of_shards": "1",
      "analyzer": {
        "my_analyzer": {
          "tokenizer": "my_tokenizer"
        }
      },
      "tokenizer": {
        "my_tokenizer": {
          "type": "char_group",
          "tokenize_on_chars": [
            "whitespace",
            "-",
            ",",
            ":",
            "\n"
          ]
        }
      }
    }
  }
}

```

My Index mapping:

```auto
{
  "mappings": {
    "properties": {
      "@timestamp": {
        "type": "date"
      },
      "Id": {
        "type": "long"
      },
      "Name": {
        "type": "search_as_you_type",
        "doc_values": false,
        "max_shingle_size": 3
      },
      "Name_chargroup": {
        "type": "text",
        "analyzer": "my_analyzer"
      },
      "tags": {
        "type": "text",
        "fields": {
          "keyword": {
            "type": "keyword",
            "ignore_above": 256
          }
        }
      }
    }
  }
}

```

My query:

```auto
{
    
    "size":200,
    "query":{
        "multi_match":{
            "query":"gss info",
            "type":"most_fields",
            "fields":["Name_chargroup "],
            "operator": "and"
            }
        }
        }

```

The result coming as null...

The document Present is Name\_chargroup : "Gss Infotech"

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [June 19, 2023, 2:31am UTC](https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212/2 "2023-06-19T02:31:51Z")

</div>

Hi @Rakhshunda_Noorein_J

> [@Rakhshunda\_Noorein\_J](#):
>
> ```auto
> "my_analyzer": {
> "tokenizer": "my_tokenizer"
> }
> 
> ```

You need to add a "lowercase" filter to lowercase the terms.

```auto
"my_analyzer": {
          "tokenizer": "my_tokenizer",
          "filter": [
            "lowercase"
          ]
        }

```

---

<div class="post-metadata">

**Author:** ![Rakhshunda\_Noorein\_J](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rakhshunda_noorein_j/32/99407_2.png) [@Rakhshunda\_Noorein\_J](https://discuss.elastic.co/u/Rakhshunda_Noorein_J)\
**Post date:** [June 19, 2023, 11:38am UTC](https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212/3 "2023-06-19T11:38:59Z")

</div>

Hello, Now the result is coming but not as expected as char\_group tokenizer works.

my field mapping:

```auto
 "Name_chargroup": {
                    "type": "text",
                    "analyzer": "my_analyzer"
                },

```

My document is: `Gss InfoTech` my search term: `Info Gss`

result for charGroup tokenizer:

```auto
POST _analyze
{
  "tokenizer": {
    "type": "char_group",
    "tokenize_on_chars": [
      "whitespace",
      "-",
      "\n"
    ]
  },
  "text": "Info Gss"
}

response:
{
    "tokens": [
        {
            "token": "Info",
            "start_offset": 0,
            "end_offset": 4,
            "type": "word",
            "position": 0
        },
        {
            "token": "Gss",
            "start_offset": 5,
            "end_offset": 8,
            "type": "word",
            "position": 1
        }
    ]
}

```

but my rersult coming as null. On the other hand `Gss Info` and `Infotech GSS` is getting the result but not `Info Gss`

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [June 19, 2023, 12:37pm UTC](https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212/4 "2023-06-19T12:37:05Z")

</div>

> [@Rakhshunda\_Noorein\_J](#):
>
> ```auto
> POST _analyze
> {
> "tokenizer": {
> "type": "char_group",
> "tokenize_on_chars": [
> "whitespace",
> "-",
> "\n"
> ]
> },
> "text": "Info Gss"
> }
> 
> ```

Why don't you use the "standard" tokenizer?

if you index "Gss InfoTech" and the search term is "gss info" and your query Match with operator "AND' you will not have results because "infotech != info".  
If you remove the "and" the match will be on the "gss" token.  
If you want to apply the match with the term "info" you will have to use the [edge\_ngram tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-edgengram-tokenizer.html).

---

<div class="post-metadata">

**Author:** ![Rakhshunda\_Noorein\_J](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rakhshunda_noorein_j/32/99407_2.png) [@Rakhshunda\_Noorein\_J](https://discuss.elastic.co/u/Rakhshunda_Noorein_J)\
**Post date:** [June 19, 2023, 2:00pm UTC](https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212/5 "2023-06-19T14:00:53Z")

</div>

> [@RabBit\_BR](#):
>
> `my_tokenizer`

previously I had used edge\_ngram tokenizer. But the problem happening with it is -  
when I am searching suppose : `Information` , I had given max\_gram:10 and min\_gram:3, so it is` breaking information as inf, info, infr,...like that`. Because of that, Information is coming below info. Meaning.. Document Like InfoEdge, Infotech coming first that Information Technology, which I dont want.

So for this reason I wanted a tokenizer which will break the words when a whitespace is encounter.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 17, 2023, 2:01pm UTC](https://discuss.elastic.co/t/character-group-tokenizer-in-elasticsearch/336212/6 "2023-07-17T14:01:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
