# Char\_filter doesn't work properly

**URL:** <https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218>\
**Category:** Elasticsearch\
**Created:** [July 25, 2023, 6:53pm UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218 "2023-07-25T18:53:40Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vladimir\_Talabko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vladimir_talabko/32/120862_2.png) [@Vladimir\_Talabko](https://discuss.elastic.co/u/Vladimir_Talabko)\
**Post date:** [July 25, 2023, 6:53pm UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/1 "2023-07-25T18:53:40Z")

</div>

Hello!  
I have an index with a char\_filter for avoiding special symbols like "-\_.":

```auto
curl -X PUT "localhost:9200/test_index?pretty" -H 'Content-Type: application/json' -d'
{
    "settings": {
        "analysis": {
            "char_filter": {
                "specials_char_filter": {
                    "type": "mapping",
                    "mappings": ["- =>", ". =>", "_ =>"]
                }
            },
            "analyzer": {
                "articul_analyzer": {
                    "type": "custom",
                    "tokenizer": "whitespace",
                    "char_filter": [
                        "html_strip", "specials_char_filter"
                    ],
                    "filter": [
                        "lowercase",
                        "trim"
                    ]
                }
            }
        }
    }
}
'

```

I've checked the analyzer with the next command:

```auto
curl -X POST "localhost:9200/test_index/_analyze?pretty" -H 'Content-Type: application/json' -d'
{
    "analyzer": "articul_analyzer",
    "text": "<p>U-_.298 </p> "
}
'

```

It works, I'm getting "token" : "u298",

After indexing a few records I'm trying the next search request:

```auto
curl -X GET "localhost:9200/test_index/_search?pretty" -H 'Content-Type: application/json' -d'
{
  "query": {
    "query_string": {
        "query": "U29",
        "default_field": "articul_indexed"
    }
  }
}
'

```

And it doesn't work. =(( It works for U-289, 289, and others, but for U298 doesn't.  
Could you give me some clue how to get it worked?

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [July 26, 2023, 1:49am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/2 "2023-07-26T01:49:35Z")

</div>

Hi @Vladimir_Talabko

Did you try to use [wildcards](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-query-string-query.html#query-string-wildcard)?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 26, 2023, 7:12am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/3 "2023-07-26T07:12:27Z")

</div>

It's because `u29` has not been indexed as a token but `u298` so this can not match.

If you want to do some prefix search, you could add edge n grams to your analyzer:

```auto
DELETE /test_index
PUT /test_index
{
  "settings": {
    "analysis": {
      "char_filter": {
        "specials_char_filter": {
          "type": "mapping",
          "mappings": [
            "- =>",
            ". =>",
            "_ =>"
          ]
        }
      },
      "filter": {
        "prefix": {
          "type": "edge_ngram",
          "min_gram": 2,
          "max_gram": 4
        }
      },
      "analyzer": {
        "articul_analyzer": {
          "type": "custom",
          "tokenizer": "whitespace",
          "char_filter": [
            "html_strip",
            "specials_char_filter"
          ],
          "filter": [
            "lowercase",
            "trim"
          ]
        },
        "articul_analyzer_prefix": {
          "type": "custom",
          "tokenizer": "whitespace",
          "char_filter": [
            "html_strip",
            "specials_char_filter"
          ],
          "filter": [
            "lowercase",
            "trim",
            "prefix"
          ]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "foo": {
        "type": "text",
        "analyzer": "articul_analyzer_prefix",
        "search_analyzer": "articul_analyzer"
      }
    }
  }
}

# At index time
POST /test_index/_analyze?pretty
{
    "analyzer": "articul_analyzer_prefix",
    "text": "<p>U-_.298 </p> "
}

# At search time
POST /test_index/_analyze?pretty
{
    "analyzer": "articul_analyzer",
    "text": "U29"
}

```

---

<div class="post-metadata">

**Author:** ![Vladimir\_Talabko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vladimir_talabko/32/120862_2.png) [@Vladimir\_Talabko](https://discuss.elastic.co/u/Vladimir_Talabko)\
**Post date:** [July 26, 2023, 7:58am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/4 "2023-07-26T07:58:54Z")

</div>

Thank you, but edge\_ngram won't help me because hyphens might be several times anywhere in my texts.  
When I was indexing my texts I set mappings like so:

```auto
$workParams['mappings'] = [
    'properties' => [
        "articul_indexed" => [
            'type' => 'text',
            'analyzer' => 'articul_analyzer'
        ]
    ]
];

```

Doesn't this mapping say to the engine to tokenyze my text?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 26, 2023, 8:15am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/5 "2023-07-26T08:15:21Z")

</div>

Did you try my example with hyphens?  
If so, please share what works and what does not as a full example which can be ran in Kibana Dev Console like I provided.

---

<div class="post-metadata">

**Author:** ![Vladimir\_Talabko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vladimir_talabko/32/120862_2.png) [@Vladimir\_Talabko](https://discuss.elastic.co/u/Vladimir_Talabko)\
**Post date:** [July 26, 2023, 8:45am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/6 "2023-07-26T08:45:07Z")

</div>

Yes, it works for the "U-298" example. However, if I index text like "aaaaaa-uuuuuu", the engine even can't find the "aaaaa" (or "uuu") part from it. =(  
I just want to be able to delete some symbols (-\_.,) from a string, and then search a request (also without such symbols) starting from any symbol in my string.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 26, 2023, 11:37am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/7 "2023-07-26T11:37:35Z")

</div>

Could you illustrate that with a full example please?

---

<div class="post-metadata">

**Author:** ![Vladimir\_Talabko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vladimir_talabko/32/120862_2.png) [@Vladimir\_Talabko](https://discuss.elastic.co/u/Vladimir_Talabko)\
**Post date:** [July 26, 2023, 3:06pm UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/8 "2023-07-26T15:06:43Z")

</div>

```auto
curl -X DELETE "localhost:9200/test_index?pretty"

curl -X PUT "localhost:9200/test_index?pretty" -H 'Content-Type: application/json' -d'
{
    "settings": {
        "analysis": {
            "char_filter": {
                "specials_char_filter": {
                    "type": "mapping",
                    "mappings": ["-=>", ".=>", "_=>"]
                }
            },
            "analyzer": {
                "articul_analyzer": {
                    "type": "custom",
                    "tokenizer": "whitespace",
                    "char_filter": [
                        "html_strip", "specials_char_filter"
                    ],
                    "filter": [
                        "lowercase",
                        "trim"
                    ]
                }
            }
        }
    },
    "mappings": {
        "properties": {
          "articul_indexed": {
            "type": "text",
            "analyzer": "articul_analyzer"
          }
        }
      }
}
'

curl -X POST "localhost:9200/test_index/_analyze?pretty" -H 'Content-Type: application/json' -d'
{
    "analyzer": "articul_analyzer",
    "text": "<p>U-_.298 </p> "
}
'

```

Here I see: "token" : "u298"  
Perfect!

Two records:

```auto
curl -X PUT "localhost:9200/test_index/_doc/1?pretty" -H 'Content-Type: application/json' -d'
{
    "articul_indexed": "U-298"
}
'

curl -X PUT "localhost:9200/test_index/_doc/2?pretty" -H 'Content-Type: application/json' -d'
{
    "articul_indexed": "aaaaa-uuuuu"
}
'

```

and two tests:

```auto
curl -X GET "localhost:9200/test_index/_search?pretty" -H 'Content-Type: application/json' -d'
{
  "query": {
    "query_string": {
        "query": "articul_indexed:U298"
    }
  }
}
'

```

This one works fine!

```auto
curl -X GET "localhost:9200/test_index/_search?pretty" -H 'Content-Type: application/json' -d'
{
  "query": {
    "query_string": {
        "query": "articul_indexed:aaaaa"
    }
  }
}
'

```

This one doesn't =(

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 26, 2023, 3:31pm UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/9 "2023-07-26T15:31:48Z")

</div>

Try replacing the special characters with whitespace instead of removing them in your analyzer.

---

<div class="post-metadata">

**Author:** ![Vladimir\_Talabko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vladimir_talabko/32/120862_2.png) [@Vladimir\_Talabko](https://discuss.elastic.co/u/Vladimir_Talabko)\
**Post date:** [July 27, 2023, 6:22am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/11 "2023-07-27T06:22:04Z")

</div>

I use "tokenizer": "whitespace",

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 27, 2023, 6:27am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/12 "2023-07-27T06:27:03Z")

</div>

Yes, and that needs whitespace to work, which is why I suggested replacing with space rather than remove the characters.

When you remove the characters `aaaaa-uuuuu` will be tokenised as `aaaaauuuuu`, which means you can not search for either component. If you instead replace with space the whitespace tokenizer will tokenize it as `aaaaa` and `uuuuu`.

---

<div class="post-metadata">

**Author:** ![Vladimir\_Talabko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vladimir_talabko/32/120862_2.png) [@Vladimir\_Talabko](https://discuss.elastic.co/u/Vladimir_Talabko)\
**Post date:** [July 27, 2023, 7:47am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/13 "2023-07-27T07:47:07Z")

</div>

If `aaaaa-uuuuu` will be tokenised as `aaaaauuuuu` I can search any type as `aaa`, `uuu`, `aauu`, isn't it?  
If I get separately `aaaaa` and `uuuuu` I won't be able to find `aauu` for example.  
I just want to dismiss some characters from vendor codes because people don't type them at most, but I want to show them results regardless `aaaaa`, `uuuuu`, or `aauu`. Only the order is matter.

p.s. As the next step, by allowing 1-2 fuzzyness symbols I'm going to expand this feature.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 27, 2023, 8:01am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/14 "2023-07-27T08:01:25Z")

</div>

> [@Vladimir\_Talabko](#):
>
> If `aaaaa-uuuuu` will be tokenised as `aaaaauuuuu` I can search any type as `aaa`, `uuu`, `aauu`, isn't it?

Not necessarily. It depends on the mapping of the field. You could find substrings like in your example, but that would require a wildcard query, which is one of the most expensive and inefficient query type you can use in Elasticsearch.

If this is how you want to query your data, you might want to look into [the wildcard field type](https://www.elastic.co/guide/en/elasticsearch/reference/8.6/keyword.html#wildcard-field-type) in order to make these queries more efficient.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 24, 2023, 8:01am UTC](https://discuss.elastic.co/t/char-filter-doesnt-work-properly/339218/15 "2023-08-24T08:01:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
