# ElasticSearch won't recongize char\_filter mapping

**URL:** <https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338>\
**Category:** Elasticsearch\
**Created:** [January 14, 2013, 10:01pm UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338 "2013-01-14T22:01:44Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [January 14, 2013, 10:01pm UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338/1 "2013-01-14T22:01:44Z")

</div>

In summary: Everything is currently working, except for char\_filter mapping.

I'm currently on ElasticSearch 19.10 because it works fine for our current  
production application, and I am not extending its use to require any of  
the bug fixes in the change logs for more recent versions.

I've isolated this issue to configured analyzers and a collection of HTTP  
\_analyze requests that easily reproduce the problem. No additional data or  
queries are needed at this point (I don't believe, anyway).

Here is the example I found at  
[http://www.elasticsearch.org/guide/reference/index-modules/analysis/mapping-charfilter.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/mapping-charfilter.html)

{  
"index" : {  
"analysis" : {  
"char\_filter" : {  
"my\_mapping" : {  
"type" : "mapping",  
"mappings" : ["ph=\>f", "qu=\>q"]  
}  
},  
"analyzer" : {  
"custom\_with\_char\_filter" : {  
"tokenizer" : "standard",  
"char\_filter" : ["my\_mapping"]  
},  
}  
}  
}  
}

Here is what my elasticsearch.yml configuration looks like. Note the  
Finnish character mappings that are typical for searching Finnish names:  
The previous example didn't quite work with what I need: A snowball  
stemming tokenizer with Finnish stemming rules, no stop words, and  
convering w to v on the input string before tokenizing. After playing  
around a little, here's what works (except for the char\_filter):

index:  
analysis:  
char\_filter:  
finnish\_char\_mapping:  
type: mapping  
mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
analyzer:  
# Default uses snowball stemming analyzer with no stop words  
# with the default language per the JVM:  
default:  
type: snowball  
stopwords: _none_  
# Per-language analyzers  
english\_standard:  
type: standard  
language: English  
stopwords: _none_  
english\_stemming:  
type: snowball  
language: English  
stopwords: _none_  
finnish\_stemming:  
type: snowball  
language: Finnish  
char\_filter: [finnish\_char\_mapping]  
stopwords: _none_

This first analyze operation returns the expected tokens. It analyzes the  
text using the standard analyzer, and stop words are included:

$ curl -XGET 'localhost:9200/sgen/\_analyze?analyzer=standard&pretty=true'  
-d 'Debby Debbie and Walter' && echo  
{  
"tokens" : [ {  
"token" : "debby",  
"start\_offset" : 0,  
"end\_offset" : 5,  
"type" : "",  
"position" : 1  
}, {  
"token" : "debbie",  
"start\_offset" : 6,  
"end\_offset" : 12,  
"type" : "",  
"position" : 2  
}, {  
"token" : "walter",  
"start\_offset" : 17,  
"end\_offset" : 23,  
"type" : "",  
"position" : 4  
} ]  
}

This also works: It uses the snowball analyzer with the English language  
and with stop words included in the list of tokens as desired:

$ curl -XGET  
'localhost:9200/sgen/\_analyze?analyzer=english\_stemming&pretty=true' -d  
'Debby Debbie and Walter' && echo  
{  
"tokens" : [ {  
"token" : "debbi",  
"start\_offset" : 0,  
"end\_offset" : 5,  
"type" : "",  
"position" : 1  
}, {  
"token" : "debbi",  
"start\_offset" : 6,  
"end\_offset" : 12,  
"type" : "",  
"position" : 2  
}, {  
"token" : "and",  
"start\_offset" : 13,  
"end\_offset" : 16,  
"type" : "",  
"position" : 3  
}, {  
"token" : "walter",  
"start\_offset" : 17,  
"end\_offset" : 23,  
"type" : "",  
"position" : 4  
} ]  
}

But this doesn't fully work. It uses the Finnish stemming rules (to the  
best of my knowledge; the tokens are different than those created using the  
English snowball stemming rules). But it does not honor the character  
mapping: I would have expected "valter" and not "walter" as the last token  
string. And of course, a search for valter won't match walter and this  
analysis token issue is likely the root cause:

$ curl -XGET  
'localhost:9200/sgen/\_analyze?analyzer=finnish\_stemming&pretty=true' -d  
'Debby Debbie and Walter' && echo  
{  
"tokens" : [ {  
"token" : "deby",  
"start\_offset" : 0,  
"end\_offset" : 5,  
"type" : "",  
"position" : 1  
}, {  
"token" : "debie",  
"start\_offset" : 6,  
"end\_offset" : 12,  
"type" : "",  
"position" : 2  
}, {  
"token" : "and",  
"start\_offset" : 13,  
"end\_offset" : 16,  
"type" : "",  
"position" : 3  
}, {  
"token" : "walter",  
"start\_offset" : 17,  
"end\_offset" : 23,  
"type" : "",  
"position" : 4  
} ]  
}

I cannot get ElasticSearch to define the analyzers and mappings when  
creating an index: There aren't any examples of both, and my  
experimentation yields mappings that can only point to configured  
analyzers. So configuring a list of analyzers is a currently acceptable  
work-around.

But honoring the char\_filter mapping is something that is necessary to  
resolve.

Thank you in advance.

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 15, 2013, 2:13am UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338/2 "2013-01-15T02:13:19Z")

</div>

I don't think you can put a char\_filter on existing analyzer, I think you  
need to create a custom analyzer instead:

index:  
analysis:  
char\_filter:  
finnish\_char\_mapping:  
type: mapping  
mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
analyzer:  
# Default uses snowball stemming analyzer with no stop words  
# with the default language per the JVM:  
default:  
type: snowball  
stopwords: _none_  
# Per-language analyzers  
english\_standard:  
type: standard  
language: English  
stopwords: _none_  
english\_stemming:  
type: snowball  
language: English  
stopwords: _none_  
finnish\_stemming:  
type: custom  
tokenizer: standard  
filter: [standard, lowercase, finnish\_snowball]  
char\_filter: [finnish\_char\_mapping]  
filter:  
finnish\_snowball:  
type: snowball  
language: Finnish

On Monday, January 14, 2013 5:01:44 PM UTC-5, InquiringMind wrote:

> In summary: Everything is currently working, except for char\_filter  
> mapping.
> 
> I'm currently on Elasticsearch 19.10 because it works fine for our current  
> production application, and I am not extending its use to require any of  
> the bug fixes in the change logs for more recent versions.
> 
> I've isolated this issue to configured analyzers and a collection of HTTP  
> \_analyze requests that easily reproduce the problem. No additional data or  
> queries are needed at this point (I don't believe, anyway).
> 
> Here is the example I found at  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/mapping-charfilter.html)
> 
> {  
> "index" : {  
> "analysis" : {  
> "char\_filter" : {  
> "my\_mapping" : {  
> "type" : "mapping",  
> "mappings" : ["ph=\>f", "qu=\>q"]  
> }  
> },  
> "analyzer" : {  
> "custom\_with\_char\_filter" : {  
> "tokenizer" : "standard",  
> "char\_filter" : ["my\_mapping"]  
> },  
> }  
> }  
> }  
> }
> 
> Here is what my elasticsearch.yml configuration looks like. Note the  
> Finnish character mappings that are typical for searching Finnish names:  
> The previous example didn't quite work with what I need: A snowball  
> stemming tokenizer with Finnish stemming rules, no stop words, and  
> convering w to v on the input string before tokenizing. After playing  
> around a little, here's what works (except for the char\_filter):
> 
> index:  
> analysis:  
> char\_filter:  
> finnish\_char\_mapping:  
> type: mapping  
> mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
> analyzer:  
> # Default uses snowball stemming analyzer with no stop words  
> # with the default language per the JVM:  
> default:  
> type: snowball  
> stopwords: _none_  
> # Per-language analyzers  
> english\_standard:  
> type: standard  
> language: English  
> stopwords: _none_  
> english\_stemming:  
> type: snowball  
> language: English  
> stopwords: _none_  
> finnish\_stemming:  
> type: snowball  
> language: Finnish  
> char\_filter: [finnish\_char\_mapping]  
> stopwords: _none_
> 
> This first analyze operation returns the expected tokens. It analyzes the  
> text using the standard analyzer, and stop words are included:
> 
> $ curl -XGET 'localhost:9200/sgen/\_analyze?analyzer=standard&pretty=true'  
> -d 'Debby Debbie and Walter' && echo  
> {  
> "tokens" : [ {  
> "token" : "debby",  
> "start\_offset" : 0,  
> "end\_offset" : 5,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "debbie",  
> "start\_offset" : 6,  
> "end\_offset" : 12,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "walter",  
> "start\_offset" : 17,  
> "end\_offset" : 23,  
> "type" : "",  
> "position" : 4  
> } ]  
> }
> 
> This also works: It uses the snowball analyzer with the English language  
> and with stop words included in the list of tokens as desired:
> 
> $ curl -XGET  
> 'localhost:9200/sgen/\_analyze?analyzer=english\_stemming&pretty=true' -d  
> 'Debby Debbie and Walter' && echo  
> {  
> "tokens" : [ {  
> "token" : "debbi",  
> "start\_offset" : 0,  
> "end\_offset" : 5,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "debbi",  
> "start\_offset" : 6,  
> "end\_offset" : 12,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "and",  
> "start\_offset" : 13,  
> "end\_offset" : 16,  
> "type" : "",  
> "position" : 3  
> }, {  
> "token" : "walter",  
> "start\_offset" : 17,  
> "end\_offset" : 23,  
> "type" : "",  
> "position" : 4  
> } ]  
> }
> 
> But this doesn't fully work. It uses the Finnish stemming rules (to the  
> best of my knowledge; the tokens are different than those created using the  
> English snowball stemming rules). But it does not honor the character  
> mapping: I would have expected "valter" and not "walter" as the last token  
> string. And of course, a search for valter won't match walter and this  
> analysis token issue is likely the root cause:
> 
> $ curl -XGET  
> 'localhost:9200/sgen/\_analyze?analyzer=finnish\_stemming&pretty=true' -d  
> 'Debby Debbie and Walter' && echo  
> {  
> "tokens" : [ {  
> "token" : "deby",  
> "start\_offset" : 0,  
> "end\_offset" : 5,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "debie",  
> "start\_offset" : 6,  
> "end\_offset" : 12,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "and",  
> "start\_offset" : 13,  
> "end\_offset" : 16,  
> "type" : "",  
> "position" : 3  
> }, {  
> "token" : "walter",  
> "start\_offset" : 17,  
> "end\_offset" : 23,  
> "type" : "",  
> "position" : 4  
> } ]  
> }
> 
> I cannot get Elasticsearch to define the analyzers and mappings when  
> creating an index: There aren't any examples of both, and my  
> experimentation yields mappings that can only point to configured  
> analyzers. So configuring a list of analyzers is a currently acceptable  
> work-around.
> 
> But honoring the char\_filter mapping is something that is necessary to  
> resolve.
> 
> Thank you in advance.

--

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [January 15, 2013, 4:12pm UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338/3 "2013-01-15T16:12:54Z")

</div>

Thank you very much, Igor! I wouldn't have guessed it, but from your  
example it now makes sense.

I added the stopwords : _none_ configuration line to the finnish\_stemminganalyzer. It seems to work (that is, it tokenizes stop words) for the few  
Finnish stop word examples I could find. Though my very limited knowledge  
of Finnish doesn't let me verify the behavior.

I also created a counterpart that is non-stemming but still performs  
character replacements. Here is my current configuration for analyzers:

index:  
analysis:  
char\_filter:  
finnish\_char\_mapping:  
type: mapping  
mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
filter:  
finnish\_standard:  
type: standard  
language: Finnish  
finnish\_snowball:  
type: snowball  
language: Finnish  
analyzer:  
# Default analyzer uses the snowball stemming analyzer with no  
# stop words, all defined by the default language per the JVM:  
default:  
type: snowball  
stopwords: _none_  
# Per-language analyzers: No stemming  
english\_standard:  
type: standard  
language: English  
stopwords: _none_  
finnish\_standard:  
type: custom  
tokenizer: standard  
filter: [standard, lowercase, finnish\_standard]  
char\_filter: [finnish\_char\_mapping]  
stopwords: _none_  
# Per-language analyzers: Stemming  
english\_stemming:  
type: snowball  
language: English  
stopwords: _none_  
finnish\_stemming:  
type: custom  
tokenizer: standard  
filter: [standard, lowercase, finnish\_snowball]  
char\_filter: [finnish\_char\_mapping]  
stopwords: _none_

Thanks again!

On Monday, January 14, 2013 9:13:19 PM UTC-5, Igor Motov wrote:

> I don't think you can put a char\_filter on existing analyzer, I think you  
> need to create a custom analyzer instead:
> 
> index:  
> analysis:  
> char\_filter:  
> finnish\_char\_mapping:  
> type: mapping  
> mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
> analyzer:  
> # Default uses snowball stemming analyzer with no stop words  
> # with the default language per the JVM:  
> default:  
> type: snowball  
> stopwords: _none_  
> # Per-language analyzers  
> english\_standard:  
> type: standard  
> language: English  
> stopwords: _none_  
> english\_stemming:  
> type: snowball  
> language: English  
> stopwords: _none_  
> finnish\_stemming:  
> type: custom  
> tokenizer: standard  
> filter: [standard, lowercase, finnish\_snowball]  
> char\_filter: [finnish\_char\_mapping]  
> filter:  
> finnish\_snowball:  
> type: snowball  
> language: Finnish
> 
> On Monday, January 14, 2013 5:01:44 PM UTC-5, InquiringMind wrote:
> 
> > In summary: Everything is currently working, except for char\_filter  
> > mapping.
> > 
> > I'm currently on Elasticsearch 19.10 because it works fine for our  
> > current production application, and I am not extending its use to require  
> > any of the bug fixes in the change logs for more recent versions.
> > 
> > I've isolated this issue to configured analyzers and a collection of HTTP  
> > \_analyze requests that easily reproduce the problem. No additional data or  
> > queries are needed at this point (I don't believe, anyway).
> > 
> > Here is the example I found at  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/mapping-charfilter.html)
> > 
> > {  
> > "index" : {  
> > "analysis" : {  
> > "char\_filter" : {  
> > "my\_mapping" : {  
> > "type" : "mapping",  
> > "mappings" : ["ph=\>f", "qu=\>q"]  
> > }  
> > },  
> > "analyzer" : {  
> > "custom\_with\_char\_filter" : {  
> > "tokenizer" : "standard",  
> > "char\_filter" : ["my\_mapping"]  
> > },  
> > }  
> > }  
> > }  
> > }
> > 
> > Here is what my elasticsearch.yml configuration looks like. Note the  
> > Finnish character mappings that are typical for searching Finnish names:  
> > The previous example didn't quite work with what I need: A snowball  
> > stemming tokenizer with Finnish stemming rules, no stop words, and  
> > convering w to v on the input string before tokenizing. After playing  
> > around a little, here's what works (except for the char\_filter):
> > 
> > index:  
> > analysis:  
> > char\_filter:  
> > finnish\_char\_mapping:  
> > type: mapping  
> > mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
> > analyzer:  
> > # Default uses snowball stemming analyzer with no stop words  
> > # with the default language per the JVM:  
> > default:  
> > type: snowball  
> > stopwords: _none_  
> > # Per-language analyzers  
> > english\_standard:  
> > type: standard  
> > language: English  
> > stopwords: _none_  
> > english\_stemming:  
> > type: snowball  
> > language: English  
> > stopwords: _none_  
> > finnish\_stemming:  
> > type: snowball  
> > language: Finnish  
> > char\_filter: [finnish\_char\_mapping]  
> > stopwords: _none_
> > 
> > This first analyze operation returns the expected tokens. It analyzes the  
> > text using the standard analyzer, and stop words are included:
> > 
> > $ curl -XGET 'localhost:9200/sgen/\_analyze?analyzer=standard&pretty=true'  
> > -d 'Debby Debbie and Walter' && echo  
> > {  
> > "tokens" : [ {  
> > "token" : "debby",  
> > "start\_offset" : 0,  
> > "end\_offset" : 5,  
> > "type" : "",  
> > "position" : 1  
> > }, {  
> > "token" : "debbie",  
> > "start\_offset" : 6,  
> > "end\_offset" : 12,  
> > "type" : "",  
> > "position" : 2  
> > }, {  
> > "token" : "walter",  
> > "start\_offset" : 17,  
> > "end\_offset" : 23,  
> > "type" : "",  
> > "position" : 4  
> > } ]  
> > }
> > 
> > This also works: It uses the snowball analyzer with the English language  
> > and with stop words included in the list of tokens as desired:
> > 
> > $ curl -XGET  
> > 'localhost:9200/sgen/\_analyze?analyzer=english\_stemming&pretty=true' -d  
> > 'Debby Debbie and Walter' && echo  
> > {  
> > "tokens" : [ {  
> > "token" : "debbi",  
> > "start\_offset" : 0,  
> > "end\_offset" : 5,  
> > "type" : "",  
> > "position" : 1  
> > }, {  
> > "token" : "debbi",  
> > "start\_offset" : 6,  
> > "end\_offset" : 12,  
> > "type" : "",  
> > "position" : 2  
> > }, {  
> > "token" : "and",  
> > "start\_offset" : 13,  
> > "end\_offset" : 16,  
> > "type" : "",  
> > "position" : 3  
> > }, {  
> > "token" : "walter",  
> > "start\_offset" : 17,  
> > "end\_offset" : 23,  
> > "type" : "",  
> > "position" : 4  
> > } ]  
> > }
> > 
> > But this doesn't fully work. It uses the Finnish stemming rules (to the  
> > best of my knowledge; the tokens are different than those created using the  
> > English snowball stemming rules). But it does not honor the character  
> > mapping: I would have expected "valter" and not "walter" as the last token  
> > string. And of course, a search for valter won't match walter and this  
> > analysis token issue is likely the root cause:
> > 
> > $ curl -XGET  
> > 'localhost:9200/sgen/\_analyze?analyzer=finnish\_stemming&pretty=true' -d  
> > 'Debby Debbie and Walter' && echo  
> > {  
> > "tokens" : [ {  
> > "token" : "deby",  
> > "start\_offset" : 0,  
> > "end\_offset" : 5,  
> > "type" : "",  
> > "position" : 1  
> > }, {  
> > "token" : "debie",  
> > "start\_offset" : 6,  
> > "end\_offset" : 12,  
> > "type" : "",  
> > "position" : 2  
> > }, {  
> > "token" : "and",  
> > "start\_offset" : 13,  
> > "end\_offset" : 16,  
> > "type" : "",  
> > "position" : 3  
> > }, {  
> > "token" : "walter",  
> > "start\_offset" : 17,  
> > "end\_offset" : 23,  
> > "type" : "",  
> > "position" : 4  
> > } ]  
> > }
> > 
> > I cannot get Elasticsearch to define the analyzers and mappings when  
> > creating an index: There aren't any examples of both, and my  
> > experimentation yields mappings that can only point to configured  
> > analyzers. So configuring a list of analyzers is a currently acceptable  
> > work-around.
> > 
> > But honoring the char\_filter mapping is something that is necessary to  
> > resolve.
> > 
> > Thank you in advance.

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 15, 2013, 6:49pm UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338/4 "2013-01-15T18:49:27Z")

</div>

The stopwords: _none_ part in the finnish\_stemming shouldn't be necessary.  
I intentionally omitted the stopword filter from the filter list. A typical  
definition of a snowball analyzer is this:

my\_stopword\_analyzer:  
type: custom  
tokenizer: standrad  
filter: [standard, lowercase, stop, snowball]

and the stop word filter (stop) has to be configured if you want  
non-default behavior. Since I didn't specify the stop in my analyzer  
definition there is really nothing to configure there.

On Tuesday, January 15, 2013 11:12:54 AM UTC-5, InquiringMind wrote:

> Thank you very much, Igor! I wouldn't have guessed it, but from your  
> example it now makes sense.
> 
> I added the stopwords : _none_ configuration line to the finnish\_stemminganalyzer. It seems to work (that is, it tokenizes stop words) for the few  
> Finnish stop word examples I could find. Though my very limited knowledge  
> of Finnish doesn't let me verify the behavior.
> 
> I also created a counterpart that is non-stemming but still performs  
> character replacements. Here is my current configuration for analyzers:
> 
> index:  
> analysis:  
> char\_filter:  
> finnish\_char\_mapping:  
> type: mapping  
> mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
> filter:  
> finnish\_standard:  
> type: standard  
> language: Finnish  
> finnish\_snowball:  
> type: snowball  
> language: Finnish  
> analyzer:  
> # Default analyzer uses the snowball stemming analyzer with no  
> # stop words, all defined by the default language per the JVM:  
> default:  
> type: snowball  
> stopwords: _none_  
> # Per-language analyzers: No stemming  
> english\_standard:  
> type: standard  
> language: English  
> stopwords: _none_  
> finnish\_standard:  
> type: custom  
> tokenizer: standard  
> filter: [standard, lowercase, finnish\_standard]  
> char\_filter: [finnish\_char\_mapping]  
> stopwords: _none_  
> # Per-language analyzers: Stemming  
> english\_stemming:  
> type: snowball  
> language: English  
> stopwords: _none_  
> finnish\_stemming:  
> type: custom  
> tokenizer: standard  
> filter: [standard, lowercase, finnish\_snowball]  
> char\_filter: [finnish\_char\_mapping]  
> stopwords: _none_
> 
> Thanks again!
> 
> On Monday, January 14, 2013 9:13:19 PM UTC-5, Igor Motov wrote:
> 
> > I don't think you can put a char\_filter on existing analyzer, I think you  
> > need to create a custom analyzer instead:
> > 
> > index:  
> > analysis:  
> > char\_filter:  
> > finnish\_char\_mapping:  
> > type: mapping  
> > mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
> > analyzer:  
> > # Default uses snowball stemming analyzer with no stop words  
> > # with the default language per the JVM:  
> > default:  
> > type: snowball  
> > stopwords: _none_  
> > # Per-language analyzers  
> > english\_standard:  
> > type: standard  
> > language: English  
> > stopwords: _none_  
> > english\_stemming:  
> > type: snowball  
> > language: English  
> > stopwords: _none_  
> > finnish\_stemming:  
> > type: custom  
> > tokenizer: standard  
> > filter: [standard, lowercase, finnish\_snowball]  
> > char\_filter: [finnish\_char\_mapping]  
> > filter:  
> > finnish\_snowball:  
> > type: snowball  
> > language: Finnish
> > 
> > On Monday, January 14, 2013 5:01:44 PM UTC-5, InquiringMind wrote:
> > 
> > > In summary: Everything is currently working, except for char\_filter  
> > > mapping.
> > > 
> > > I'm currently on Elasticsearch 19.10 because it works fine for our  
> > > current production application, and I am not extending its use to require  
> > > any of the bug fixes in the change logs for more recent versions.
> > > 
> > > I've isolated this issue to configured analyzers and a collection of  
> > > HTTP \_analyze requests that easily reproduce the problem. No additional  
> > > data or queries are needed at this point (I don't believe, anyway).
> > > 
> > > Here is the example I found at  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/mapping-charfilter.html)
> > > 
> > > {  
> > > "index" : {  
> > > "analysis" : {  
> > > "char\_filter" : {  
> > > "my\_mapping" : {  
> > > "type" : "mapping",  
> > > "mappings" : ["ph=\>f", "qu=\>q"]  
> > > }  
> > > },  
> > > "analyzer" : {  
> > > "custom\_with\_char\_filter" : {  
> > > "tokenizer" : "standard",  
> > > "char\_filter" : ["my\_mapping"]  
> > > },  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > Here is what my elasticsearch.yml configuration looks like. Note the  
> > > Finnish character mappings that are typical for searching Finnish names:  
> > > The previous example didn't quite work with what I need: A snowball  
> > > stemming tokenizer with Finnish stemming rules, no stop words, and  
> > > convering w to v on the input string before tokenizing. After playing  
> > > around a little, here's what works (except for the char\_filter):
> > > 
> > > index:  
> > > analysis:  
> > > char\_filter:  
> > > finnish\_char\_mapping:  
> > > type: mapping  
> > > mappings: ["Å=\>O", "å=\>o", "W=\>V", "w=\>v"]  
> > > analyzer:  
> > > # Default uses snowball stemming analyzer with no stop words  
> > > # with the default language per the JVM:  
> > > default:  
> > > type: snowball  
> > > stopwords: _none_  
> > > # Per-language analyzers  
> > > english\_standard:  
> > > type: standard  
> > > language: English  
> > > stopwords: _none_  
> > > english\_stemming:  
> > > type: snowball  
> > > language: English  
> > > stopwords: _none_  
> > > finnish\_stemming:  
> > > type: snowball  
> > > language: Finnish  
> > > char\_filter: [finnish\_char\_mapping]  
> > > stopwords: _none_
> > > 
> > > This first analyze operation returns the expected tokens. It analyzes  
> > > the text using the standard analyzer, and stop words are included:
> > > 
> > > $ curl -XGET  
> > > 'localhost:9200/sgen/\_analyze?analyzer=standard&pretty=true' -d 'Debby  
> > > Debbie and Walter' && echo  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "debby",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 5,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "debbie",  
> > > "start\_offset" : 6,  
> > > "end\_offset" : 12,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "walter",  
> > > "start\_offset" : 17,  
> > > "end\_offset" : 23,  
> > > "type" : "",  
> > > "position" : 4  
> > > } ]  
> > > }
> > > 
> > > This also works: It uses the snowball analyzer with the English language  
> > > and with stop words included in the list of tokens as desired:
> > > 
> > > $ curl -XGET  
> > > 'localhost:9200/sgen/\_analyze?analyzer=english\_stemming&pretty=true' -d  
> > > 'Debby Debbie and Walter' && echo  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "debbi",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 5,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "debbi",  
> > > "start\_offset" : 6,  
> > > "end\_offset" : 12,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "and",  
> > > "start\_offset" : 13,  
> > > "end\_offset" : 16,  
> > > "type" : "",  
> > > "position" : 3  
> > > }, {  
> > > "token" : "walter",  
> > > "start\_offset" : 17,  
> > > "end\_offset" : 23,  
> > > "type" : "",  
> > > "position" : 4  
> > > } ]  
> > > }
> > > 
> > > But this doesn't fully work. It uses the Finnish stemming rules (to the  
> > > best of my knowledge; the tokens are different than those created using the  
> > > English snowball stemming rules). But it does not honor the character  
> > > mapping: I would have expected "valter" and not "walter" as the last token  
> > > string. And of course, a search for valter won't match walter and this  
> > > analysis token issue is likely the root cause:
> > > 
> > > $ curl -XGET  
> > > 'localhost:9200/sgen/\_analyze?analyzer=finnish\_stemming&pretty=true' -d  
> > > 'Debby Debbie and Walter' && echo  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "deby",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 5,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "debie",  
> > > "start\_offset" : 6,  
> > > "end\_offset" : 12,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "and",  
> > > "start\_offset" : 13,  
> > > "end\_offset" : 16,  
> > > "type" : "",  
> > > "position" : 3  
> > > }, {  
> > > "token" : "walter",  
> > > "start\_offset" : 17,  
> > > "end\_offset" : 23,  
> > > "type" : "",  
> > > "position" : 4  
> > > } ]  
> > > }
> > > 
> > > I cannot get Elasticsearch to define the analyzers and mappings when  
> > > creating an index: There aren't any examples of both, and my  
> > > experimentation yields mappings that can only point to configured  
> > > analyzers. So configuring a list of analyzers is a currently acceptable  
> > > work-around.
> > > 
> > > But honoring the char\_filter mapping is something that is necessary to  
> > > resolve.
> > > 
> > > Thank you in advance.

--

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [January 15, 2013, 7:59pm UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338/5 "2013-01-15T19:59:59Z")

</div>

Aha! Now it's falling into place. Your last comment was needed (for me,  
anyway) to make the on-line guide clearer. Yes, the stopwords statement is  
not needed because the analyzer is being constructed without the stopfilter.

I went back and re-created my analyzers based on your guidance,  
constructing them out of their individual parts (almost, lowercase is  
really a higher performance combination!), and omitting the stop filter. My  
remaining questions are:

1. Is the _language_ statement (with the language name in lowercase) still  
recommended even for a non-stemming analyzer that doesn't include a stopfilter. For example:

2. When adding the type and field mappings to an index create command, what  
is the JSON format for adding the analyzer definitions? No matter how much  
I play around with the format, the mappings don't seem to be able to find  
any analyzer that I defined in the API; only the pre-configured analyzers  
are seen. This isn't as high priority, but it would be nice to allow the  
most flexibility without the need to pre-configure every supported language.

And thanks again! You've really helped make the existing documentation come  
to life!

On Tuesday, January 15, 2013 1:49:27 PM UTC-5, Igor Motov wrote:

> The stopwords: _none_ part in the finnish\_stemming shouldn't be  
> necessary. I intentionally omitted the stopword filter from the filter  
> list. A typical definition of a snowball analyzer is this:
> 
> my\_stopword\_analyzer:  
> type: custom  
> tokenizer: standrad  
> filter: [standard, lowercase, stop, snowball]
> 
> and the stop word filter (stop) has to be configured if you want  
> non-default behavior. Since I didn't specify the stop in my analyzer  
> definition there is really nothing to configure there.
> 
> > >

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 15, 2013, 11:09pm UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338/6 "2013-01-15T23:09:53Z")

</div>

1. You are configuring a custom analyzer from pieces here. Custom analyzer  
doesn't use the language parameter because it doesn't have any  
language-specific components. It just glue with which you can create  
analyzer out of pieces: char filters, tokenizer and token filters. The  
language parameter should be used when you configure individual pieces such  
as stopword or snowball filters or built-in analyzers such as snowball.

2. I think I already answered this earlier today.

On Tuesday, January 15, 2013 2:59:59 PM UTC-5, InquiringMind wrote:

> Aha! Now it's falling into place. Your last comment was needed (for me,  
> anyway) to make the on-line guide clearer. Yes, the stopwords statement  
> is not needed because the analyzer is being constructed without the stopfilter.
> 
> I went back and re-created my analyzers based on your guidance,  
> constructing them out of their individual parts (almost, lowercase is  
> really a higher performance combination!), and omitting the stop filter. My  
> remaining questions are:
> 
> 1. Is the _language_ statement (with the language name in lowercase)  
> still recommended even for a non-stemming analyzer that doesn't include a  
> stop filter. For example:
> 
> 2. When adding the type and field mappings to an index create command,  
> what is the JSON format for adding the analyzer definitions? No matter how  
> much I play around with the format, the mappings don't seem to be able to  
> find any analyzer that I defined in the API; only the pre-configured  
> analyzers are seen. This isn't as high priority, but it would be nice to  
> allow the most flexibility without the need to pre-configure every  
> supported language.
> 
> And thanks again! You've really helped make the existing documentation  
> come to life!
> 
> On Tuesday, January 15, 2013 1:49:27 PM UTC-5, Igor Motov wrote:
> 
> > The stopwords: _none_ part in the finnish\_stemming shouldn't be  
> > necessary. I intentionally omitted the stopword filter from the filter  
> > list. A typical definition of a snowball analyzer is this:
> > 
> > my\_stopword\_analyzer:  
> > type: custom  
> > tokenizer: standrad  
> > filter: [standard, lowercase, stop, snowball]
> > 
> > and the stop word filter (stop) has to be configured if you want  
> > non-default behavior. Since I didn't specify the stop in my analyzer  
> > definition there is really nothing to configure there.
> > 
> > > >

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:56am UTC](https://discuss.elastic.co/t/elasticsearch-wont-recongize-char-filter-mapping/10338/7 "2017-07-06T02:56:15Z")

</div>


