# Using a char\_filter in combination with a lowercase filter

**URL:** <https://discuss.elastic.co/t/using-a-char-filter-in-combination-with-a-lowercase-filter/19316>\
**Category:** Elasticsearch\
**Created:** [August 18, 2014, 9:14am UTC](https://discuss.elastic.co/t/using-a-char-filter-in-combination-with-a-lowercase-filter/19316 "2014-08-18T09:14:23Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Matthias\_Hogerheijde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matthias_hogerheijde/32/1328_2.png) [@Matthias\_Hogerheijde](https://discuss.elastic.co/u/Matthias_Hogerheijde)\
**Post date:** [August 18, 2014, 9:14am UTC](https://discuss.elastic.co/t/using-a-char-filter-in-combination-with-a-lowercase-filter/19316/1 "2014-08-18T09:14:23Z")

</div>

Hi,

We're using Elasticsearch with an Analyzer to map the `y` character to  
`ij`, (_char\_fitler_ named "char\_mapper") since in Dutch these two are  
"somewhat" interchangeable. We're also using a _lowercase filter_.

This is the configuration:

{  
"analysis": {  
"analyzer": {  
"index": {  
"type": "custom",  
"tokenizer": "standard",  
"filter": [  
"lowercase",  
"synonym\_twoway",  
"standard",  
"asciifolding"  
],  
"char\_filter": [  
"char\_mapper"  
]  
},  
"index\_prefix": {  
"type": "custom",  
"tokenizer": "standard",  
"filter": [  
"lowercase",  
"synonym\_twoway",  
"standard",  
"asciifolding",  
"prefixes"  
],  
"char\_filter": [  
"char\_mapper"  
]  
},  
"search": {  
"alias": [  
"default"  
],  
"type": "custom",  
"tokenizer": "standard",  
"filter": [  
"lowercase",  
"synonym",  
"synonym\_twoway",  
"standard",  
"asciifolding"  
],  
"char\_filter": [  
"char\_mapper"  
]  
},  
"postal\_code": {  
"tokenizer": "keyword",  
"filter": [  
"lowercase"  
]  
}  
},  
"tokenizer": {  
"standard": {  
"stopwords": [

```
    ]
  }
},
"filter": {
  "synonym": {
    "type": "synonym",
    "synonyms": [
      "st => sint",
      "jp => jan pieterszoon",
      "mh => maarten harpertszoon"
    ]
  },
  "synonym_twoway": {
    "type": "synonym",
    "synonyms": [
      "den haag, s gravenhage",
      "den bosch, s hertogenbosch"
    ]
  },
  "prefixes": {
    "type": "edgeNGram",
    "side": "front",
    "min_gram": 1,
    "max_gram": 30
  }
},
"char_filter": {
  "char_mapper": {
    "type": "mapping",
    "mappings": [
      "y => ij"
    ]
  }
}

```

}  
}

When indexing cities, we're using this mapping:

{  
"properties": {  
"city": {  
"type": "multi\_field",  
"fields": {  
"city": {  
"type": "string"  
},  
"prefix": {  
"type": "string",  
"boost": 0.5,  
"index\_analyzer": "index\_prefix"  
}  
}  
},  
"province\_code": {  
"type": "string"  
},  
"unique\_name": {  
"type": "boolean"  
},  
"point": {  
"type": "geo\_point"  
},  
"search\_terms": {  
"type": "multi\_field",  
"fields": {  
"search\_terms": {  
"type": "string"  
},  
"prefix": {  
"boost": 0.5,  
"index\_analyzer": "index\_prefix",  
"type": "string"  
}  
}  
}  
},  
"search\_analyzer": "search",  
"index\_analyzer": "index"  
}

When we index all the (Dutch) cities from our data-source, there are cities  
starting with both `IJ` and `Y`. (for example, these citiy names exist:  
_IJssel_, _IJsselstein_, _Yerseke_ and _Ysselsteyn._) It seems that these  
characters are not lowercased before the char\_mapping is applied.

Querying the index, results in

/top/city/\_search?q=ijsselstein -\> works, returns the document for  
IJsselstein  
/top/city/\_search?q=Ijsselstein -\> works, returns the document for  
IJsselstein  
/top/city/\_search?q=yerseke -\> \*doesn't \*work, returns nothing  
/top/city/\_search?q=Yerseke -\> \*does \*work, returns the document for Yerseke  
/top/city/\_search?q=YsselsteYn -\> \*doesn't \*work, returns nothing  
/top/city/\_search?q=Ysselsteyn -\> \*does \*work, returns the document for  
Ysselsteyn

Changing the case of any other letter doesn't affect the results.

I've worked around this issue by adding the mapping "Y =\> ij", i.e.:

"char\_filter": {  
"char\_mapper": {  
"type": "mapping",  
"mappings": [  
"y =\> ij",  
"Y =\> ij"  
]  
}  
}

This solves the problem, but I'd rather see that the lowercase filter is  
applied before the mapping, or, that I can make the order explicit. Is  
there any stance on this issue? Or is this intended behaviour?

Regards,  
Matthias Hogerheijde

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 19, 2014, 4:37am UTC](https://discuss.elastic.co/t/using-a-char-filter-in-combination-with-a-lowercase-filter/19316/2 "2014-08-19T04:37:16Z")

</div>

Char filters are applied before the text is tokenized, and therefore they  
are applied before the "normal" filters are used, which is why they are a  
separate class of filter. With Lucene, the order is:

char filters -\> tokenizer -\> filters

Have you looked into the ICU analyzer?

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

I have no idea how well it works with Dutch.

Cheers,

Ivan

On Mon, Aug 18, 2014 at 2:14 AM, Matthias Hogerheijde \<  
[matthias.hogerheijde@goabout.com](mailto:matthias.hogerheijde@goabout.com)\> wrote:

> Hi,
> 
> We're using Elasticsearch with an Analyzer to map the `y` character to  
> `ij`, (_char\_fitler_ named "char\_mapper") since in Dutch these two are  
> "somewhat" interchangeable. We're also using a _lowercase filter_.
> 
> This is the configuration:
> 
> {  
> "analysis": {  
> "analyzer": {  
> "index": {  
> "type": "custom",  
> "tokenizer": "standard",  
> "filter": [  
> "lowercase",  
> "synonym\_twoway",  
> "standard",  
> "asciifolding"  
> ],  
> "char\_filter": [  
> "char\_mapper"  
> ]  
> },  
> "index\_prefix": {  
> "type": "custom",  
> "tokenizer": "standard",  
> "filter": [  
> "lowercase",  
> "synonym\_twoway",  
> "standard",  
> "asciifolding",  
> "prefixes"  
> ],  
> "char\_filter": [  
> "char\_mapper"  
> ]  
> },  
> "search": {  
> "alias": [  
> "default"  
> ],  
> "type": "custom",  
> "tokenizer": "standard",  
> "filter": [  
> "lowercase",  
> "synonym",  
> "synonym\_twoway",  
> "standard",  
> "asciifolding"  
> ],  
> "char\_filter": [  
> "char\_mapper"  
> ]  
> },  
> "postal\_code": {  
> "tokenizer": "keyword",  
> "filter": [  
> "lowercase"  
> ]  
> }  
> },  
> "tokenizer": {  
> "standard": {  
> "stopwords": [
> 
> ```
> ]
> }
> },
> "filter": {
> "synonym": {
> "type": "synonym",
> "synonyms": [
> "st => sint",
> "jp => jan pieterszoon",
> "mh => maarten harpertszoon"
> ]
> },
> "synonym_twoway": {
> "type": "synonym",
> "synonyms": [
> "den haag, s gravenhage",
> "den bosch, s hertogenbosch"
> ]
> },
> "prefixes": {
> "type": "edgeNGram",
> "side": "front",
> "min_gram": 1,
> "max_gram": 30
> }
> },
> "char_filter": {
> "char_mapper": {
> "type": "mapping",
> "mappings": [
> "y => ij"
> ]
> }
> }
> 
> ```
> 
> }  
> }
> 
> When indexing cities, we're using this mapping:
> 
> {  
> "properties": {  
> "city": {  
> "type": "multi\_field",  
> "fields": {  
> "city": {  
> "type": "string"  
> },  
> "prefix": {  
> "type": "string",  
> "boost": 0.5,  
> "index\_analyzer": "index\_prefix"  
> }  
> }  
> },  
> "province\_code": {  
> "type": "string"  
> },  
> "unique\_name": {  
> "type": "boolean"  
> },  
> "point": {  
> "type": "geo\_point"  
> },  
> "search\_terms": {  
> "type": "multi\_field",  
> "fields": {  
> "search\_terms": {  
> "type": "string"  
> },  
> "prefix": {  
> "boost": 0.5,  
> "index\_analyzer": "index\_prefix",  
> "type": "string"  
> }  
> }  
> }  
> },  
> "search\_analyzer": "search",  
> "index\_analyzer": "index"  
> }
> 
> When we index all the (Dutch) cities from our data-source, there are  
> cities starting with both `IJ` and `Y`. (for example, these citiy names  
> exist: _IJssel_, _IJsselstein_, _Yerseke_ and _Ysselsteyn._) It seems  
> that these characters are not lowercased before the char\_mapping is  
> applied.
> 
> Querying the index, results in
> 
> /top/city/\_search?q=ijsselstein -\> works, returns the document for  
> IJsselstein  
> /top/city/\_search?q=Ijsselstein -\> works, returns the document for  
> IJsselstein  
> /top/city/\_search?q=yerseke -\> \*doesn't \*work, returns nothing  
> /top/city/\_search?q=Yerseke -\> \*does \*work, returns the document for  
> Yerseke  
> /top/city/\_search?q=YsselsteYn -\> \*doesn't \*work, returns nothing  
> /top/city/\_search?q=Ysselsteyn -\> \*does \*work, returns the document for  
> Ysselsteyn
> 
> Changing the case of any other letter doesn't affect the results.
> 
> I've worked around this issue by adding the mapping "Y =\> ij", i.e.:
> 
> "char\_filter": {  
> "char\_mapper": {  
> "type": "mapping",  
> "mappings": [  
> "y =\> ij",  
> "Y =\> ij"  
> ]  
> }  
> }
> 
> This solves the problem, but I'd rather see that the lowercase filter is  
> applied before the mapping, or, that I can make the order explicit. Is  
> there any stance on this issue? Or is this intended behaviour?
> 
> Regards,  
> Matthias Hogerheijde
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQAzTpAxXiZtkpXh3JLga%3DmvX3MThcsFV-2YPOXDBWSphg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQAzTpAxXiZtkpXh3JLga%3DmvX3MThcsFV-2YPOXDBWSphg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Matthias\_Hogerheijde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matthias_hogerheijde/32/1328_2.png) [@Matthias\_Hogerheijde](https://discuss.elastic.co/u/Matthias_Hogerheijde)\
**Post date:** [August 19, 2014, 7:46am UTC](https://discuss.elastic.co/t/using-a-char-filter-in-combination-with-a-lowercase-filter/19316/3 "2014-08-19T07:46:24Z")

</div>

Thanks for your reply. I see that I didn't fully understand that  
CharFilters are ran first, which makes it logical to special-case the  
different cases. I was originally thrown off-scent that searching with an  
uppercase 'Y' worked and thought that the lowercase filter was not applied  
to the 'Y', but now I see that searching for a 'y' will cause the mapper to  
search for 'ij' in stead.

I don't understand the full extend of the icu analysers, but it seems to me  
that in our case this is semantically different, since we regard 'Y' and  
'IJ' as different letters? (note that we actually regard 'ij' to be a  
single character.) It's not like removing the accents from 'ä', or  
transcribing a Cyrillic number into it's Roman equivalent, or am I wrong to  
that regard?

Regards,  
Matthias

On Tuesday, August 19, 2014 6:37:29 AM UTC+2, Ivan Brusic wrote:

> Char filters are applied before the text is tokenized, and therefore they  
> are applied before the "normal" filters are used, which is why they are a  
> separate class of filter. With Lucene, the order is:
> 
> char filters -\> tokenizer -\> filters
> 
> Have you looked into the ICU analyzer?  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/analysis-icu-plugin.html)  
> [http://www.google.com/url?q=http%3A%2F%2Fwww.elasticsearch.org%2Fguide%2Fen%2Felasticsearch%2Freference%2Fcurrent%2Fanalysis-icu-plugin.html&sa=D&sntz=1&usg=AFQjCNGvdkiBOpv0quMGWpUHS15nSr8aug](http://www.google.com/url?q=http%3A%2F%2Fwww.elasticsearch.org%2Fguide%2Fen%2Felasticsearch%2Freference%2Fcurrent%2Fanalysis-icu-plugin.html&sa=D&sntz=1&usg=AFQjCNGvdkiBOpv0quMGWpUHS15nSr8aug)
> 
> I have no idea how well it works with Dutch.
> 
> Cheers,
> 
> Ivan
> 
> On Mon, Aug 18, 2014 at 2:14 AM, Matthias Hogerheijde \<  
> [matthias.h...@goabout.com](mailto:matthias.h...@goabout.com) \<javascript:\>\> wrote:
> 
> > Hi,
> > 
> > We're using Elasticsearch with an Analyzer to map the `y` character to  
> > `ij`, (_char\_fitler_ named "char\_mapper") since in Dutch these two are  
> > "somewhat" interchangeable. We're also using a _lowercase filter_.
> > 
> > This is the configuration:
> > 
> > {  
> > "analysis": {  
> > "analyzer": {  
> > "index": {  
> > "type": "custom",  
> > "tokenizer": "standard",  
> > "filter": [  
> > "lowercase",  
> > "synonym\_twoway",  
> > "standard",  
> > "asciifolding"  
> > ],  
> > "char\_filter": [  
> > "char\_mapper"  
> > ]  
> > },  
> > "index\_prefix": {  
> > "type": "custom",  
> > "tokenizer": "standard",  
> > "filter": [  
> > "lowercase",  
> > "synonym\_twoway",  
> > "standard",  
> > "asciifolding",  
> > "prefixes"  
> > ],  
> > "char\_filter": [  
> > "char\_mapper"  
> > ]  
> > },  
> > "search": {  
> > "alias": [  
> > "default"  
> > ],  
> > "type": "custom",  
> > "tokenizer": "standard",  
> > "filter": [  
> > "lowercase",  
> > "synonym",  
> > "synonym\_twoway",  
> > "standard",  
> > "asciifolding"  
> > ],  
> > "char\_filter": [  
> > "char\_mapper"  
> > ]  
> > },  
> > "postal\_code": {  
> > "tokenizer": "keyword",  
> > "filter": [  
> > "lowercase"  
> > ]  
> > }  
> > },  
> > "tokenizer": {  
> > "standard": {  
> > "stopwords": [
> > 
> > ```
> > ]
> > }
> > },
> > "filter": {
> > "synonym": {
> > "type": "synonym",
> > "synonyms": [
> > "st => sint",
> > "jp => jan pieterszoon",
> > "mh => maarten harpertszoon"
> > ]
> > },
> > "synonym_twoway": {
> > "type": "synonym",
> > "synonyms": [
> > "den haag, s gravenhage",
> > "den bosch, s hertogenbosch"
> > ]
> > },
> > "prefixes": {
> > "type": "edgeNGram",
> > "side": "front",
> > "min_gram": 1,
> > "max_gram": 30
> > }
> > },
> > "char_filter": {
> > "char_mapper": {
> > "type": "mapping",
> > "mappings": [
> > "y => ij"
> > ]
> > }
> > }
> > 
> > ```
> > 
> > }  
> > }
> > 
> > When indexing cities, we're using this mapping:
> > 
> > {  
> > "properties": {  
> > "city": {  
> > "type": "multi\_field",  
> > "fields": {  
> > "city": {  
> > "type": "string"  
> > },  
> > "prefix": {  
> > "type": "string",  
> > "boost": 0.5,  
> > "index\_analyzer": "index\_prefix"  
> > }  
> > }  
> > },  
> > "province\_code": {  
> > "type": "string"  
> > },  
> > "unique\_name": {  
> > "type": "boolean"  
> > },  
> > "point": {  
> > "type": "geo\_point"  
> > },  
> > "search\_terms": {  
> > "type": "multi\_field",  
> > "fields": {  
> > "search\_terms": {  
> > "type": "string"  
> > },  
> > "prefix": {  
> > "boost": 0.5,  
> > "index\_analyzer": "index\_prefix",  
> > "type": "string"  
> > }  
> > }  
> > }  
> > },  
> > "search\_analyzer": "search",  
> > "index\_analyzer": "index"  
> > }
> > 
> > When we index all the (Dutch) cities from our data-source, there are  
> > cities starting with both `IJ` and `Y`. (for example, these citiy names  
> > exist: _IJssel_, _IJsselstein_, _Yerseke_ and _Ysselsteyn._) It seems  
> > that these characters are not lowercased before the char\_mapping is  
> > applied.
> > 
> > Querying the index, results in
> > 
> > /top/city/\_search?q=ijsselstein -\> works, returns the document for  
> > IJsselstein  
> > /top/city/\_search?q=Ijsselstein -\> works, returns the document for  
> > IJsselstein  
> > /top/city/\_search?q=yerseke -\> \*doesn't \*work, returns nothing  
> > /top/city/\_search?q=Yerseke -\> \*does \*work, returns the document for  
> > Yerseke  
> > /top/city/\_search?q=YsselsteYn -\> \*doesn't \*work, returns nothing  
> > /top/city/\_search?q=Ysselsteyn -\> \*does \*work, returns the document for  
> > Ysselsteyn
> > 
> > Changing the case of any other letter doesn't affect the results.
> > 
> > I've worked around this issue by adding the mapping "Y =\> ij", i.e.:
> > 
> > "char\_filter": {  
> > "char\_mapper": {  
> > "type": "mapping",  
> > "mappings": [  
> > "y =\> ij",  
> > "Y =\> ij"  
> > ]  
> > }  
> > }
> > 
> > This solves the problem, but I'd rather see that the lowercase filter is  
> > applied before the mapping, or, that I can make the order explicit. Is  
> > there any stance on this issue? Or is this intended behaviour?
> > 
> > Regards,  
> > Matthias Hogerheijde
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/e18b3d66-0cec-49ae-9bea-af699ce5a97c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e18b3d66-0cec-49ae-9bea-af699ce5a97c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [August 19, 2014, 4:36pm UTC](https://discuss.elastic.co/t/using-a-char-filter-in-combination-with-a-lowercase-filter/19316/4 "2014-08-19T16:36:50Z")

</div>

The plugin uses collation to identify characters which are equivalent. It  
does far more than simple replacement/folding, so sometimes the sort order  
matters.

> **[Collation](https://en.wikipedia.org/wiki/Collation)**
>
> Collation is the assembly of written information into a standard order. Many systems of collation are based on numerical order or alphabetical order, or extensions and combinations thereof. Collation is a fundamental element of most office filing systems, library catalogs, and reference books.
> Collation differs from classification in that the classes themselves are not necessarily ordered. However, even if the order of the classes is irrelevant, the identifiers of the classes may be members of a...

> **[ICU User Guide](https://unicode-org.github.io/icu/userguide/)**
>
> ICU User Guide

Take a look at the plugin's test to figure out how it is used. I only work  
with English/Mandarin, so I do not know how useful it is with Dutch.

[https://github.com/elasticsearch/elasticsearch-analysis-icu/tree/master/src/test/java/org/elasticsearch/index/analysis](https://github.com/elasticsearch/elasticsearch-analysis-icu/tree/master/src/test/java/org/elasticsearch/index/analysis)

Cheers,

Ivan

On Tue, Aug 19, 2014 at 12:46 AM, Matthias Hogerheijde \<  
[matthias.hogerheijde@goabout.com](mailto:matthias.hogerheijde@goabout.com)\> wrote:

> Thanks for your reply. I see that I didn't fully understand that  
> CharFilters are ran first, which makes it logical to special-case the  
> different cases. I was originally thrown off-scent that searching with an  
> uppercase 'Y' worked and thought that the lowercase filter was not applied  
> to the 'Y', but now I see that searching for a 'y' will cause the mapper to  
> search for 'ij' in stead.
> 
> I don't understand the full extend of the icu analysers, but it seems to  
> me that in our case this is semantically different, since we regard 'Y' and  
> 'IJ' as different letters? (note that we actually regard 'ij' to be a  
> single character.) It's not like removing the accents from 'ä', or  
> transcribing a Cyrillic number into it's Roman equivalent, or am I wrong to  
> that regard?
> 
> Regards,  
> Matthias
> 
> On Tuesday, August 19, 2014 6:37:29 AM UTC+2, Ivan Brusic wrote:
> 
> > Char filters are applied before the text is tokenized, and therefore they  
> > are applied before the "normal" filters are used, which is why they are a  
> > separate class of filter. With Lucene, the order is:
> > 
> > char filters -\> tokenizer -\> filters
> > 
> > Have you looked into the ICU analyzer? [http://www](http://www).  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://elasticsearch.org/guide/en/elasticsearch/reference/)  
> > current/analysis-icu-plugin.html  
> > [http://www.google.com/url?q=http%3A%2F%2Fwww.elasticsearch.org%2Fguide%2Fen%2Felasticsearch%2Freference%2Fcurrent%2Fanalysis-icu-plugin.html&sa=D&sntz=1&usg=AFQjCNGvdkiBOpv0quMGWpUHS15nSr8aug](http://www.google.com/url?q=http%3A%2F%2Fwww.elasticsearch.org%2Fguide%2Fen%2Felasticsearch%2Freference%2Fcurrent%2Fanalysis-icu-plugin.html&sa=D&sntz=1&usg=AFQjCNGvdkiBOpv0quMGWpUHS15nSr8aug)
> > 
> > I have no idea how well it works with Dutch.
> > 
> > Cheers,
> > 
> > Ivan
> > 
> > On Mon, Aug 18, 2014 at 2:14 AM, Matthias Hogerheijde \<  
> > [matthias.h...@goabout.com](mailto:matthias.h...@goabout.com)\> wrote:
> > 
> > > Hi,
> > > 
> > > We're using Elasticsearch with an Analyzer to map the `y` character to  
> > > `ij`, (_char\_fitler_ named "char\_mapper") since in Dutch these two are  
> > > "somewhat" interchangeable. We're also using a _lowercase filter_.
> > > 
> > > This is the configuration:
> > > 
> > > {  
> > > "analysis": {  
> > > "analyzer": {  
> > > "index": {  
> > > "type": "custom",  
> > > "tokenizer": "standard",  
> > > "filter": [  
> > > "lowercase",  
> > > "synonym\_twoway",  
> > > "standard",  
> > > "asciifolding"  
> > > ],  
> > > "char\_filter": [  
> > > "char\_mapper"  
> > > ]  
> > > },  
> > > "index\_prefix": {  
> > > "type": "custom",  
> > > "tokenizer": "standard",  
> > > "filter": [  
> > > "lowercase",  
> > > "synonym\_twoway",  
> > > "standard",  
> > > "asciifolding",  
> > > "prefixes"  
> > > ],  
> > > "char\_filter": [  
> > > "char\_mapper"  
> > > ]  
> > > },  
> > > "search": {  
> > > "alias": [  
> > > "default"  
> > > ],  
> > > "type": "custom",  
> > > "tokenizer": "standard",  
> > > "filter": [  
> > > "lowercase",  
> > > "synonym",  
> > > "synonym\_twoway",  
> > > "standard",  
> > > "asciifolding"  
> > > ],  
> > > "char\_filter": [  
> > > "char\_mapper"  
> > > ]  
> > > },  
> > > "postal\_code": {  
> > > "tokenizer": "keyword",  
> > > "filter": [  
> > > "lowercase"  
> > > ]  
> > > }  
> > > },  
> > > "tokenizer": {  
> > > "standard": {  
> > > "stopwords": [
> > > 
> > > ```
> > > ]
> > > }
> > > },
> > > "filter": {
> > > "synonym": {
> > > "type": "synonym",
> > > "synonyms": [
> > > "st => sint",
> > > "jp => jan pieterszoon",
> > > "mh => maarten harpertszoon"
> > > ]
> > > },
> > > "synonym_twoway": {
> > > "type": "synonym",
> > > "synonyms": [
> > > "den haag, s gravenhage",
> > > "den bosch, s hertogenbosch"
> > > ]
> > > },
> > > "prefixes": {
> > > "type": "edgeNGram",
> > > "side": "front",
> > > "min_gram": 1,
> > > "max_gram": 30
> > > }
> > > },
> > > "char_filter": {
> > > "char_mapper": {
> > > "type": "mapping",
> > > "mappings": [
> > > "y => ij"
> > > ]
> > > }
> > > }
> > > 
> > > ```
> > > 
> > > }  
> > > }
> > > 
> > > When indexing cities, we're using this mapping:
> > > 
> > > {  
> > > "properties": {  
> > > "city": {  
> > > "type": "multi\_field",  
> > > "fields": {  
> > > "city": {  
> > > "type": "string"  
> > > },  
> > > "prefix": {  
> > > "type": "string",  
> > > "boost": 0.5,  
> > > "index\_analyzer": "index\_prefix"  
> > > }  
> > > }  
> > > },  
> > > "province\_code": {  
> > > "type": "string"  
> > > },  
> > > "unique\_name": {  
> > > "type": "boolean"  
> > > },  
> > > "point": {  
> > > "type": "geo\_point"  
> > > },  
> > > "search\_terms": {  
> > > "type": "multi\_field",  
> > > "fields": {  
> > > "search\_terms": {  
> > > "type": "string"  
> > > },  
> > > "prefix": {  
> > > "boost": 0.5,  
> > > "index\_analyzer": "index\_prefix",  
> > > "type": "string"  
> > > }  
> > > }  
> > > }  
> > > },  
> > > "search\_analyzer": "search",  
> > > "index\_analyzer": "index"  
> > > }
> > > 
> > > When we index all the (Dutch) cities from our data-source, there are  
> > > cities starting with both `IJ` and `Y`. (for example, these citiy names  
> > > exist: _IJssel_, _IJsselstein_, _Yerseke_ and _Ysselsteyn._) It seems  
> > > that these characters are not lowercased before the char\_mapping is  
> > > applied.
> > > 
> > > Querying the index, results in
> > > 
> > > /top/city/\_search?q=ijsselstein -\> works, returns the document for  
> > > IJsselstein  
> > > /top/city/\_search?q=Ijsselstein -\> works, returns the document for  
> > > IJsselstein  
> > > /top/city/\_search?q=yerseke -\> \*doesn't \*work, returns nothing  
> > > /top/city/\_search?q=Yerseke -\> \*does \*work, returns the document for  
> > > Yerseke  
> > > /top/city/\_search?q=YsselsteYn -\> \*doesn't \*work, returns nothing  
> > > /top/city/\_search?q=Ysselsteyn -\> \*does \*work, returns the document for  
> > > Ysselsteyn
> > > 
> > > Changing the case of any other letter doesn't affect the results.
> > > 
> > > I've worked around this issue by adding the mapping "Y =\> ij", i.e.:
> > > 
> > > "char\_filter": {  
> > > "char\_mapper": {  
> > > "type": "mapping",  
> > > "mappings": [  
> > > "y =\> ij",  
> > > "Y =\> ij"  
> > > ]  
> > > }  
> > > }
> > > 
> > > This solves the problem, but I'd rather see that the lowercase filter is  
> > > applied before the mapping, or, that I can make the order explicit. Is  
> > > there any stance on this issue? Or is this intended behaviour?
> > > 
> > > Regards,  
> > > Matthias Hogerheijde
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).
> > > 
> > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%  
> > > [40googlegroups.com](http://40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/c60de452-2a3f-42f7-a677-956f81ecec17%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .  
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/e18b3d66-0cec-49ae-9bea-af699ce5a97c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e18b3d66-0cec-49ae-9bea-af699ce5a97c%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/e18b3d66-0cec-49ae-9bea-af699ce5a97c%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/e18b3d66-0cec-49ae-9bea-af699ce5a97c%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCjkqWBqs8u8QyGCtZ7UBZjPA346j2uMbZM8wpXKha1OA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCjkqWBqs8u8QyGCtZ7UBZjPA346j2uMbZM8wpXKha1OA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:07am UTC](https://discuss.elastic.co/t/using-a-char-filter-in-combination-with-a-lowercase-filter/19316/5 "2017-07-06T01:07:40Z")

</div>


