# Custom normalisation and filtering?

**URL:** <https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882>\
**Category:** Elasticsearch\
**Created:** [February 4, 2011, 12:53pm UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882 "2011-02-04T12:53:33Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Barsk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/barsk/32/13734_2.png) [@Barsk](https://discuss.elastic.co/u/Barsk)\
**Post date:** [February 4, 2011, 12:53pm UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/1 "2011-02-04T12:53:33Z")

</div>

```
I have spent the last day or so trying to get my head around how
normalization is done in ElasticSearch and how to customize it for
my needs.

I am responsible for a webbappliction that indexes library
catalogues. i.e card catalogues that all libraries had prior to the
digitized era. These catalogues may be really old, like spanning
from 1600-1974 and are often sorted according to some specific
rules. For instance all accents should be removed, but not for those
letters that are part of our alphabet in Sweden åäö, ÅÄÖ. For all
the rest the accents are removed e.g é=e etc. Also, some catalogues
have some special rules such as v=w, i=j etc . All my indexes are
ISO-8859-1.

In my webapp I have made my own normalization handling based on
these rules and I store the index in an SQL database.

All fine.

But now we are going to OCR process all those cards which we have
scanned already and create a free text search on <b>all</b> the
text on the cards, not just the main entry that the card is sorted
under (author or title). So I am looking at Elastic Search to help
me with this, and the features so far is awesome. I aim to replace
the search engine in my webapp with elastic search.

However the analyzer/normalization part raises some questions.

1) How do I create a custom analyzer that has a filter that removes
the accents according to these rules? Is there an API to build upon?
What I need to do is close to the ISOLatin1AccentFilter in Lucene,
but with some customisation.

2) Filter according to specific rules, e.g v=w, i=j etc

3) Nordic stemming (swedish, norwegian, finnish), seems not to be
available. It is a part of the Snowball classes that I saw is about
to be introduced in 0.15, but only German, English and Dutch was
supported there. How do I go about to add Swedish stemming in ES?

ICU is also an option, the docs on their homepage is far from light
though. But it seems they have normalization features that are
configurable. However the icu-plugin only handles the default
formats and no custom. I think tough that ICU handling is more than
I need really.
```

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [February 4, 2011, 2:40pm UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/2 "2011-02-04T14:40:24Z")

</div>

Hi,

as far as I can understand it should be possible to implement your ES plugin  
with whatever filters and analyzers you need if they are missing now.  
Language analyzers that are now present in ES are based on Lucene 3.0.3.  
Note that analysis module has undergone significant development since then  
in Lucene project, now you can find all "sv", "fi" and "no" analyzers in  
Lucene 3.1-dev version (in trunk) and it is not hard to backport those new  
analyzers (and stemmers) from 3.1 into ES. I have done this for Czech  
Stemmer some time ago (check here for inspiration: my original pull  
request[https://github.com/lukas-vlcek/elasticsearch/commit/0c0e43db76d0fdcfc7c711e1b0a210b1aba71b09](https://github.com/lukas-vlcek/elasticsearch/commit/0c0e43db76d0fdcfc7c711e1b0a210b1aba71b09)and  
here for Shay's  
cleanup[https://github.com/elasticsearch/elasticsearch/commit/034a66263a345c29d1efe1a7fb4c25e2e0f2fb4d](https://github.com/elasticsearch/elasticsearch/commit/034a66263a345c29d1efe1a7fb4c25e2e0f2fb4d)).  
The same could be done with other analyzers as well (though not very nice  
practice it is probably better then switching to 3.1-dev version of Lucene  
now).

If you need some customized analyzer or filter then it might be better idea  
to implement ES plugin for it. In this case you can try to look at some of  
the existing plugins (ICU could be a good candidate?) and start from there.

Regards,  
Lukas

On Fri, Feb 4, 2011 at 1:53 PM, Kristian Jörg [krjg@devo.se](mailto:krjg@devo.se) wrote:

> I have spent the last day or so trying to get my head around how  
> normalization is done in Elasticsearch and how to customize it for my needs.
> 
> I am responsible for a webbappliction that indexes library catalogues. i.e  
> card catalogues that all libraries had prior to the digitized era. These  
> catalogues may be really old, like spanning from 1600-1974 and are often  
> sorted according to some specific rules. For instance all accents should be  
> removed, but not for those letters that are part of our alphabet in Sweden  
> åäö, ÅÄÖ. For all the rest the accents are removed e.g é=e etc. Also, some  
> catalogues have some special rules such as v=w, i=j etc . All my indexes are  
> ISO-8859-1.
> 
> In my webapp I have made my own normalization handling based on these rules  
> and I store the index in an SQL database.  
> All fine.
> 
> But now we are going to OCR process all those cards which we have scanned  
> already and create a free text search on _all_ the text on the cards, not  
> just the main entry that the card is sorted under (author or title). So I am  
> looking at Elastic Search to help me with this, and the features so far is  
> awesome. I aim to replace the search engine in my webapp with elastic  
> search.  
> However the analyzer/normalization part raises some questions.
> 
> 1. How do I create a custom analyzer that has a filter that removes the  
> accents according to these rules? Is there an API to build upon? What I need  
> to do is close to the ISOLatin1AccentFilter in Lucene, but with some  
> customisation.
> 2. Filter according to specific rules, e.g v=w, i=j etc
> 3. Nordic stemming (swedish, norwegian, finnish), seems not to be  
> available. It is a part of the Snowball classes that I saw is about to be  
> introduced in 0.15, but only German, English and Dutch was supported there.  
> How do I go about to add Swedish stemming in ES?
> 
> ICU is also an option, the docs on their homepage is far from light though.  
> But it seems they have normalization features that are configurable. However  
> the icu-plugin only handles the default formats and no custom. I think tough  
> that ICU handling is more than I need really.

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [February 4, 2011, 8:39pm UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/3 "2011-02-04T20:39:23Z")

</div>

On Fri, Feb 4, 2011 at 7:53 AM, Kristian Jörg [krjg@devo.se](mailto:krjg@devo.se) wrote:

> I have spent the last day or so trying to get my head around how  
> normalization is done in Elasticsearch and how to customize it for my needs.
> 
> I am responsible for a webbappliction that indexes library catalogues. i.e  
> card catalogues that all libraries had prior to the digitized era. These  
> catalogues may be really old, like spanning from 1600-1974 and are often  
> sorted according to some specific rules. For instance all accents should be  
> removed, but not for those letters that are part of our alphabet in Sweden  
> åäö, ÅÄÖ. For all the rest the accents are removed e.g é=e etc. Also, some  
> catalogues have some special rules such as v=w, i=j etc . All my indexes are  
> ISO-8859-1.

Hello: there are a number of ways you can do this in ICU: collation,  
normalization, and transliteration.

But if your goal is to achieve correct sort order for a sort field,  
and not for search, I would recommend using collation.  
[http://lucene.apache.org/java/3\_0\_3/api/contrib-collation/index.html](http://lucene.apache.org/java/3_0_3/api/contrib-collation/index.html)

In particular, I would use the ICU variants here, you get support for  
many more locales, smaller sort keys, and faster indexing performance.

The collation filters here will normalize your text into a 'collation  
key' at index time for sorting, so that at runtime, you just sort on  
the field in binary order and results come back in language-sensitive  
order, just like how this is often done in databases.

---

<div class="post-metadata">

**Author:** ![Barsk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/barsk/32/13734_2.png) [@Barsk](https://discuss.elastic.co/u/Barsk)\
**Post date:** [February 7, 2011, 7:31am UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/4 "2011-02-07T07:31:17Z")

</div>

Robert Muir skrev 2011-02-04 21:39:

> On Fri, Feb 4, 2011 at 7:53 AM, Kristian JÃ¶rg[krjg@devo.se](mailto:krjg@devo.se) wrote:
> 
> > I have spent the last day or so trying to get my head around how  
> > normalization is done in Elasticsearch and how to customize it for my needs.
> > 
> > I am responsible for a webbappliction that indexes library catalogues. i.e  
> > card catalogues that all libraries had prior to the digitized era. These  
> > catalogues may be really old, like spanning from 1600-1974 and are often  
> > sorted according to some specific rules. For instance all accents should be  
> > removed, but not for those letters that are part of our alphabet in Sweden  
> > Ã¥Ã¤Ã¶, ÃÃÃ. For all the rest the accents are removed e.g Ã©=e etc. Also, some  
> > catalogues have some special rules such as v=w, i=j etc . All my indexes are  
> > ISO-8859-1.  
> > Hello: there are a number of ways you can do this in ICU: collation,  
> > normalization, and transliteration.
> 
> But if your goal is to achieve correct sort order for a sort field,  
> and not for search, I would recommend using collation.  
> [Lucene 3.0.3 API](http://lucene.apache.org/java/3_0_3/api/contrib-collation/index.html)
> 
> In particular, I would use the ICU variants here, you get support for  
> many more locales, smaller sort keys, and faster indexing performance.
> 
> The collation filters here will normalize your text into a 'collation  
> key' at index time for sorting, so that at runtime, you just sort on  
> the field in binary order and results come back in language-sensitive  
> order, just like how this is often done in databases.  
> Yes, I guess the ICU way is the correct one if we get a broader scope  
> for the library catalogues. For now we are only focusing on the nordic  
> countries and Sweden i particular so ISO-8859-1 handling is enough. If  
> we internationalize fully, UTF-8 and the ICU support seems like a  
> perfect way.  
> However right now the product is limited to ISO-8859-1 in most other  
> respects so having just one part (free text search) being full UTF-8  
> compliant is of limited use.

My question in particular was not which normalisation/collation package  
to use, it was HOW to get them into ES. The current support seems  
limited and poorly documented. It look like I have to grab the full  
source and hack around? There should be a better way to "plug in"  
whatever lucene or solr you may need. At what I have grasped so far from  
the source most or all of the analyzers and filters are pure lucene  
stuff with some wrapper code on them. Could'nt this be done dynamically  
in runtime, perhaps with the help of reflection etc?

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [February 7, 2011, 9:00am UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/5 "2011-02-07T09:00:38Z")

</div>

On Mon, Feb 7, 2011 at 2:31 AM, Kristian Jörg [krjg@devo.se](mailto:krjg@devo.se) wrote:

> However right now the product is limited to ISO-8859-1 in most other  
> respects so having just one part (free text search) being full UTF-8  
> compliant is of limited use.

I don't understand what you are saying here. All lucene indexes are  
UTF-8, that includes yours too.

---

<div class="post-metadata">

**Author:** ![Barsk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/barsk/32/13734_2.png) [@Barsk](https://discuss.elastic.co/u/Barsk)\
**Post date:** [February 7, 2011, 1:10pm UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/6 "2011-02-07T13:10:19Z")

</div>

Robert Muir skrev 2011-02-07 10:00:

> On Mon, Feb 7, 2011 at 2:31 AM, Kristian JÃ¶rg[krjg@devo.se](mailto:krjg@devo.se) wrote:
> 
> > However right now the product is limited to ISO-8859-1 in most other  
> > respects so having just one part (free text search) being full UTF-8  
> > compliant is of limited use.
> 
> I don't understand what you are saying here. All lucene indexes are  
> UTF-8, that includes yours too.  
> Ah, right.  
> My point is, as I understood the ICU docs, it is specialized in handling  
> UTF-8 locales with regard to analyzing and collation.  
> My needs is, for the forseable future, only with the nordic languages in  
> mind so I will only be using a very small portion of what ICU is capble  
> of. And if using ICU is complicated, sticking to the "normal" lucene  
> stuff may well work for my needs.

So I am still looking for what is the best route to follow. What I think  
I need to do is filter for the special library sorting rules ( like i=j,  
v=w etc) with a custom filter that I create as a plugin and then run it  
through an ordinary analyzer like snowball with swedish stemming. I  
think I have figured out how to do it from the latest contributions to  
the source (snowball filter is part of 0.15). Collation is another  
question I am investigating. What controls that? I need swedish collation...

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [February 7, 2011, 4:29pm UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/7 "2011-02-07T16:29:20Z")

</div>

On Mon, Feb 7, 2011 at 8:10 AM, Kristian Jörg [krjg@devo.se](mailto:krjg@devo.se) wrote:

> Ah, right.  
> My point is, as I understood the ICU docs, it is specialized in handling  
> UTF-8 locales with regard to analyzing and collation.  
> My needs is, for the forseable future, only with the nordic languages in  
> mind so I will only be using a very small portion of what ICU is capble of.  
> And if using ICU is complicated, sticking to the "normal" lucene stuff may  
> well work for my needs.
> 
> So I am still looking for what is the best route to follow. What I think I  
> need to do is filter for the special library sorting rules ( like i=j, v=w  
> etc) with a custom filter that I create as a plugin and then run it through  
> an ordinary analyzer like snowball with swedish stemming. I think I have  
> figured out how to do it from the latest contributions to the source  
> (snowball filter is part of 0.15). Collation is another question I am  
> investigating. What controls that? I need swedish collation...

Collation and locales don't have anything to do with UTF-8... as far  
as using ICU here its just as easy as using the JDK support! Just drop  
in the extra jar file.  
You can see some examples here:  
[http://lucene.apache.org/java/3\_0\_3/api/contrib-collation/org/apache/lucene/collation/package-summary.html](http://lucene.apache.org/java/3_0_3/api/contrib-collation/org/apache/lucene/collation/package-summary.html)

I wouldn't recommend trying to normalize text yourself to make your  
own collation keys... i would use the built in support  
In lucene this is just as easy as new  
ICUCollationKeyAnalyzer(Collator.getInstance(new ULocale("sv")));  
then you are indexing sort keys for swedish collation.

if your library truly does have special sorting rules, you can take an  
existing collator and customize it, here's an example:  
[http://wiki.apache.org/solr/UnicodeCollation#Sorting\_text\_with\_custom\_rules](http://wiki.apache.org/solr/UnicodeCollation#Sorting_text_with_custom_rules)

But first, i would explore the built-in rules to make sure they don't  
satisfy your requirements first (you can do this with ICU's locale  
explorer, e.g.):

> **[ICU Demonstration - Locale Explorer](https://icu4c-demos.unicode.org/icu-bin/locexp?_=sv_SE&d_=en&x=col)**
>
> Here is a demonstration of how ICU locale data works.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 7, 2011, 6:31pm UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/8 "2011-02-07T18:31:57Z")

</div>

There is the ICU plugin that provides the ICU level token filters that you are after: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/icu-plugin.html/)  
On Monday, February 7, 2011 at 6:29 PM, Robert Muir wrote:

> On Mon, Feb 7, 2011 at 8:10 AM, Kristian JÃ¶rg [krjg@devo.se](mailto:krjg@devo.se) wrote:
> 
> > Ah, right.  
> > My point is, as I understood the ICU docs, it is specialized in handling  
> > UTF-8 locales with regard to analyzing and collation.  
> > My needs is, for the forseable future, only with the nordic languages in  
> > mind so I will only be using a very small portion of what ICU is capble of.  
> > And if using ICU is complicated, sticking to the "normal" lucene stuff may  
> > well work for my needs.
> > 
> > So I am still looking for what is the best route to follow. What I think I  
> > need to do is filter for the special library sorting rules ( like i=j, v=w  
> > etc) with a custom filter that I create as a plugin and then run it through  
> > an ordinary analyzer like snowball with swedish stemming. I think I have  
> > figured out how to do it from the latest contributions to the source  
> > (snowball filter is part of 0.15). Collation is another question I am  
> > investigating. What controls that? I need swedish collation...
> 
> Collation and locales don't have anything to do with UTF-8... as far  
> as using ICU here its just as easy as using the JDK support! Just drop  
> in the extra jar file.  
> You can see some examples here:  
> [org.apache.lucene.collation (Lucene 3.0.3 API)](http://lucene.apache.org/java/3_0_3/api/contrib-collation/org/apache/lucene/collation/package-summary.html)
> 
> I wouldn't recommend trying to normalize text yourself to make your  
> own collation keys... i would use the built in support  
> In lucene this is just as easy as new  
> ICUCollationKeyAnalyzer(Collator.getInstance(new ULocale("sv")));  
> then you are indexing sort keys for swedish collation.
> 
> if your library truly does have special sorting rules, you can take an  
> existing collator and customize it, here's an example:  
> [UnicodeCollation - Solr - Apache Software Foundation](http://wiki.apache.org/solr/UnicodeCollation#Sorting_text_with_custom_rules)
> 
> But first, i would explore the built-in rules to make sure they don't  
> satisfy your requirements first (you can do this with ICU's locale  
> explorer, e.g.):  
> [ICU Demonstration - Locale Explorer](http://demo.icu-project.org/icu-bin/locexp?_=sv_SE&d_=en&x=col)

---

<div class="post-metadata">

**Author:** ![Barsk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/barsk/32/13734_2.png) [@Barsk](https://discuss.elastic.co/u/Barsk)\
**Post date:** [February 8, 2011, 8:12am UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/9 "2011-02-08T08:12:20Z")

</div>

Robert Muir skrev 2011-02-07 17:29:

> On Mon, Feb 7, 2011 at 8:10 AM, Kristian JÃ¶rg[krjg@devo.se](mailto:krjg@devo.se) wrote:
> 
> > Ah, right.  
> > My point is, as I understood the ICU docs, it is specialized in handling  
> > UTF-8 locales with regard to analyzing and collation.  
> > My needs is, for the forseable future, only with the nordic languages in  
> > mind so I will only be using a very small portion of what ICU is capble of.  
> > And if using ICU is complicated, sticking to the "normal" lucene stuff may  
> > well work for my needs.
> > 
> > So I am still looking for what is the best route to follow. What I think I  
> > need to do is filter for the special library sorting rules ( like i=j, v=w  
> > etc) with a custom filter that I create as a plugin and then run it through  
> > an ordinary analyzer like snowball with swedish stemming. I think I have  
> > figured out how to do it from the latest contributions to the source  
> > (snowball filter is part of 0.15). Collation is another question I am  
> > investigating. What controls that? I need swedish collation...
> 
> Collation and locales don't have anything to do with UTF-8... as far  
> as using ICU here its just as easy as using the JDK support! Just drop  
> in the extra jar file.  
> You can see some examples here:  
> [org.apache.lucene.collation (Lucene 3.0.3 API)](http://lucene.apache.org/java/3_0_3/api/contrib-collation/org/apache/lucene/collation/package-summary.html)
> 
> I wouldn't recommend trying to normalize text yourself to make your  
> own collation keys... i would use the built in support  
> In lucene this is just as easy as new  
> ICUCollationKeyAnalyzer(Collator.getInstance(new ULocale("sv")));  
> then you are indexing sort keys for swedish collation.
> 
> if your library truly does have special sorting rules, you can take an  
> existing collator and customize it, here's an example:  
> [UnicodeCollation - Solr - Apache Software Foundation](http://wiki.apache.org/solr/UnicodeCollation#Sorting_text_with_custom_rules)
> 
> But first, i would explore the built-in rules to make sure they don't  
> satisfy your requirements first (you can do this with ICU's locale  
> explorer, e.g.):  
> [ICU Demonstration - Locale Explorer](http://demo.icu-project.org/icu-bin/locexp?_=sv_SE&d_=en&x=col)  
> Thanks Robert,

a very nice answer to my questions, and I did also miss some of the info  
in your first reply that covers a fair bit. I was a bit preoccupied then  
I guess.  
I will look into all of this in depth during the day and and try to  
build the index with normalizing and collation as I need it and see how  
it goes.

One thing that is still not crystal clear to me is how to actually USE  
the customized filters I need in ES. In Lucene it is kind of built-in.  
Here it looks like I need to build a wrapper and deploy it as a JAR. But  
things will clear once I get started I suppose.

Thanks again for all your splendid support

/Kristian

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [February 8, 2011, 11:41am UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/10 "2011-02-08T11:41:59Z")

</div>

On Tue, Feb 8, 2011 at 3:12 AM, Kristian Jörg [krjg@devo.se](mailto:krjg@devo.se) wrote:

> One thing that is still not crystal clear to me is how to actually USE the  
> customized filters I need in ES. In Lucene it is kind of built-in. Here it  
> looks like I need to build a wrapper and deploy it as a JAR. But things will  
> clear once I get started I suppose.

Did you see Shay Banon's response? It appears elasticsearch already  
has a nice integration with these filters:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:12am UTC](https://discuss.elastic.co/t/custom-normalisation-and-filtering/3882/11 "2017-07-06T04:12:23Z")

</div>


