# unicodeSetFilter in analysis-icu ignored

**URL:** <https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432>\
**Category:** Elasticsearch\
**Created:** [March 25, 2016, 11:24am UTC](https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432 "2016-03-25T11:24:05Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Barsk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/barsk/32/13734_2.png) [@Barsk](https://discuss.elastic.co/u/Barsk)\
**Post date:** [March 25, 2016, 11:24am UTC](https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432/1 "2016-03-25T11:24:05Z")

</div>

Upgraded to ES 2.2.1 from a very old 0.18 installation and I have run inte a problem.  
I have two environments where I have configured an analysis-icu analyzer. In the first envoronment (win 7) everytyhing works perfectly, but in my other environment (Liniux) the unicodeSetFilter parameter is ignored. The elasticsearch.yml file looks like this and is utf-8 encoded:

```
index :
    analysis :
        analyzer : 
           swedishIcuFoldingAnalyzer :
               type : custom
               tokenizer : standard
               filter : [icuFolding, lowercase, swedish_stop]
   
        filter :
           swedish_stop :
                 type : stop
                 stopwords : _swedish_ 
           icuFolding :
                type : icu_folding
                unicodeSetFilter : "[^åäöÅÄÖ]"

```

When I query for the text "modrar" it will match "mödrar" in the Linux environment while on Windows it will not - as expected.

Any clues...?

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [March 25, 2016, 5:37pm UTC](https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432/2 "2016-03-25T17:37:35Z")

</div>

I don't think it makes sense to use the `standard` tokenizer with ICU. Instead, I recommend the `icu_tokenizer`

Also, as a side node, you should always apply stop word filter first, before lowercase or folding.

If the ICU plugin by Elastic really does not work, which would be very strange, I can offer an alternative implementation at [https://github.com/jprante/elasticsearch-plugin-bundle/](https://github.com/jprante/elasticsearch-plugin-bundle/) where I just added an ICU folding filter test that succeeds like you have described [https://github.com/jprante/elasticsearch-plugin-bundle/blob/403480349d5caf055e835c0dc3cf8bc798ec9359/src/test/java/org/xbib/elasticsearch/index/analysis/icu/IcuFoldingFilterTests.java](https://github.com/jprante/elasticsearch-plugin-bundle/blob/403480349d5caf055e835c0dc3cf8bc798ec9359/src/test/java/org/xbib/elasticsearch/index/analysis/icu/IcuFoldingFilterTests.java)

---

<div class="post-metadata">

**Author:** ![Barsk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/barsk/32/13734_2.png) [@Barsk](https://discuss.elastic.co/u/Barsk)\
**Post date:** [March 26, 2016, 11:22am UTC](https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432/3 "2016-03-26T11:22:38Z")

</div>

You may be correct with the icu\_tokenizer instead of standard, but it makes no difference here.  
Also the ordering of filters are a good point, but I believe lowercase should come before stopwords, right? I mean, all the stopwords are in lowercase.

When it comes to the analysis-icu plugin. It does work, The problem is that it is ignoring, or misinterpreting the unicodeSetFilter parameter. This is strange.

---

<div class="post-metadata">

**Author:** ![Barsk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/barsk/32/13734_2.png) [@Barsk](https://discuss.elastic.co/u/Barsk)\
**Post date:** [March 29, 2016, 11:29am UTC](https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432/4 "2016-03-29T11:29:14Z")

</div>

Well, it turned out to be a stupid error on my part. I had two instances of ES running apparently on the same machine and restarting the service had no effect since there was another instance blocking the reloads...

It seems ES is quietly just finding the next unoccupied port number and will not even give a warning that the standard port is blocked.

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [March 30, 2016, 12:43pm UTC](https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432/5 "2016-03-30T12:43:22Z")

</div>

> [@Barsk](#):
>
> It seems ES is quietly just finding the next unoccupied port number and will not even give a warning that the standard port is blocked.

That's correct, this is the intended behavior. There is no "standard port" but a port range (9200-9300, 9300-9400). It is supposed to ease demonstrations with multicast, so you can just start many ES processes on the same machine as you wish, or on the same network, they will find and form a cluster. Since multicast is gone, this behavior seems weird.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:04pm UTC](https://discuss.elastic.co/t/unicodesetfilter-in-analysis-icu-ignored/45432/6 "2017-07-05T23:04:05Z")

</div>


