# Facet filters with ICU folding?

**URL:** <https://discuss.elastic.co/t/facet-filters-with-icu-folding/12791>\
**Category:** Elasticsearch\
**Created:** [July 15, 2013, 3:37pm UTC](https://discuss.elastic.co/t/facet-filters-with-icu-folding/12791 "2013-07-15T15:37:32Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Guillermo\_Arias\_del\_](https://avatars.discourse-cdn.com/v4/letter/g/cab0a1/32.png) [@Guillermo\_Arias\_del\_](https://discuss.elastic.co/u/Guillermo_Arias_del_)\
**Post date:** [July 15, 2013, 3:37pm UTC](https://discuss.elastic.co/t/facet-filters-with-icu-folding/12791/1 "2013-07-15T15:37:32Z")

</div>

Hi, all!

I have an index with a field "\_tokens" which has relevant tokens associated  
with a document. This field is configured as follows:

"\_token" : {  
"type" : "multi\_field",  
"fields" : {  
"\_token" : {  
"type" : "string",  
"index" : "not\_analyzed",  
...  
},  
"folded" : {  
"type" : "string",  
"analyzer" : "folded",  
...  
},  
"folded\_edge\_ngram" : {  
"type" : "string",  
"index\_analyzer" : "folded\_edge\_ngram",  
"search\_analyzer" : "folded",  
...  
}  
}  
}  
}

The analyzer "folded" and "folded\_edge\_ngram" are ICU folded and the latter  
has edge\_ngram as well.

I'm tring to do a search using the following code:

{  
"size": 0,  
"query": {  
"bool": {  
"must": [  
{  
"term": {  
"\_token.folded\_edge\_ngram": "bar"  
}  
}  
]  
}  
},  
"facets" : {  
"tokens" : {  
"terms" : {  
"field" : "\_token"  
}  
}  
}  
}

It returns all tokens beginning with "bar" with ICU folding, such as "Bär"  
or "bar". But it also returns related tokens (remember that there can be  
more than one token in "\_tokens"), so I want to restrict the facets with  
something like:

"exclude": doesn't work, because it only supports a full term match  
"regex": it works to an extent (match beginning, case insensitive), but it  
doesn't do ICU folding  
"scripts": OMG, how does this work?

So, my question is: Is there a form to reduce the facets based on a match  
with the ICU folding analyer? Or, am I totally wrong and should be using  
something else (more probable)?

P.S. : Afterwards, I also need the opposite. That is: search all documents  
containing a (ICU folded) word and do a faceting among the other terms  
(this has to do with autocompletion).

Thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [July 15, 2013, 9:19pm UTC](https://discuss.elastic.co/t/facet-filters-with-icu-folding/12791/2 "2013-07-15T21:19:58Z")

</div>

Think of facet entries like visual entities. You can run a facet query  
on a ICU folded field, but ICU folded terms are not really suitable for  
being visual entities. If you facet them, you receive "bar" for "bar and  
"bar" for "Bär". So far, so bad.

For this, I always use keyword-analyzed fields for faceting, like you do  
with multifielding. So I get two entries for "bar" and "Bär", as in the  
original document.

The challenge I have is the facet entries being sorted by ICU  
collations, so I once openend a pull request

> <https://github.com/elastic/elasticsearch-analysis-icu/pull/7/>
>
> Hi, 
> 
> I have implemented an ICU term facet that allows sorting the term entries …by ICU collation rules. It is useful for linguistic applications. An example for german phonebook sorting is included.
> 
> Cheers, Jörg

Or do you want to collapse "bar" and "Bär" into one facet entry by  
intention?

Jörg

Am 15.07.13 17:37, schrieb Guillermo Arias del Río:

> So, my question is: Is there a form to reduce the facets based on a  
> match with the ICU folding analyer?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Guillermo\_Arias\_del\_](https://avatars.discourse-cdn.com/v4/letter/g/cab0a1/32.png) [@Guillermo\_Arias\_del\_](https://discuss.elastic.co/u/Guillermo_Arias_del_)\
**Post date:** [July 16, 2013, 7:19am UTC](https://discuss.elastic.co/t/facet-filters-with-icu-folding/12791/3 "2013-07-16T07:19:42Z")

</div>

Hi, Jörg,

I am trying to perform autocompletion, and I want to do it with ICU  
folding. If the user types "bar", I want tokens like "Bär", "bar" and  
"barman". But it gets more interesting: I want the user to be able to give  
me more than one token. For example:

- User types: "hello wor"
- I match "hello" against "\_tokens.folded" and match "wor" against  
"\_tokens.folded\_edge\_ngram"
- ES gives me the documents, for instance: document1.\_tokens = [ "Hello"  
"World" ], document2.\_tokens = ["hello" "word" "blabla"]
- I want to exclude "Hello", "hello", and "blabla"; and retain "World"  
and "word"

If it were a search against an unanalyzed field, I could accomplish this  
with "exclude" and "regex", but I can't. So now, I am looping through the  
results and filtering myself, which means calling \_analyze for each  
result...

Maybe I should try with another index structure, I don't know.

Guillermo.

El lunes, 15 de julio de 2013 23:19:58 UTC+2, Jörg Prante escribió:

> Think of facet entries like visual entities. You can run a facet query  
> on a ICU folded field, but ICU folded terms are not really suitable for  
> being visual entities. If you facet them, you receive "bar" for "bar and  
> "bar" for "Bär". So far, so bad.
> 
> For this, I always use keyword-analyzed fields for faceting, like you do  
> with multifielding. So I get two entries for "bar" and "Bär", as in the  
> original document.
> 
> The challenge I have is the facet entries being sorted by ICU  
> collations, so I once openend a pull request  
> [Adding ICU collation based sorting for facets by jprante · Pull Request #7 · elastic/elasticsearch-analysis-icu · GitHub](https://github.com/elasticsearch/elasticsearch-analysis-icu/pull/7/)
> 
> Or do you want to collapse "bar" and "Bär" into one facet entry by  
> intention?
> 
> Jörg
> 
> Am 15.07.13 17:37, schrieb Guillermo Arias del Río:
> 
> > So, my question is: Is there a form to reduce the facets based on a  
> > match with the ICU folding analyer?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [July 16, 2013, 7:35am UTC](https://discuss.elastic.co/t/facet-filters-with-icu-folding/12791/4 "2013-07-16T07:35:49Z")

</div>

I still don't get what your proposed role of facet filters is.

Autocompletion works best on edge-n-gram fields. You can combine edge  
n-gram and ICU folding. No need to take care for exclude, regex, and  
\_analyze.

But, it seems you want autosuggest, not autocomplete. That is, you want  
to deliver a list of ranked suggestions for a start of word/phrase.  
Check if the Suggest API can help you:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Jörg

Am 16.07.13 09:19, schrieb Guillermo Arias del Río:

> Hi, Jörg,
> 
> I am trying to perform autocompletion, and I want to do it with ICU  
> folding. If the user types "bar", I want tokens like "Bär", "bar" and  
> "barman". But it gets more interesting: I want the user to be able to  
> give me more than one token. For example:
> 
> - User types: "hello wor"
> - I match "hello" against "\_tokens.folded" and match "wor" against  
> "\_tokens.folded\_edge\_ngram"
> - ES gives me the documents, for instance: document1.\_tokens = [  
> "Hello" "World" ], document2.\_tokens = ["hello" "word" "blabla"]
> - I want to exclude "Hello", "hello", and "blabla"; and retain  
> "World" and "word"
> 
> If it were a search against an unanalyzed field, I could accomplish  
> this with "exclude" and "regex", but I can't. So now, I am looping  
> through the results and filtering myself, which means calling \_analyze  
> for each result...
> 
> Maybe I should try with another index structure, I don't know.
> 
> Guillermo.
> 
> El lunes, 15 de julio de 2013 23:19:58 UTC+2, Jörg Prante escribió:
> 
> ```
> Think of facet entries like visual entities. You can run a facet
> query
> on a ICU folded field, but ICU folded terms are not really
> suitable for
> being visual entities. If you facet them, you receive "bar" for
> "bar and
> "bar" for "Bär". So far, so bad.
> 
> For this, I always use keyword-analyzed fields for faceting, like
> you do
> with multifielding. So I get two entries for "bar" and "Bär", as
> in the
> original document.
> 
> The challenge I have is the facet entries being sorted by ICU
> collations, so I once openend a pull request
> https://github.com/elasticsearch/elasticsearch-analysis-icu/pull/7/ <https://github.com/elasticsearch/elasticsearch-analysis-icu/pull/7/>
> 
> Or do you want to collapse "bar" and "Bär" into one facet entry by
> intention?
> 
> Jörg
> 
> Am 15.07.13 17:37, schrieb Guillermo Arias del Río:
> > So, my question is: Is there a form to reduce the facets based on a
> > match with the ICU folding analyer?
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google  
> Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send  
> an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:26am UTC](https://discuss.elastic.co/t/facet-filters-with-icu-folding/12791/5 "2017-07-06T02:26:19Z")

</div>


