# How to so sort multy token string and distinct feature + utf8 support?

**URL:** <https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124>\
**Category:** Elasticsearch\
**Created:** [December 12, 2011, 7:35am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124 "2011-12-12T07:35:32Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pascal\_Pensa](https://avatars.discourse-cdn.com/v4/letter/p/48db29/32.png) [@Pascal\_Pensa](https://discuss.elastic.co/u/Pascal_Pensa)\
**Post date:** [December 12, 2011, 7:35am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/1 "2011-12-12T07:35:32Z")

</div>

Hi,

I'm new to ES ans trying to sort strings i always get an error if  
string contains more than one word.

Another question is about dynamic deduplication/distinct/unique based  
on a field, i've searched along the ES wiki and trying to search in  
this group without success, does ES provides a "unique" feature or  
something equivalent removing duplicates answers given a field ?  
something like:

..?q=some+key+words&unique=reference

removing any duplicates from the resultset based on the "reference"  
tag.

And the last one, my test data contains accents, i tried various  
configurations of analysers, installed the icu plugin and set it into  
a filter, set langage to french, but it seems accents are not removed  
from tokenized items.

I'm actually using sphinxsearch and accents need to be manually table-  
mapped into the configuration file, is there an quivalent into ES ?

Thanks !  
Pascal

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 12, 2011, 9:02am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/2 "2011-12-12T09:02:13Z")

</div>

Hi

> I'm new to ES ans trying to sort strings i always get an error if  
> string contains more than one word.

You'll need to index them via keyword analyzer

> Another question is about dynamic deduplication/distinct/unique based  
> on a field

issue 256 regarding group by feature is not yet implemented. You'll  
need to do it on the client side.

> And the last one, my test data contains accents, i tried various  
> configurations of analysers, installed the icu plugin and set it into  
> a filter, set langage to french, but it seems accents are not removed  
> from tokenized items.

Did you tried the custom rules of the icu plugin? I read an article  
that it should be somehow possible ... I'll check

Regards,  
Peter.

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 12, 2011, 9:54am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/3 "2011-12-12T09:54:51Z")

</div>

Hi

> Did you tried the custom rules of the icu plugin? I read an article  
> that it should be somehow possible ... I'll check

Hmmh, strange for German umlauts it is done in the filter:

[http://web.archiveorange.com/archive/v/xJxT8VzgTaUwuBXnP9gJ](http://web.archiveorange.com/archive/v/xJxT8VzgTaUwuBXnP9gJ)

for french not I think:

[http://grepcode.com/file/repository.grepcode.com/java/eclipse.org/3.7/org.apache.lucene/analysis/2.9.1/org/apache/lucene/analysis/fr/FrenchStemmer.java](http://grepcode.com/file/repository.grepcode.com/java/eclipse.org/3.7/org.apache.lucene/analysis/2.9.1/org/apache/lucene/analysis/fr/FrenchStemmer.java)

I fear you will have to patch it or include another stemmer. Ah, but I  
saw you already asked it at the right place 🙂 (french elasticsearch  
group)

Regards,  
Peter.

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 12, 2011, 10:10am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/4 "2011-12-12T10:10:45Z")

</div>

This stemmer should do the work:

[https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/common/src/java/org/apache/lucene/analysis/fr/FrenchLightStemmer.java#L230](https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/common/src/java/org/apache/lucene/analysis/fr/FrenchLightStemmer.java#L230)

but you'll need to include it (and raise an issue?) like I did with a  
custom filter:

> <https://github.com/karussell/Jetwick/blob/master/es/config/elasticsearch.json>

using this filter&factory and setting the stemmer.

[https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/common/src/java/org/apache/lucene/analysis/fr/FrenchStemFilter.java](https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/common/src/java/org/apache/lucene/analysis/fr/FrenchStemFilter.java)  
[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/FrenchStemTokenFilterFactory.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/FrenchStemTokenFilterFactory.java)

Something like

public class MyFrenchStemTokenFilterFactory extends  
FrenchStemTokenFilterFactory {

```
private final Set<?> exclusions;

@Inject
public FrenchStemTokenFilterFactory(Index index, @IndexSettings

```

Settings indexSettings, @Assisted String name, @Assisted Settings  
settings) {  
super(index, indexSettings, name, settings);  
}

```
@Override
public TokenStream create(TokenStream tokenStream) {
    return new FrenchStemFilter(tokenStream,

```

exclusions).setStemmer(new FrenchLightStemmer());  
}  
}

Peter.

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 12, 2011, 10:13am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/5 "2011-12-12T10:13:59Z")

</div>

Ok, sorry to bubble up once you should be able to simply use:

light\_french

as filter. found it in the code:

[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/StemmerTokenFilterFactory.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/analysis/StemmerTokenFilterFactory.java)

Regards,  
Peter.

On 12 Dez., 11:10, Karussell [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com) wrote:

> This stemmer should do the work:
> 
> [https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/com](https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/com)...
> 
> but you'll need to include it (and raise an issue?) like I did with a  
> custom filter:
> 
> [https://github.com/karussell/Jetwick/blob/master/es/config/elasticsea](https://github.com/karussell/Jetwick/blob/master/es/config/elasticsea)...
> 
> using this filter&factory and setting the stemmer.
> 
> [https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/com...https://github.com/elasticsearch/elasticsearch/blob/master/src/main/j](https://github.com/apache/lucene-solr/blob/trunk/modules/analysis/com...https://github.com/elasticsearch/elasticsearch/blob/master/src/main/j)...
> 
> Something like
> 
> public class MyFrenchStemTokenFilterFactory extends  
> FrenchStemTokenFilterFactory {
> 
> ```
> private final Set<?> exclusions;
> 
> @Inject
> public FrenchStemTokenFilterFactory(Index index, @IndexSettings
> 
> ```
> 
> Settings indexSettings, @Assisted String name, @Assisted Settings  
> settings) {  
> super(index, indexSettings, name, settings);  
> }
> 
> ```
> @Override
> public TokenStream create(TokenStream tokenStream) {
> return new FrenchStemFilter(tokenStream,
> 
> ```
> 
> exclusions).setStemmer(new FrenchLightStemmer());  
> }
> 
> }
> 
> Peter.

---

<div class="post-metadata">

**Author:** ![Pascal\_Pensa](https://avatars.discourse-cdn.com/v4/letter/p/48db29/32.png) [@Pascal\_Pensa](https://discuss.elastic.co/u/Pascal_Pensa)\
**Post date:** [December 13, 2011, 6:45am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/6 "2011-12-13T06:45:06Z")

</div>

Thanks, i'll try your suggestions and give feedback,

For unicity (or grouping as named in sphinxsearch) it's because we  
have products and videos duplicated in various sub catalogs /  
categories.  
We don't recombine similar entries as they have their own keywords,  
target url and so on depending on the portal they belong to, and our  
search is possible in a given portal or cross portal.  
In cross universe search only one result is displayed sorted by  
various factors (freshness, relevance, ...) others similar results are  
throwed, today everythnig is done by the search engine.

Throwing data client side is complicated as we have to get many  
results to build the navigation bar by removing duplicates and  
counting, imagine we may have thousand results it'll be a pain to  
paginate results, we're using ajax and we prefer search engine powered  
pagination to limit data transfer and platform load.

Pascal

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 13, 2011, 10:48am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/7 "2011-12-13T10:48:59Z")

</div>

Not sure if I completely followed your usecase but IMO one option in  
your case would be to use only one product with an array for the urls  
+categories and decide (via middle layer or client) which one to  
display.

> Throwing data client side is complicated as we have to get many  
> results to build the navigation bar by removing duplicates and  
> counting, imagine we may have thousand results it'll be a pain to  
> paginate results, we're using ajax and we prefer search engine powered  
> pagination to limit data transfer and platform load.

Ok, I more meant with 'client side' the middle layer (if there is  
any).

Regards,  
Peter.

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 13, 2011, 10:49am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/8 "2011-12-13T10:49:47Z")

</div>

Also have a look into parent/child if that could solve your problem:

[http://www.elasticsearch.org/guide/reference/query-dsl/top-children-query.html](http://www.elasticsearch.org/guide/reference/query-dsl/top-children-query.html)

---

<div class="post-metadata">

**Author:** ![Pascal\_Pensa](https://avatars.discourse-cdn.com/v4/letter/p/48db29/32.png) [@Pascal\_Pensa](https://discuss.elastic.co/u/Pascal_Pensa)\
**Post date:** [December 13, 2011, 6:20pm UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/9 "2011-12-13T18:20:29Z")

</div>

thanks,

Found the way to remove accents: added asciifolding filter as french  
stemmer doesn't

Pascal

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:45am UTC](https://discuss.elastic.co/t/how-to-so-sort-multy-token-string-and-distinct-feature-utf8-support/6124/10 "2017-07-06T03:45:33Z")

</div>


