# Why does hyphenation\_decompounder require word\_list?

**URL:** <https://discuss.elastic.co/t/why-does-hyphenation-decompounder-require-word-list/114567>\
**Category:** Elasticsearch\
**Created:** [January 8, 2018, 4:56pm UTC](https://discuss.elastic.co/t/why-does-hyphenation-decompounder-require-word-list/114567 "2018-01-08T16:56:22Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![hbruch](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hbruch/32/26310_2.png) [@hbruch](https://discuss.elastic.co/u/hbruch)\
**Post date:** [January 8, 2018, 4:56pm UTC](https://discuss.elastic.co/t/why-does-hyphenation-decompounder-require-word-list/114567/1 "2018-01-08T16:56:23Z")

</div>

HyphenationCompoundWordTokenFilterFactory inherits from AbstractCompoundWordTokenFilterFactory , which performs a mandatory check for a supplied word\_list.

As the underlying lucene HyphenationCompoundWordTokenFilter does not require a word\_list, is there a specific requirement, why it must be supplied for elasticsearch?

In my use case, I'd like to avoid specifying in advance all possible matching subwords.

Regards,  
Holger

---

<div class="post-metadata">

**Author:** ![hbruch](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hbruch/32/26310_2.png) [@hbruch](https://discuss.elastic.co/u/hbruch)\
**Post date:** [January 13, 2018, 2:24pm UTC](https://discuss.elastic.co/t/why-does-hyphenation-decompounder-require-word-list/114567/2 "2018-01-13T14:24:40Z")

</div>

Ok, I managed to work around this creating a custom analysis plugin that creates the HyphenationCompoundWordTokenFilter without wordlist.

However, applying the decompunder on index and search time I got unexpected results: at query time, all decompounded subwords seem to be treated as synonyms, so all documents containing just one subword get the same score as documents containing more(?).

I expected the token to be split in mulitple terms which are scored individually so documents containing both of the are ranked higher. This is same expectation as @singer had in [#11749](https://discuss.elastic.co/t/decompounder-in-query-string-analyzer/11749), I suppose.

explain results seem to indicate, that any subword is treated as a synonym for the complete compounded word(?):

`> "description" : "weight(Synonym(collector.default:scherenbosteler collector.default:scherenbostelerstrasse collector.default:strasse) in 912) [PerFieldSimilarity], result of:",`

How could I change this behaviour?

See also this [stackoverflow question](https://stackoverflow.com/questions/47710944/what-does-weightsynonym-mean-in-elasticsearch), if you could provide an explanation for weight(Synonym())

---

<div class="post-metadata">

**Author:** ![hbruch](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hbruch/32/26310_2.png) [@hbruch](https://discuss.elastic.co/u/hbruch)\
**Post date:** [January 13, 2018, 4:41pm UTC](https://discuss.elastic.co/t/why-does-hyphenation-decompounder-require-word-list/114567/3 "2018-01-13T16:41:56Z")

</div>

Seems that [reusing the original start/end offset](https://github.com/apache/lucene-solr/blob/master/lucene/analysis/common/src/java/org/apache/lucene/analysis/compound/CompoundWordTokenFilterBase.java#L142) advices the QueryBuild to [build a SynonymQuery](https://github.com/apache/lucene-solr/blob/master/lucene/core/src/java/org/apache/lucene/util/QueryBuilder.java#L407) as it [collects all terms with a zero position increment](https://github.com/apache/lucene-solr/blob/master/lucene/core/src/java/org/apache/lucene/util/QueryBuilder.java#L423) in the uncleared currentQuery.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 10, 2018, 4:42pm UTC](https://discuss.elastic.co/t/why-does-hyphenation-decompounder-require-word-list/114567/4 "2018-02-10T16:42:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
