# \[Ann\] Elasticsearch Word Decompound Plugin

**URL:** <https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773>\
**Category:** Elasticsearch\
**Created:** [November 20, 2012, 11:47pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773 "2012-11-20T23:47:54Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [November 20, 2012, 11:47pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/1 "2012-11-20T23:47:54Z")

</div>

Hi,

in my spare time this evening, while I'm still wrangling with some NLP  
plugins (Stanford , UIMA, OpenNLP), and eagerly awaiting Lucene 4, I  
reworked a Compact Patricia Trie implementation of Chris Biemann for a  
german word decompounding Elasticsearch analysis plugin.

It can decompound german words like "Rechtsanwaltskanzleien" into "Recht,  
anwalt, kanzlei" or "Jahresfeier" into "Jahr, feier". The best thing is,  
you don't need to provide a word list.

You can find it  
here: [https://github.com/jprante/elasticsearch-analysis-decompound](https://github.com/jprante/elasticsearch-analysis-decompound)

Have fun!

Jörg

--

---

<div class="post-metadata">

**Author:** ![jarib](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jarib/32/882_2.png) [@jarib](https://discuss.elastic.co/u/jarib)\
**Post date:** [November 23, 2012, 12:10am UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/2 "2012-11-23T00:10:05Z")

</div>

Very cool!

Do you expect this to work for other languages as well? I see Norwegian  
mentioned in the README, which is exactly what I'm after.

I was debugging memory issues with ES _today_ which turned out to be caused  
by my (probably way too large) dictionary, so if this works out it's a  
godsend.

Jari

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [November 23, 2012, 1:13am UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/3 "2012-11-23T01:13:29Z")

</div>

Hi Jari,

just give it a shot. The CPT data provided is derived from the Leipzig  
Wortschatz, which is german, so I doubt it works flawlessly for Norwegian.  
I could try to ask Chris Biemann if he knows how to build Norwegian  
decompounder CPTs.

Best regards,

Jörg

On Friday, November 23, 2012 1:10:05 AM UTC+1, jarib wrote:

> Very cool!
> 
> Do you expect this to work for other languages as well? I see Norwegian  
> mentioned in the README, which is exactly what I'm after.
> 
> I was debugging memory issues with ES _today_ which turned out to be  
> caused by my (probably way too large) dictionary, so if this works out it's  
> a godsend.
> 
> Jari

--

---

<div class="post-metadata">

**Author:** ![jarib](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jarib/32/882_2.png) [@jarib](https://discuss.elastic.co/u/jarib)\
**Post date:** [November 23, 2012, 1:24am UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/4 "2012-11-23T01:24:30Z")

</div>

Hi Jörg,

I did some simple tests which appear to work ok, but if it's possible to  
improve I'd be interested in working on it. Please ask Chris (or let me  
know how to contact him)!

Jari

On Fri, Nov 23, 2012 at 2:13 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com) wrote:

> Hi Jari,
> 
> just give it a shot. The CPT data provided is derived from the Leipzig  
> Wortschatz, which is german, so I doubt it works flawlessly for Norwegian.  
> I could try to ask Chris Biemann if he knows how to build Norwegian  
> decompounder CPTs.
> 
> Best regards,
> 
> Jörg
> 
> On Friday, November 23, 2012 1:10:05 AM UTC+1, jarib wrote:
> 
> > Very cool!
> > 
> > Do you expect this to work for other languages as well? I see Norwegian  
> > mentioned in the README, which is exactly what I'm after.
> > 
> > I was debugging memory issues with ES _today_ which turned out to be  
> > caused by my (probably way too large) dictionary, so if this works out it's  
> > a godsend.
> > 
> > Jari
> 
> --

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [November 23, 2012, 9:44am UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/5 "2012-11-23T09:44:43Z")

</div>

Hi Jari,

the reason that you find it to work ok for norwegian is because there is a  
strong relationship between norwegian and german language. Chris Biemann  
confirmed, it is required to train the three Compact Patricia Tries (CPTs)  
for other languages for correct decompounding. If an already decompounded  
word list for norwegian can be provided, you are lucky. If not, he  
suggested a rough approach, by using the Morfessor tool of  
[http://www.cis.hut.fi/projects/morpho/](http://www.cis.hut.fi/projects/morpho/) that can automatically generate  
decompounded word lists out of existing word lists, as they are provided  
by [http://corpora.informatik.uni-leipzig.de/download.html](http://corpora.informatik.uni-leipzig.de/download.html)

I will see if I can provide a script in the plugin distribution ZIP that  
can train CPTs for other languages beside german that have compounded and  
agglutinated forms as well (scandinavian languages)

Cheers,

Jörg

On Friday, November 23, 2012 2:24:53 AM UTC+1, jarib wrote:

> Hi Jörg,
> 
> I did some simple tests which appear to work ok, but if it's possible to  
> improve I'd be interested in working on it. Please ask Chris (or let me  
> know how to contact him)!
> 
> Jari
> 
> On Fri, Nov 23, 2012 at 2:13 AM, Jörg Prante \<[joerg...@gmail.com](mailto:joerg...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > Hi Jari,
> > 
> > just give it a shot. The CPT data provided is derived from the Leipzig  
> > Wortschatz, which is german, so I doubt it works flawlessly for Norwegian.  
> > I could try to ask Chris Biemann if he knows how to build Norwegian  
> > decompounder CPTs.
> > 
> > Best regards,
> > 
> > Jörg
> > 
> > On Friday, November 23, 2012 1:10:05 AM UTC+1, jarib wrote:
> > 
> > > Very cool!
> > > 
> > > Do you expect this to work for other languages as well? I see Norwegian  
> > > mentioned in the README, which is exactly what I'm after.
> > > 
> > > I was debugging memory issues with ES _today_ which turned out to be  
> > > caused by my (probably way too large) dictionary, so if this works out it's  
> > > a godsend.
> > > 
> > > Jari
> > 
> > --

--

---

<div class="post-metadata">

**Author:** ![jarib](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jarib/32/882_2.png) [@jarib](https://discuss.elastic.co/u/jarib)\
**Post date:** [November 23, 2012, 2:40pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/6 "2012-11-23T14:40:39Z")

</div>

On Fri, Nov 23, 2012 at 10:44 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com) wrote:

> If an already decompounded word list for norwegian can be provided, you  
> are lucky.

Do you have an example of what this file should look like? The norwegian  
spell check project at [http://no.speling.org/](http://no.speling.org/) has a lot of relevant data.  
I'll have to dig in to see if they have a proper decompounded word list.

> I will see if I can provide a script in the plugin distribution ZIP that  
> can train CPTs for other languages beside german that have compounded and  
> agglutinated forms as well (scandinavian languages)

That would be fantastic. If I find the data, would you want new language  
trees included in the plugin?

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [November 23, 2012, 4:58pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/7 "2012-11-23T16:58:03Z")

</div>

Hi Jari,

On Friday, November 23, 2012 3:41:02 PM UTC+1, jarib wrote:

> On Fri, Nov 23, 2012 at 10:44 AM, Jörg Prante \<[joerg...@gmail.com](mailto:joerg...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > If an already decompounded word list for norwegian can be provided, you  
> > are lucky.
> 
> Do you have an example of what this file should look like? The norwegian  
> spell check project at [http://no.speling.org/](http://no.speling.org/) has a lot of relevant data.  
> I'll have to dig in to see if they have a proper decompounded word list.

Please refer to the Morfessor paper

> **[Creutz05tr.pdf](https://users.ics.aalto.fi/mcreutz/papers/Creutz05tr.pdf)**
>
> 174.60 KB

where a decompounded word list look like

_Smørbrød_  
_Midtsommernattsdrøm_  
...

-\>

_Smør + brød_  
\*  
Midt + sommer + natt + drøm  
...

- 

> > I will see if I can provide a script in the plugin distribution ZIP that  
> > can train CPTs for other languages beside german that have compounded and  
> > agglutinated forms as well (scandinavian languages)
> 
> That would be fantastic. If I find the data, would you want new language  
> trees included in the plugin?

Yes, I would do an update of the plugin, sure. With a ISO-639 language  
parameter, you could select the CPTs for the language. Beside the script I  
plan to develop so you could build CPTs for yourself... I don't think I can  
handle Korean for example.

Cheers,

Jörg

--

---

<div class="post-metadata">

**Author:** ![fabik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fabik/32/122740_2.png) [@fabik](https://discuss.elastic.co/u/fabik)\
**Post date:** [November 28, 2012, 11:23am UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/8 "2012-11-28T11:23:36Z")

</div>

Hey,

i just discovered your plugin after i was looking for something else to use than the nativ ES one. The problem i have, i am buildung a search for a products and running into the problem that some products have "herrenschuhe" and others "schuhe für herren" in the title. So my idea was to just run the filter against the titles and against the search query. But when running it against the search query it would break "herrenschuhe" into "herrenschuhe" + "herren" + "schuhe". To have this work the best way i would need the filter to drop the original "herrenschuhe". Would it be possible to add something like in the WordDelimiter, the preserve\_original param?

---

<div class="post-metadata">

**Author:** ![Bruce\_Ritchie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruce_ritchie/32/9370_2.png) [@Bruce\_Ritchie](https://discuss.elastic.co/u/Bruce_Ritchie)\
**Post date:** [January 9, 2013, 5:12pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/9 "2013-01-09T17:12:42Z")

</div>

Jörg,

This plugin is interesting to me for the German normalization so thanks for  
putting it together! Any chance you could get this uploaded to the new  
[download.elasticsearch.org](http://download.elasticsearch.org) service (or maven) ? I had to manually hack the  
github url to get at the 1.1.0 zip file for this plugin otherwise people  
would have to build manually atm.

Regards,

Bruce Ritchie

On Tuesday, November 20, 2012 6:47:54 PM UTC-5, Jörg Prante wrote:

> Hi,
> 
> in my spare time this evening, while I'm still wrangling with some NLP  
> plugins (Stanford , UIMA, OpenNLP), and eagerly awaiting Lucene 4, I  
> reworked a Compact Patricia Trie implementation of Chris Biemann for a  
> german word decompounding Elasticsearch analysis plugin.
> 
> It can decompound german words like "Rechtsanwaltskanzleien" into "Recht,  
> anwalt, kanzlei" or "Jahresfeier" into "Jahr, feier". The best thing is,  
> you don't need to provide a word list.
> 
> You can find it here:  
> [GitHub - jprante/elasticsearch-analysis-decompound: Decompounding Plugin for Elasticsearch](https://github.com/jprante/elasticsearch-analysis-decompound)
> 
> Have fun!
> 
> Jörg

--

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 9, 2013, 5:34pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/10 "2013-01-09T17:34:14Z")

</div>

You can always download the file and install it locally (-url file://....).  
No longer a cleaner one step process, but better than changing the source  
(IMHO).

--  
Ivan

On Wed, Jan 9, 2013 at 9:12 AM, Bruce Ritchie [bruce.ritchie@gmail.com](mailto:bruce.ritchie@gmail.com)wrote:

> This plugin is interesting to me for the German normalization so thanks  
> for putting it together! Any chance you could get this uploaded to the new  
> [download.elasticsearch.org](http://download.elasticsearch.org) service (or maven) ? I had to manually hack  
> the github url to get at the 1.1.0 zip file for this plugin otherwise  
> people would have to build manually atm.
> 
> Regards,
> 
> Bruce Ritchie
> 
> On Tuesday, November 20, 2012 6:47:54 PM UTC-5, Jörg Prante wrote:
> 
> > Hi,
> > 
> > in my spare time this evening, while I'm still wrangling with some NLP  
> > plugins (Stanford , UIMA, OpenNLP), and eagerly awaiting Lucene 4, I  
> > reworked a Compact Patricia Trie implementation of Chris Biemann for a  
> > german word decompounding Elasticsearch analysis plugin.
> > 
> > It can decompound german words like "Rechtsanwaltskanzleien" into "Recht,  
> > anwalt, kanzlei" or "Jahresfeier" into "Jahr, feier". The best thing is,  
> > you don't need to provide a word list.
> > 
> > You can find it here: [https://github.com/\*\*jprante/elasticsearch-](https://github.com/**jprante/elasticsearch-)\*\*  
> > analysis-decompound[https://github.com/jprante/elasticsearch-analysis-decompound](https://github.com/jprante/elasticsearch-analysis-decompound)
> > 
> > Have fun!
> > 
> > Jörg
> 
> --

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [January 9, 2013, 8:33pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/11 "2013-01-09T20:33:39Z")

</div>

Bruce,

thanks for your interest - right now there is no other method than  
downloading with a full URL. Github will remove the ZIP files soon.

I have no access to the maven search site URL download or to  
download.elasticsearch.org.

To improve the situation, I am reorganizing all my plugins now for better  
distribution, more to be announced on this list. My plan is to setup a  
Maven, RPM and deb distribution service at the brand new [bintray.com](http://bintray.com)  
service.

Regards,

Jörg

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:57am UTC](https://discuss.elastic.co/t/ann-elasticsearch-word-decompound-plugin/9773/12 "2017-07-06T02:57:09Z")

</div>


