# \[Ann\] Elasticsearch Analysis Baseform Plugin 1.0.0

**URL:** <https://discuss.elastic.co/t/ann-elasticsearch-analysis-baseform-plugin-1-0-0/14047>\
**Category:** Elasticsearch\
**Created:** [October 21, 2013, 9:21pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-analysis-baseform-plugin-1-0-0/14047 "2013-10-21T21:21:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [October 21, 2013, 9:21pm UTC](https://discuss.elastic.co/t/ann-elasticsearch-analysis-baseform-plugin-1-0-0/14047/1 "2013-10-21T21:21:28Z")

</div>

Hi,

I have started a lexicon-based analyzer for linguistic processing of full  
word forms to their base form (right now, only german lexicon is provided)

> **[jprante/elasticsearch-analysis-baseform](https://github.com/jprante/elasticsearch-analysis-baseform)**
>
> elasticsearch-analysis-baseform - Baseform lemmatization for Elasticsearch

With this plugin, full word forms are reduced to base forms in the  
tokenization process. This is also known as lemmatization.

Why is lemmatization better than stemming? With this plugin, you can  
generate additional baseform tokens also for irregular word forms. Example:  
for the word "zurückgezogen", the base form is "zurückziehen". Algorithmic  
stemming would be rather limited for such cases.

Thanks to Dawid Weiss for the FSA and Daniel Naber for the german  
fullform/baseform lexicon.

Cheers,

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [October 22, 2013, 2:25am UTC](https://discuss.elastic.co/t/ann-elasticsearch-analysis-baseform-plugin-1-0-0/14047/2 "2013-10-22T02:25:58Z")

</div>

Fantastiq! 😉

Would it make sense to contribute the core of this to Lucene, where I'm  
sure this sort of thing would thrive?

## Thanks, Otis

Performance Monitoring \* Log Analytics \* Search Analytics  
Solr & Elasticsearch Support \* [http://sematext.com/](http://sematext.com/)

On Monday, October 21, 2013 5:21:28 PM UTC-4, Jörg Prante wrote:

> Hi,
> 
> I have started a lexicon-based analyzer for linguistic processing of full  
> word forms to their base form (right now, only german lexicon is provided)
> 
> [GitHub - jprante/elasticsearch-analysis-baseform: Baseform lemmatization for Elasticsearch](https://github.com/jprante/elasticsearch-analysis-baseform)
> 
> With this plugin, full word forms are reduced to base forms in the  
> tokenization process. This is also known as lemmatization.
> 
> Why is lemmatization better than stemming? With this plugin, you can  
> generate additional baseform tokens also for irregular word forms. Example:  
> for the word "zurückgezogen", the base form is "zurückziehen". Algorithmic  
> stemming would be rather limited for such cases.
> 
> Thanks to Dawid Weiss for the FSA and Daniel Naber for the german  
> fullform/baseform lexicon.
> 
> Cheers,
> 
> Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [October 22, 2013, 7:02am UTC](https://discuss.elastic.co/t/ann-elasticsearch-analysis-baseform-plugin-1-0-0/14047/3 "2013-10-22T07:02:31Z")

</div>

It already is (for polish) [https://issues.apache.org/jira/browse/LUCENE-2341](https://issues.apache.org/jira/browse/LUCENE-2341)

My version is a stripped down version of Dawid Weiss' morfologik FSA,  
attached with a reader for Daniel Naber's german lexicon, only for  
lemmatization. Morfologik can do much more (POS tagging).

It should be possible to create something like morfologik-german,  
morfologik-english morofologik-french etc. but I did not dig into it yet.

For Elasticsearch, Dariusz Gertych already implemented a morfologik plugin  
for polish stemming based on Lucene

> **[monterail/elasticsearch-analysis-morfologik](https://github.com/monterail/elasticsearch-analysis-morfologik)**
>
> elasticsearch-analysis-morfologik - Morfologik (Polish) Analysis Plugin for ElasticSearch

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:11am UTC](https://discuss.elastic.co/t/ann-elasticsearch-analysis-baseform-plugin-1-0-0/14047/4 "2017-07-06T02:11:15Z")

</div>


