# Importing language analyzers

**URL:** <https://discuss.elastic.co/t/importing-language-analyzers/4903>\
**Category:** Elasticsearch\
**Created:** [July 20, 2011, 8:28am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903 "2011-07-20T08:28:56Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pawel\_Konieczny](https://avatars.discourse-cdn.com/v4/letter/p/dc4da7/32.png) [@Pawel\_Konieczny](https://discuss.elastic.co/u/Pawel_Konieczny)\
**Post date:** [July 20, 2011, 8:28am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/1 "2011-07-20T08:28:56Z")

</div>

Hey!

Is it possible to import language analyzers from Lucene since ES is  
built on top of it (and Lucene has definitely more languages supported  
out of box)?  
The list of languages supported by EA is pretty extensive too, but it  
lacks polish language which I need 🙂  
EA seems much more flexible and has great potential (and I would like  
to use it in a project I'm developing), but without support for a  
given language it just won't do.  
Also, who adds language support to EA, its developers or the  
community?

Cheers,  
Pawel

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 20, 2011, 4:50pm UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/2 "2011-07-20T16:50:51Z")

</div>

Yes, you can hook your own analyzer, but you will need to implement a custom  
class that provides it. Check for example the GermanAnalyzerProvider. What  
is the name of the polish analyzer? I might have missed it and did not  
include it out of the box.

2011/7/20 Paweł Konieczny [koniecznypw@gmail.com](mailto:koniecznypw@gmail.com)

> Hey!
> 
> Is it possible to import language analyzers from Lucene since ES is  
> built on top of it (and Lucene has definitely more languages supported  
> out of box)?  
> The list of languages supported by EA is pretty extensive too, but it  
> lacks polish language which I need 🙂  
> EA seems much more flexible and has great potential (and I would like  
> to use it in a project I'm developing), but without support for a  
> given language it just won't do.  
> Also, who adds language support to EA, its developers or the  
> community?
> 
> Cheers,  
> Pawel

---

<div class="post-metadata">

**Author:** ![Pawel\_Konieczny](https://avatars.discourse-cdn.com/v4/letter/p/dc4da7/32.png) [@Pawel\_Konieczny](https://discuss.elastic.co/u/Pawel_Konieczny)\
**Post date:** [July 21, 2011, 10:56am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/3 "2011-07-21T10:56:24Z")

</div>

From what I understand, it's called Stempel and it's included in  
Lucene.

On Jul 20, 6:50 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Yes, you can hook your own analyzer, but you will need to implement a custom  
> class that provides it. Check for example the GermanAnalyzerProvider. What  
> is the name of the polish analyzer? I might have missed it and did not  
> include it out of the box.
> 
> 2011/7/20 Paweł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> 
> > Hey!
> 
> > Is it possible to import language analyzers from Lucene since ES is  
> > built on top of it (and Lucene has definitely more languages supported  
> > out of box)?  
> > The list of languages supported by EA is pretty extensive too, but it  
> > lacks polish language which I need 🙂  
> > EA seems much more flexible and has great potential (and I would like  
> > to use it in a project I'm developing), but without support for a  
> > given language it just won't do.  
> > Also, who adds language support to EA, its developers or the  
> > community?
> 
> > Cheers,  
> > Pawel

---

<div class="post-metadata">

**Author:** ![ofavre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ofavre/32/116819_2.png) [@ofavre](https://discuss.elastic.co/u/ofavre)\
**Post date:** [July 21, 2011, 2:55pm UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/4 "2011-07-21T14:55:21Z")

</div>

I think we should review all the available analyzers available, and identify  
the missing ones (ie not wrapped in ES).

I also found that one on the Internet for Chinese:

> **[Google Code Archive - Long-term storage for Google Code Project Hosting.](https://code.google.com/archive/p/ik-analyzer)**

With an ES plugin (at least a stub):

> <https://github.com/medcl/elasticsearch/blob/21abad12a0096173e8836dd042ca403751ab7ad1/plugins/analysis/ik/src/main/java/org/elasticsearch/index/analysis/IkAnalyzer.java>

But it's not part of Lucene-contrib.

Here is what I found, part of Lucene-contrib (apparently only those few \*  
language\* analyzers are missing) :

- Polish analyzer, with Stempel stemmer:  
[Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-stempel/index.html)
- Smart Chinese analyzer, with 2 flavors:  
[Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-smartcn/index.html)
- Latvian:  
[org.apache.lucene.analysis.lv (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/lucene/analysis/lv/package-summary.html)

Some other findings:

- KpStemmer (Kraaij-Pohlmann stemming algorithm for  
Dutch[1][http://snowball.tartarus.org/algorithms/kraaij\_pohlmann/stemmer.html](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html)  
?):  
[KpStemmer (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartarus/snowball/ext/KpStemmer.html)
- LovinsStemmer, alike PorterStemmer, but less interesting according to  
[2] [http://en.wikipedia.org/wiki/Stemming#History](http://en.wikipedia.org/wiki/Stemming#History):  
[LovinsStemmer (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartarus/snowball/ext/LovinsStemmer.html)
- Wikipedia-syntax-aware tokenizer:  
[org.apache.lucene.analysis.wikipedia (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/lucene/analysis/wikipedia/package-summary.html)
- WordNet synonym injector :  
[Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-wordnet/index.html)

The most interesting lists come from Solr itself:

- 

[http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-summary.html](http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-summary.html)

- [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis)

I didn't look very thoroughly at those last two links, but it looks that we  
may be missing:

- ClassicTokenizer (may be deprecated or superseded by the  
StandardTokenizer, I have no idea):  
[https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucene/analysis/standard/ClassicTokenizer.html?is-external=true](https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucene/analysis/standard/ClassicTokenizer.html?is-external=true)
- CommonGrams:  
[http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGramsFilter.html](http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGramsFilter.html)
- Lao, Myanmar, Khmer - seem to only split in syllables:  
[LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Lao.2C_Myanmar.2C_Khmer)

Towards a small easy pull-request?

```
[1] http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html
     Seen from

```

## [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Notes_about_solr.SnowballPorterFilterFactory) [2] [Stemming - Wikipedia](http://en.wikipedia.org/wiki/Stemming#History)

Olivier Favre

[www.yakaz.com](http://www.yakaz.com)

2011/7/21 Paweł Konieczny [koniecznypw@gmail.com](mailto:koniecznypw@gmail.com)

> From what I understand, it's called Stempel and it's included in  
> Lucene.
> 
> On Jul 20, 6:50 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Yes, you can hook your own analyzer, but you will need to implement a  
> > custom  
> > class that provides it. Check for example the GermanAnalyzerProvider.  
> > What  
> > is the name of the polish analyzer? I might have missed it and did not  
> > include it out of the box.
> > 
> > 2011/7/20 Paweł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> > 
> > > Hey!
> > 
> > > Is it possible to import language analyzers from Lucene since ES is  
> > > built on top of it (and Lucene has definitely more languages supported  
> > > out of box)?  
> > > The list of languages supported by EA is pretty extensive too, but it  
> > > lacks polish language which I need 🙂  
> > > EA seems much more flexible and has great potential (and I would like  
> > > to use it in a project I'm developing), but without support for a  
> > > given language it just won't do.  
> > > Also, who adds language support to EA, its developers or the  
> > > community?
> > 
> > > Cheers,  
> > > Pawel

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 22, 2011, 12:44am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/5 "2011-07-22T00:44:38Z")

</div>

Heya,

Yea, open issues for the missing analyzers, the stempel one, for example,  
should be simple to add (its in a different lib).

On Thu, Jul 21, 2011 at 5:55 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com) wrote:

> I think we should review all the available analyzers available, and  
> identify the missing ones (ie not wrapped in ES).
> 
> I also found that one on the Internet for Chinese:  
> [Google Code Archive - Long-term storage for Google Code Project Hosting.](http://code.google.com/p/ik-analyzer/)  
> With an ES plugin (at least a stub):  
> [https://github.com/medcl/elasticsearch/blob/21abad12a0096173e8836dd042ca403751ab7ad1/plugins/analysis/ik/src/main/java/org/elasticsearch/index/analysis/IkAnalyzer.java](https://github.com/medcl/elasticsearch/blob/21abad12a0096173e8836dd042ca403751ab7ad1/plugins/analysis/ik/src/main/java/org/elasticsearch/index/analysis/IkAnalyzer.java)  
> But it's not part of Lucene-contrib.
> 
> Here is what I found, part of Lucene-contrib (apparently only those few \*  
> language\* analyzers are missing) :
> 
> - Polish analyzer, with Stempel stemmer:  
> [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-stempel/index.html)
> - Smart Chinese analyzer, with 2 flavors:  
> [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-smartcn/index.html)
> - Latvian:  
> [org.apache.lucene.analysis.lv (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/lucene/analysis/lv/package-summary.html)
> 
> Some other findings:
> 
> - KpStemmer (Kraaij-Pohlmann stemming algorithm for Dutch[1][http://snowball.tartarus.org/algorithms/kraaij\_pohlmann/stemmer.html](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html)  
> ?):  
> [KpStemmer (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartarus/snowball/ext/KpStemmer.html)
> - LovinsStemmer, alike PorterStemmer, but less interesting according  
> to [2] [http://en.wikipedia.org/wiki/Stemming#History](http://en.wikipedia.org/wiki/Stemming#History):  
> [LovinsStemmer (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartarus/snowball/ext/LovinsStemmer.html)
> - Wikipedia-syntax-aware tokenizer:  
> [org.apache.lucene.analysis.wikipedia (Lucene 3.3.0 API)](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/lucene/analysis/wikipedia/package-summary.html)
> - WordNet synonym injector :  
> [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-wordnet/index.html)
> 
> The most interesting lists come from Solr itself:
> 
> - 
> 
> [http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-summary.html](http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-summary.html)
> 
> - [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis)
> 
> I didn't look very thoroughly at those last two links, but it looks that we  
> may be missing:
> 
> - ClassicTokenizer (may be deprecated or superseded by the  
> StandardTokenizer, I have no idea):  
> [https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucene/analysis/standard/ClassicTokenizer.html?is-external=true](https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucene/analysis/standard/ClassicTokenizer.html?is-external=true)
> - CommonGrams:  
> [http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGramsFilter.html](http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGramsFilter.html)
> - Lao, Myanmar, Khmer - seem to only split in syllables:  
> [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Lao.2C_Myanmar.2C_Khmer)
> 
> Towards a small easy pull-request?
> 
> ```
> [1]
> 
> ```
> 
> ## [The Kraaij-Pohlmann stemming algorithm](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html) Seen from [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Notes_about_solr.SnowballPorterFilterFactory) [2] [Stemming - Wikipedia](http://en.wikipedia.org/wiki/Stemming#History)
> 
> Olivier Favre
> 
> [www.yakaz.com](http://www.yakaz.com)
> 
> 2011/7/21 Paweł Konieczny [koniecznypw@gmail.com](mailto:koniecznypw@gmail.com)
> 
> > From what I understand, it's called Stempel and it's included in  
> > Lucene.
> > 
> > On Jul 20, 6:50 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Yes, you can hook your own analyzer, but you will need to implement a  
> > > custom  
> > > class that provides it. Check for example the GermanAnalyzerProvider.  
> > > What  
> > > is the name of the polish analyzer? I might have missed it and did not  
> > > include it out of the box.
> > > 
> > > 2011/7/20 Paweł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> > > 
> > > > Hey!
> > > 
> > > > Is it possible to import language analyzers from Lucene since ES is  
> > > > built on top of it (and Lucene has definitely more languages supported  
> > > > out of box)?  
> > > > The list of languages supported by EA is pretty extensive too, but it  
> > > > lacks polish language which I need 🙂  
> > > > EA seems much more flexible and has great potential (and I would like  
> > > > to use it in a project I'm developing), but without support for a  
> > > > given language it just won't do.  
> > > > Also, who adds language support to EA, its developers or the  
> > > > community?
> > > 
> > > > Cheers,  
> > > > Pawel

---

<div class="post-metadata">

**Author:** ![Pawel\_Konieczny](https://avatars.discourse-cdn.com/v4/letter/p/dc4da7/32.png) [@Pawel\_Konieczny](https://discuss.elastic.co/u/Pawel_Konieczny)\
**Post date:** [July 22, 2011, 7:52am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/6 "2011-07-22T07:52:46Z")

</div>

So will the next release have them included out of box? I'm in no  
hurry and I'd rather wait until someone does it properly.

Cheers,  
Pawel

On Jul 22, 2:44 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Heya,
> 
> Yea, open issues for the missing analyzers, the stempel one, for example,  
> should be simple to add (its in a different lib).
> 
> On Thu, Jul 21, 2011 at 5:55 PM, Olivier Favre [oliv...@yakaz.com](mailto:oliv...@yakaz.com) wrote:
> 
> > I think we should review all the available analyzers available, and  
> > identify the missing ones (ie not wrapped in ES).
> 
> > I also found that one on the Internet for Chinese:  
> > [Google Code Archive - Long-term storage for Google Code Project Hosting.](http://code.google.com/p/ik-analyzer/)  
> > With an ES plugin (at least a stub):  
> > [GitHub - medcl/elasticsearch at 21abad12a0096173e8836dd042ca403751ab7ad1](https://github.com/medcl/elasticsearch/blob/21abad12a0096173e8836dd04)...  
> > But it's not part of Lucene-contrib.
> 
> > Here is what I found, part of Lucene-contrib (apparently only those few \*  
> > language\* analyzers are missing) :
> 
> > - Polish analyzer, with Stempel stemmer:  
> > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-stempel/index.html)
> > - Smart Chinese analyzer, with 2 flavors:  
> > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-smartcn/index.html)
> > - Latvian:  
> > [Index of /\_\_root/docs.lucene.apache.org/core/3\_3\_0/api/contrib-analyzers/org/apache](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/)...
> 
> > Some other findings:
> 
> > - KpStemmer (Kraaij-Pohlmann stemming algorithm for Dutch[1][http://snowball.tartarus.org/algorithms/kraaij\_pohlmann/stemmer.html](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html)  
> > ?):  
> > [http://lucene.apache.org/java/3\_3\_0/api/contrib-analyzers/org/tartaru](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartaru)...
> > - LovinsStemmer, alike PorterStemmer, but less interesting according  
> > to [2] [http://en.wikipedia.org/wiki/Stemming#History](http://en.wikipedia.org/wiki/Stemming#History):  
> > [http://lucene.apache.org/java/3\_3\_0/api/contrib-analyzers/org/tartaru](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartaru)...
> > - Wikipedia-syntax-aware tokenizer:  
> > [Index of /\_\_root/docs.lucene.apache.org/core/3\_3\_0/api/contrib-analyzers/org/apache](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/)...
> > - WordNet synonym injector :  
> > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-wordnet/index.html)
> 
> > The most interesting lists come from Solr itself:
> 
> > - 
> > 
> > [http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-su](http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-su)...  
> > -[LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis)
> 
> > I didn't look very thoroughly at those last two links, but it looks that we  
> > may be missing:
> 
> > - ClassicTokenizer (may be deprecated or superseded by the  
> > StandardTokenizer, I have no idea):  
> > [https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucen](https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucen)...
> > - CommonGrams:  
> > [http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGram](http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGram)...
> > - Lao, Myanmar, Khmer - seem to only split in syllables:  
> > [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Lao.2C_Myanmar.2C_Khmer)
> 
> > Towards a small easy pull-request?
> 
> > ```
> > [1]
> > 
> > ```
> > 
> > ## [The Kraaij-Pohlmann stemming algorithm](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html) Seen from [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Notes_about_solr.Snowbal)... [2][Stemming - Wikipedia](http://en.wikipedia.org/wiki/Stemming#History)
> > 
> > Olivier Favre
> 
> > [www.yakaz.com](http://www.yakaz.com)
> 
> > 2011/7/21 Paweł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> 
> > > From what I understand, it's called Stempel and it's included in  
> > > Lucene.
> 
> > > On Jul 20, 6:50 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Yes, you can hook your own analyzer, but you will need to implement a  
> > > > custom  
> > > > class that provides it. Check for example the GermanAnalyzerProvider.  
> > > > What  
> > > > is the name of the polish analyzer? I might have missed it and did not  
> > > > include it out of the box.
> 
> > > > 2011/7/20 Paweł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> 
> > > > > Hey!
> 
> > > > > Is it possible to import language analyzers from Lucene since ES is  
> > > > > built on top of it (and Lucene has definitely more languages supported  
> > > > > out of box)?  
> > > > > The list of languages supported by EA is pretty extensive too, but it  
> > > > > lacks polish language which I need 🙂  
> > > > > EA seems much more flexible and has great potential (and I would like  
> > > > > to use it in a project I'm developing), but without support for a  
> > > > > given language it just won't do.  
> > > > > Also, who adds language support to EA, its developers or the  
> > > > > community?
> 
> > > > > Cheers,  
> > > > > Pawel

---

<div class="post-metadata">

**Author:** ![medcl\_net](https://avatars.discourse-cdn.com/v4/letter/m/90ced4/32.png) [@medcl\_net](https://discuss.elastic.co/u/medcl_net)\
**Post date:** [July 22, 2011, 11:05am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/7 "2011-07-22T11:05:17Z")

</div>

hey,i just write a post about how to customize an es plugin a few days  
ago,but in chinese~ ☹

[http://log.medcl.net/item/2011/07/diving-into-elasticsearch-3-编写自定义分词插件/](http://log.medcl.net/item/2011/07/diving-into-elasticsearch-3-%E7%BC%96%E5%86%99%E8%87%AA%E5%AE%9A%E4%B9%89%E5%88%86%E8%AF%8D%E6%8F%92%E4%BB%B6/)

-----Original Message-----  
From: PaweÂł Konieczny  
Sent: Friday, July 22, 2011 3:52 PM  
To: users  
Subject: Re: Importing language analyzers

So will the next release have them included out of box? I'm in no  
hurry and I'd rather wait until someone does it properly.

Cheers,  
Pawel

On Jul 22, 2:44 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Heya,
> 
> Yea, open issues for the missing analyzers, the stempel one, for  
> example,  
> should be simple to add (its in a different lib).
> 
> On Thu, Jul 21, 2011 at 5:55 PM, Olivier Favre [oliv...@yakaz.com](mailto:oliv...@yakaz.com) wrote:
> 
> > I think we should review all the available analyzers available, and  
> > identify the missing ones (ie not wrapped in ES).
> 
> > I also found that one on the Internet for Chinese:  
> > [Google Code Archive - Long-term storage for Google Code Project Hosting.](http://code.google.com/p/ik-analyzer/)  
> > With an ES plugin (at least a stub):  
> > [GitHub - medcl/elasticsearch at 21abad12a0096173e8836dd042ca403751ab7ad1](https://github.com/medcl/elasticsearch/blob/21abad12a0096173e8836dd04)...  
> > But it's not part of Lucene-contrib.
> 
> > Here is what I found, part of Lucene-contrib (apparently only those few  
> > \*  
> > language\* analyzers are missing) :
> 
> > - Polish analyzer, with Stempel stemmer:  
> > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-stempel/index.html)
> > - Smart Chinese analyzer, with 2 flavors:  
> > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-smartcn/index.html)
> > - Latvian:
> > 
> > [Index of /\_\_root/docs.lucene.apache.org/core/3\_3\_0/api/contrib-analyzers/org/apache](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/)...
> 
> > Some other findings:
> 
> > - KpStemmer (Kraaij-Pohlmann stemming algorithm for  
> > Dutch[1][http://snowball.tartarus.org/algorithms/kraaij\_pohlmann/stemmer.html](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html)  
> > ?):
> > 
> > [http://lucene.apache.org/java/3\_3\_0/api/contrib-analyzers/org/tartaru](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartaru)...
> > 
> > - LovinsStemmer, alike PorterStemmer, but less interesting according  
> > to [2] [http://en.wikipedia.org/wiki/Stemming#History](http://en.wikipedia.org/wiki/Stemming#History):
> > 
> > [http://lucene.apache.org/java/3\_3\_0/api/contrib-analyzers/org/tartaru](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartaru)...
> > 
> > - Wikipedia-syntax-aware tokenizer:
> > 
> > [Index of /\_\_root/docs.lucene.apache.org/core/3\_3\_0/api/contrib-analyzers/org/apache](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/)...
> > 
> > - WordNet synonym injector :  
> > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-wordnet/index.html)
> 
> > The most interesting lists come from Solr itself:
> 
> > - 
> > 
> > [http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-su](http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-su)...  
> > -[LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis)
> 
> > I didn't look very thoroughly at those last two links, but it looks that  
> > we  
> > may be missing:
> 
> > - ClassicTokenizer (may be deprecated or superseded by the  
> > StandardTokenizer, I have no idea):
> > 
> > [https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucen](https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucen)...
> > 
> > - CommonGrams:
> > 
> > [http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGram](http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGram)...
> > 
> > - Lao, Myanmar, Khmer - seem to only split in syllables:  
> > [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Lao.2C_Myanmar.2C_Khmer)
> 
> > Towards a small easy pull-request?
> 
> > ```
> > [1]
> > 
> > ```
> > 
> > ## [The Kraaij-Pohlmann stemming algorithm](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html) Seen from [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Notes_about_solr.Snowbal)... [2][Stemming - Wikipedia](http://en.wikipedia.org/wiki/Stemming#History)
> > 
> > Olivier Favre
> 
> > [www.yakaz.com](http://www.yakaz.com)
> 
> > 2011/7/21 PaweÂł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> 
> > > From what I understand, it's called Stempel and it's included in  
> > > Lucene.
> 
> > > On Jul 20, 6:50 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Yes, you can hook your own analyzer, but you will need to implement a  
> > > > custom  
> > > > class that provides it. Check for example the GermanAnalyzerProvider.  
> > > > What  
> > > > is the name of the polish analyzer? I might have missed it and did  
> > > > not  
> > > > include it out of the box.
> 
> > > > 2011/7/20 PaweÂł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> 
> > > > > Hey!
> 
> > > > > Is it possible to import language analyzers from Lucene since ES  
> > > > > is  
> > > > > built on top of it (and Lucene has definitely more languages  
> > > > > supported  
> > > > > out of box)?  
> > > > > The list of languages supported by EA is pretty extensive too, but  
> > > > > it  
> > > > > lacks polish language which I need 🙂  
> > > > > EA seems much more flexible and has great potential (and I would  
> > > > > like  
> > > > > to use it in a project I'm developing), but without support for a  
> > > > > given language it just won't do.  
> > > > > Also, who adds language support to EA, its developers or the  
> > > > > community?
> 
> > > > > Cheers,  
> > > > > Pawel

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 24, 2011, 2:17am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/8 "2011-07-24T02:17:09Z")

</div>

It can have it, sure, just make sure to open issues for the relevant  
analyzers.

2011/7/22 Paweł Konieczny [koniecznypw@gmail.com](mailto:koniecznypw@gmail.com)

> So will the next release have them included out of box? I'm in no  
> hurry and I'd rather wait until someone does it properly.
> 
> Cheers,  
> Pawel
> 
> On Jul 22, 2:44 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Heya,
> > 
> > Yea, open issues for the missing analyzers, the stempel one, for  
> > example,  
> > should be simple to add (its in a different lib).
> > 
> > On Thu, Jul 21, 2011 at 5:55 PM, Olivier Favre [oliv...@yakaz.com](mailto:oliv...@yakaz.com)  
> > wrote:
> > 
> > > I think we should review all the available analyzers available, and  
> > > identify the missing ones (ie not wrapped in ES).
> > 
> > > I also found that one on the Internet for Chinese:  
> > > [Google Code Archive - Long-term storage for Google Code Project Hosting.](http://code.google.com/p/ik-analyzer/)  
> > > With an ES plugin (at least a stub):  
> > > [GitHub - medcl/elasticsearch at 21abad12a0096173e8836dd042ca403751ab7ad1](https://github.com/medcl/elasticsearch/blob/21abad12a0096173e8836dd04).  
> > > ..  
> > > But it's not part of Lucene-contrib.
> > 
> > > Here is what I found, part of Lucene-contrib (apparently only those few
> 
> - 
> 
> > > language\* analyzers are missing) :
> > 
> > > - Polish analyzer, with Stempel stemmer:  
> > > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-stempel/index.html)
> > > - Smart Chinese analyzer, with 2 flavors:  
> > > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-smartcn/index.html)
> > > - Latvian:
> 
> [Index of /\_\_root/docs.lucene.apache.org/core/3\_3\_0/api/contrib-analyzers/org/apache](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/)...
> 
> > > Some other findings:
> > 
> > > - KpStemmer (Kraaij-Pohlmann stemming algorithm for Dutch[1]\<  
> > > [The Kraaij-Pohlmann stemming algorithm](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html)\>  
> > > ?):
> 
> [http://lucene.apache.org/java/3\_3\_0/api/contrib-analyzers/org/tartaru](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartaru)...
> 
> > > - LovinsStemmer, alike PorterStemmer, but less interesting according  
> > > to [2] [http://en.wikipedia.org/wiki/Stemming#History](http://en.wikipedia.org/wiki/Stemming#History):
> 
> [http://lucene.apache.org/java/3\_3\_0/api/contrib-analyzers/org/tartaru](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/tartaru)...
> 
> > > - Wikipedia-syntax-aware tokenizer:
> 
> [Index of /\_\_root/docs.lucene.apache.org/core/3\_3\_0/api/contrib-analyzers/org/apache](http://lucene.apache.org/java/3_3_0/api/contrib-analyzers/org/apache/)...
> 
> > > - WordNet synonym injector :  
> > > [Lucene 3.3.0 API](http://lucene.apache.org/java/3_3_0/api/contrib-wordnet/index.html)
> > 
> > > The most interesting lists come from Solr itself:
> > 
> > > -
> 
> [http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-su](http://lucene.apache.org/solr/api/org/apache/solr/analysis/package-su)...
> 
> > > -[LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis)
> > 
> > > I didn't look very thoroughly at those last two links, but it looks  
> > > that we  
> > > may be missing:
> > 
> > > - ClassicTokenizer (may be deprecated or superseded by the  
> > > StandardTokenizer, I have no idea):
> 
> [https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucen](https://builds.apache.org/job/Lucene-3.x/javadoc/all/org/apache/lucen)...
> 
> > > - CommonGrams:
> 
> [http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGram](http://lucene.apache.org/solr/api/org/apache/solr/analysis/CommonGram)...
> 
> > > - Lao, Myanmar, Khmer - seem to only split in syllables:
> 
> [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Lao.2C_Myanmar.2C_Khmer)
> 
> > > Towards a small easy pull-request?
> > 
> > > ```
> > > [1]
> > > 
> > > ```
> > > 
> > > ## [The Kraaij-Pohlmann stemming algorithm](http://snowball.tartarus.org/algorithms/kraaij_pohlmann/stemmer.html) Seen from [LanguageAnalysis - Solr - Apache Software Foundation](http://wiki.apache.org/solr/LanguageAnalysis#Notes_about_solr.Snowbal). .. [2][Stemming - Wikipedia](http://en.wikipedia.org/wiki/Stemming#History)
> > > 
> > > Olivier Favre
> > 
> > > [www.yakaz.com](http://www.yakaz.com)
> > 
> > > 2011/7/21 Paweł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> > 
> > > > From what I understand, it's called Stempel and it's included in  
> > > > Lucene.
> > 
> > > > On Jul 20, 6:50 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > 
> > > > > Yes, you can hook your own analyzer, but you will need to implement  
> > > > > a  
> > > > > custom  
> > > > > class that provides it. Check for example the  
> > > > > GermanAnalyzerProvider.  
> > > > > What  
> > > > > is the name of the polish analyzer? I might have missed it and did  
> > > > > not  
> > > > > include it out of the box.
> > 
> > > > > 2011/7/20 Paweł Konieczny [konieczn...@gmail.com](mailto:konieczn...@gmail.com)
> > 
> > > > > > Hey!
> > 
> > > > > > Is it possible to import language analyzers from Lucene since ES  
> > > > > > is  
> > > > > > built on top of it (and Lucene has definitely more languages  
> > > > > > supported  
> > > > > > out of box)?  
> > > > > > The list of languages supported by EA is pretty extensive too, but  
> > > > > > it  
> > > > > > lacks polish language which I need 🙂  
> > > > > > EA seems much more flexible and has great potential (and I would  
> > > > > > like  
> > > > > > to use it in a project I'm developing), but without support for a  
> > > > > > given language it just won't do.  
> > > > > > Also, who adds language support to EA, its developers or the  
> > > > > > community?
> > 
> > > > > > Cheers,  
> > > > > > Pawel

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:59am UTC](https://discuss.elastic.co/t/importing-language-analyzers/4903/9 "2017-07-06T03:59:41Z")

</div>


