# Using the Snowball stemmers

**URL:** <https://discuss.elastic.co/t/using-the-snowball-stemmers/3685>\
**Category:** Elasticsearch\
**Created:** [December 21, 2010, 2:03pm UTC](https://discuss.elastic.co/t/using-the-snowball-stemmers/3685 "2010-12-21T14:03:02Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Joaquin\_Cuenca\_Abela](https://avatars.discourse-cdn.com/v4/letter/j/e480ec/32.png) [@Joaquin\_Cuenca\_Abela](https://discuss.elastic.co/u/Joaquin_Cuenca_Abela)\
**Post date:** [December 21, 2010, 2:03pm UTC](https://discuss.elastic.co/t/using-the-snowball-stemmers/3685/1 "2010-12-21T14:03:02Z")

</div>

Hi,

I want to index some spanish text, and it seems that I need to use  
some external stemmer to do this, as ES/Lucene doesn't include by  
default any stemmer for spanish.

Is there any docs on how to use an external stemmer (for instance the  
Snowball ones) with ElasticSearch?

Thanks!

--  
Joaquin Cuenca Abela

---

<div class="post-metadata">

**Author:** ![Sebastian\_Gavarini](https://avatars.discourse-cdn.com/v4/letter/s/db5fbb/32.png) [@Sebastian\_Gavarini](https://discuss.elastic.co/u/Sebastian_Gavarini)\
**Post date:** [December 21, 2010, 5:46pm UTC](https://discuss.elastic.co/t/using-the-snowball-stemmers/3685/2 "2010-12-21T17:46:57Z")

</div>

Hi Joaquin,

If I remember correctly Lucene includes the Snowball family of  
stemmers in it's contrib package. PorterStemmer is included in ES but  
it's English only. I'll try to draft an exmaple for you, I would  
suggest that you look at how PorterStemmer is included in ES and  
create similar classes to use Lucene's Snowball, take a look at the  
class:

- PorterStemTokenFilterFactory: it's the factory responsible of  
creating and configuring the stemmer

Create a similar factory, eg: SnowballTokenFilterFactory

package org.elasticsearch.index.analysis;  
imports....  
public class SnowballTokenFilterFactory extends  
AbstractTokenFilterFactory {

```
private String language;

@Inject public SnowballTokenFilterFactory(Index index,

```

@IndexSettings Settings indexSettings, @Assisted String name,  
@Assisted Settings settings) {  
super(index, indexSettings, name);  
this.language = settings.get("language");  
}

```
@Override public TokenStream create(TokenStream tokenStream) {
    return new SnowballFilter(tokenStream, language);
}

```

}

I haven't tried it, but it should work pretty much like that, with the  
correct imports. It's important (or at least it was when I looked into  
it sometime ago) to use the package "org.elasticsearch.\*" for these  
factories. Then the variable "language" should appear in  
elasticsearch.yml, like:

index:  
analysis:  
analyzer:  
my\_analyzer:  
type: custom  
tokenizer: whitespace  
filter: [lowercase, asciifolding, snowball]  
filter:  
snowball:  
type:  
org.elasticsearch.index.analysis.SnowballTokenFilterFactory  
language: Spanish

You need to include your custom classes, in this case  
SnowballTokenFilterFactory, in a jar in the lib directory of ES.

Please give it a try and let me know if you have some problems.

Regards,  
Sebastian.

On Dec 21, 11:03 am, Joaquin Cuenca Abela [joaq...@cuencaabela.com](mailto:joaq...@cuencaabela.com)  
wrote:

> Hi,
> 
> I want to index some spanish text, and it seems that I need to use  
> some external stemmer to do this, as ES/Lucene doesn't include by  
> default any stemmer for spanish.
> 
> Is there any docs on how to use an external stemmer (for instance the  
> Snowball ones) with Elasticsearch?
> 
> Thanks!
> 
> --  
> Joaquin Cuenca Abela

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:15am UTC](https://discuss.elastic.co/t/using-the-snowball-stemmers/3685/3 "2017-07-06T04:15:02Z")

</div>


