# Supporting as many languages as possible

**URL:** <https://discuss.elastic.co/t/supporting-as-many-languages-as-possible/14681>\
**Category:** Elasticsearch\
**Created:** [December 3, 2013, 3:43pm UTC](https://discuss.elastic.co/t/supporting-as-many-languages-as-possible/14681 "2013-12-03T15:43:11Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [December 3, 2013, 3:43pm UTC](https://discuss.elastic.co/t/supporting-as-many-languages-as-possible/14681/1 "2013-12-03T15:43:11Z")

</div>

tl/dr: asking for input on the "right" way to support lots of languages

The time has come for me to support more languages! I have some ideas  
about how to do this but I'd love some advice. Background: my install base  
has ~300 languages and my searching supports the concept of both "plain"  
and "aggressive" analyzers. Each index only supports a single language.  
My thoughts:

For languages for which I don't have a special case I'll make a single  
analyzer:  
{  
"type": "custom",  
"tokenizer": "standard",  
"filter": ["standard", "icu\_normalizer", "lowercase"]  
}

For languages that have a stemmer I'll use the that analyzer above as a  
"plain" analyzer and something like this (custom per language) for the  
"aggressive" analyzer:  
{  
"type": "custom",  
"tokenizer": "standard",  
"filter": [ "standard", "icu\_normalizer", "possessive\_english",  
"lowercase", "stop", "kstem", "asciifolding" ]  
}

Not all languages will want asciifolding (but we're used to it in English)  
and not all languages have stop words. Bonus: Some of my install base uses  
word\_delimiter in the aggressive analyzer as well!

For some languages I think I'll need to replace the plain analyzer with a  
weakened version of the custom analyzer - Japanese will need the kuromoji  
tokenizer, for example.

Questions:  
Does this make sense?  
Should I spend time investigating using the icu\_tokenizer instead of the  
standard tokenizer?  
Are there any analysis plugins that I should look beyond ICU, Smart-CN,  
Stempel, and Kuromoji?  
I saw some talk about a Hebrew plugin but that plugin isn't listed on  
Elasticsearch's plugins page. Is it useful/ready?

Assertion:  
I'm happy to use any plugin so long as it has some open source license, is  
actively supported by someone who speaks the language, and has instructions  
in English. I assume plugins always have instructions in their native  
language, but I need some in English too.

Thanks for reading! Please tell me all the mistakes I'm about to make!

Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd0R8DG%3DOUtPzW\_277%3DSChykrmm0U5rZu\_p8-qJB9id2zQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd0R8DG%3DOUtPzW_277%3DSChykrmm0U5rZu_p8-qJB9id2zQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:03am UTC](https://discuss.elastic.co/t/supporting-as-many-languages-as-possible/14681/2 "2017-07-06T02:03:36Z")

</div>


