# Best way to index multiple languages

**URL:** <https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410>\
**Category:** Elasticsearch\
**Created:** [January 17, 2012, 1:22pm UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410 "2012-01-17T13:22:52Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexandre/32/3014_2.png) [@Alexandre](https://discuss.elastic.co/u/Alexandre)\
**Post date:** [January 17, 2012, 1:22pm UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/1 "2012-01-17T13:22:52Z")

</div>

Hi,

I am trying to figure out gow to index the following in ES.

I have documents representing location names ([geonames.org](http://geonames.org)). Each  
document has a category such as Airport, restaurant, river, beach  
etc ..  
I have translated the categories in 25 languages. For each document I  
want to add the translated categories.

Should I :

1. create 1 field per category translation and use a different  
analyzer with different language for each categroy field ?
2. create 1 general category object field with an array of  
translations ... in that case how should I set analysis ?
3. do it some other way I am not aware of 🙂 ?

Bonus question : I also need to do multi language querying ... so a  
french person will query "Aéroport de Genève" but an english person  
will query "Geneva Airport" and chinese person will query "" in their own language etc ... Is there  
something special I need to do to build the query so that we get the  
best combination of location name and category hit?

Many thanks for your answers !

---

<div class="post-metadata">

**Author:** ![Jan\_Fiedler](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jan_fiedler/32/2518_2.png) [@Jan\_Fiedler](https://discuss.elastic.co/u/Jan_Fiedler)\
**Post date:** [January 17, 2012, 4:20pm UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/2 "2012-01-17T16:20:50Z")

</div>

No real production experience yet but I am using the approach of having one  
field per language (with language specific analyzer configurations attached  
to them).

For the bonus question: Often you will have some context in your app that  
would define the language (e.g. user selecting language for their browsing  
session as the remaining page content most likely will have to show the  
correct language too). In this case it would be trivial to select the  
correct field in the query. Without session context you could use language  
detection on the user input and select the correct field based on that. If  
you do not have a language detection library, you could try to run the  
search across all language fields (this may generate some noise though).

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [January 18, 2012, 5:50am UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/3 "2012-01-18T05:50:16Z")

</div>

Alexandre,

Check the ML archive, I just asked a similar question the other day  
and the approach we'll be taking is the one where we have a single set  
of fields and at index-time we explicitly specify the analyzer for  
each field depending on the field or document language. We'll be  
using our own Language Identifier library for that (http://  
[Cloud Monitoring Tools & Services | Sematext](http://sematext.com/products/language-identifier/index.html)). You could use  
a Language Identifier to detect query language, too. Precision may  
suffer if queries are very short or ambiguous (is "die" an English  
verb? Or English noun? Or a German article?), though this can be  
addressed through UI, giving people options to select from one or a  
few guessed languages, allowing people to permanently store/remember  
their language selection and such.

## Otis

Sematext is hiring Elasticsearch / Solr developers --

> **[Jobs](https://sematext.com/jobs/)**
>
> We’re Hiring We are always looking for smart, passionate, motivated, and independent people regardless of where on the planet they may be. Learn more about the company Agent & Backend Engineer Full Stack Developer Backend Engineer Frontend...

On Jan 17, 11:20 am, Jan Fiedler [fiedler....@gmail.com](mailto:fiedler....@gmail.com) wrote:

> No real production experience yet but I am using the approach of having one  
> field per language (with language specific analyzer configurations attached  
> to them).
> 
> For the bonus question: Often you will have some context in your app that  
> would define the language (e.g. user selecting language for their browsing  
> session as the remaining page content most likely will have to show the  
> correct language too). In this case it would be trivial to select the  
> correct field in the query. Without session context you could use language  
> detection on the user input and select the correct field based on that. If  
> you do not have a language detection library, you could try to run the  
> search across all language fields (this may generate some noise though).

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 18, 2012, 9:06pm UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/4 "2012-01-18T21:06:58Z")

</div>

It really depends, there are several ways to do it. You can create an index  
per language, with its own mapping that has lang analyzer specific on the  
relevant field. Another option is to use multiple field names for each  
language, each with its own analyzer associated with it. Those are usually  
the best two options.

On Tue, Jan 17, 2012 at 6:20 PM, Jan Fiedler [fiedler.jan@gmail.com](mailto:fiedler.jan@gmail.com) wrote:

> No real production experience yet but I am using the approach of having  
> one field per language (with language specific analyzer configurations  
> attached to them).
> 
> For the bonus question: Often you will have some context in your app that  
> would define the language (e.g. user selecting language for their browsing  
> session as the remaining page content most likely will have to show the  
> correct language too). In this case it would be trivial to select the  
> correct field in the query. Without session context you could use language  
> detection on the user input and select the correct field based on that. If  
> you do not have a language detection library, you could try to run the  
> search across all language fields (this may generate some noise though).

---

<div class="post-metadata">

**Author:** ![Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexandre/32/3014_2.png) [@Alexandre](https://discuss.elastic.co/u/Alexandre)\
**Post date:** [January 21, 2012, 1:35pm UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/5 "2012-01-21T13:35:47Z")

</div>

Many thanks to all for your insights !

I started working through the field per language approach. I guess  
I'll be using variants of snowball analyzer for each language when  
available and a regular LanguageAnalyzer when not.

For the query part I can play both with a user setting and query  
language identification (whenever possible).

You were really helpful !

Cheers!

Alex

---

<div class="post-metadata">

**Author:** ![Jussi\_Arpalahti](https://avatars.discourse-cdn.com/v4/letter/j/f4b2a3/32.png) [@Jussi\_Arpalahti](https://discuss.elastic.co/u/Jussi_Arpalahti)\
**Post date:** [January 22, 2012, 10:41am UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/6 "2012-01-22T10:41:44Z")

</div>

On 21 January 2012 15:35, Alexandre [azlist1@gmail.com](mailto:azlist1@gmail.com) wrote:

> Many thanks to all for your insights !
> 
> I started working through the field per language approach. I guess  
> I'll be using variants of snowball analyzer for each language when  
> available and a regular LanguageAnalyzer when not.
> 
> For the query part I can play both with a user setting and query  
> language identification (whenever possible).
> 
> Hi.

Perhaps this is only relevant for a language like finnish, but we have  
indexed two versions of every field. One type with no language analyzing  
for exact matches and the other with language analyzer for expanded matches.

Documents are build like this:  
title: "some text" -\> to default analyzer  
title\_en: "some text" -\> to english analyzer  
language: english

When searching the language independent field is given a slight boost over  
the linguistically analyzed field. Our documents only containt text in one  
language and are indexed using a language field. Searches are then filtered  
by this field on the user's chosen language.

FYI, finnish is a somewhat complex language to search for. As wikipedia  
says "it modifies inflects [http://en.wikipedia.org/wiki/Inflection](http://en.wikipedia.org/wiki/Inflection) the  
forms of nouns [http://en.wikipedia.org/wiki/Noun](http://en.wikipedia.org/wiki/Noun),  
adjectives[http://en.wikipedia.org/wiki/Adjective](http://en.wikipedia.org/wiki/Adjective),  
pronouns [http://en.wikipedia.org/wiki/Pronoun](http://en.wikipedia.org/wiki/Pronoun),  
numerals[http://en.wikipedia.org/wiki/Number\_names](http://en.wikipedia.org/wiki/Number_names)and  
verbs [http://en.wikipedia.org/wiki/Verb](http://en.wikipedia.org/wiki/Verb), depending on their roles in the  
sentence [http://en.wikipedia.org/wiki/Sentence\_(linguistics)](http://en.wikipedia.org/wiki/Sentence_%28linguistics%29)." Thus  
we need to index the word as it is in the document but also in its base  
form so user does not have to match the dozens of variations of the word.  
We have also used this indexing strategy for english and swedish words.  
However we don't yet have enough experience with the search service to say  
if this is a good approach for these languages.

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [January 23, 2012, 7:53pm UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/7 "2012-01-23T19:53:22Z")

</div>

Hi Shay,

There is also the option of specifying an Analyzer for each field at  
index-time, right?  
Are there some drawback to this approach that makes you say that index  
per language and field set per language are usually the best the  
options?

## Thanks, Otis

Sematext is hiring Elasticsearch / Solr developers - [Jobs](http://sematext.com/about/jobs.html)

On Jan 18, 4:06 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> It really depends, there are several ways to do it. You can create an index  
> per language, with its own mapping that has lang analyzer specific on the  
> relevant field. Another option is to use multiple field names for each  
> language, each with its own analyzer associated with it. Those are usually  
> the best two options.
> 
> On Tue, Jan 17, 2012 at 6:20 PM, Jan Fiedler [fiedler....@gmail.com](mailto:fiedler....@gmail.com) wrote:
> 
> > No real production experience yet but I am using the approach of having  
> > one field per language (with language specific analyzer configurations  
> > attached to them).
> 
> > For the bonus question: Often you will have some context in your app that  
> > would define the language (e.g. user selecting language for their browsing  
> > session as the remaining page content most likely will have to show the  
> > correct language too). In this case it would be trivial to select the  
> > correct field in the query. Without session context you could use language  
> > detection on the user input and select the correct field based on that. If  
> > you do not have a language detection library, you could try to run the  
> > search across all language fields (this may generate some noise though).

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 23, 2012, 8:52pm UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/8 "2012-01-23T20:52:05Z")

</div>

Its just the fact that a field will now have its terms produced by  
different analyzers. Can certainly be used, but I would prefer to separate  
it.

On Mon, Jan 23, 2012 at 9:53 PM, Otis Gospodnetic \<  
[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

> Hi Shay,
> 
> There is also the option of specifying an Analyzer for each field at  
> index-time, right?  
> Are there some drawback to this approach that makes you say that index  
> per language and field set per language are usually the best the  
> options?
> 
> ## Thanks, Otis
> 
> Sematext is hiring Elasticsearch / Solr developers -  
> [Jobs - Sematext](http://sematext.com/about/jobs.html)
> 
> On Jan 18, 4:06 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > It really depends, there are several ways to do it. You can create an  
> > index  
> > per language, with its own mapping that has lang analyzer specific on the  
> > relevant field. Another option is to use multiple field names for each  
> > language, each with its own analyzer associated with it. Those are  
> > usually  
> > the best two options.
> > 
> > On Tue, Jan 17, 2012 at 6:20 PM, Jan Fiedler [fiedler....@gmail.com](mailto:fiedler....@gmail.com)  
> > wrote:
> > 
> > > No real production experience yet but I am using the approach of having  
> > > one field per language (with language specific analyzer configurations  
> > > attached to them).
> > 
> > > For the bonus question: Often you will have some context in your app  
> > > that  
> > > would define the language (e.g. user selecting language for their  
> > > browsing  
> > > session as the remaining page content most likely will have to show the  
> > > correct language too). In this case it would be trivial to select the  
> > > correct field in the query. Without session context you could use  
> > > language  
> > > detection on the user input and select the correct field based on  
> > > that. If  
> > > you do not have a language detection library, you could try to run the  
> > > search across all language fields (this may generate some noise  
> > > though).

---

<div class="post-metadata">

**Author:** ![Mark\_Waddle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_waddle/32/2608_2.png) [@Mark\_Waddle](https://discuss.elastic.co/u/Mark_Waddle)\
**Post date:** [January 24, 2012, 3:49am UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/9 "2012-01-24T03:49:02Z")

</div>

I agree with Shay. I would think that having separate indices would be  
better, especially if you ever plan to use the \_all field[http://www.elasticsearch.org/guide/reference/mapping/all-field.html](http://www.elasticsearch.org/guide/reference/mapping/all-field.html)  
.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:41am UTC](https://discuss.elastic.co/t/best-way-to-index-multiple-languages/6410/10 "2017-07-06T03:41:44Z")

</div>


