# MultiLingual Index

**URL:** <https://discuss.elastic.co/t/multilingual-index/38150>\
**Category:** Elasticsearch\
**Created:** [December 30, 2015, 5:26am UTC](https://discuss.elastic.co/t/multilingual-index/38150 "2015-12-30T05:26:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![hash\_include](https://avatars.discourse-cdn.com/v4/letter/h/b782af/32.png) [@hash\_include](https://discuss.elastic.co/u/hash_include)\
**Post date:** [December 30, 2015, 5:26am UTC](https://discuss.elastic.co/t/multilingual-index/38150/1 "2015-12-30T05:26:01Z")

</div>

Hi All

I have document corpus with few documents in chinese, few in German and others in english. I cannot create multiple indexes based on languages owing to current infrastructure.

I need to have one index with multiple analyzers on fields.  
My current thoughts:  
If there is a field "title" then we need to have title.german(german analyzer), title.chinese(cjk analyzer), title.english(english analyzer), title.general (standard). But this approach will have all documents analyzed in all possible analyzers bloating up the index size and index time. Is there a way to apply specific analyzers to specific documents based on language field?.

I am looking into ICUFolding and other aspects of multilingual search as well. Please guide me in this regards.

Thanks  
Sri Harsha

---

<div class="post-metadata">

**Author:** ![loren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loren/32/44942_2.png) [@loren](https://discuss.elastic.co/u/loren)\
**Post date:** [December 30, 2015, 6:55pm UTC](https://discuss.elastic.co/t/multilingual-index/38150/2 "2015-12-30T18:55:50Z")

</div>

I ran into the same problem. The way I went about it was based on the [One Language per Field](https://www.elastic.co/guide/en/elasticsearch/guide/current/one-lang-fields.html) approach. I use a custom serializer to look at the document language at index time and then copy the `title` field over to a `title_#{language}` field. So a French document would end up with `title` and `title_fr` fields. I set up the index template to use a French analyzer for `*_fr` fields, and so on. For search, I use both fields to influence the score.

It's a similar approach to what you were suggesting, but here you only end up with one extra field per document.

If it's helpful, the code that handles all of this is part of [this project](https://github.com/GSA/i14y). Relevant files are [here](https://github.com/GSA/i14y/blob/master/app/models/document.rb), [here](https://github.com/GSA/i14y/blob/master/lib/serde.rb), and [here](https://github.com/GSA/i14y/blob/master/app/templates/documents.rb).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 30, 2015, 8:23pm UTC](https://discuss.elastic.co/t/multilingual-index/38150/3 "2015-12-30T20:23:46Z")

</div>

> [@hash\_include](#):
>
> Is there a way to apply specific analyzers to specific documents based on language field?

Not any more.

In 1.x, you could select the analyzer from a path. So, you could index the language code based on your input, and the analyzer would be automatically set to german, english, 中文, whatever.

In 2.x this feature was removed.

> **[\_analyzer | Elasticsearch Guide \[2.1\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/2.1/mapping-analyzer-field.html)**

Maybe I can find a trick to implement this again in my language detection plugin.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:27pm UTC](https://discuss.elastic.co/t/multilingual-index/38150/4 "2017-07-05T23:27:44Z")

</div>


