# Use case of multiple Language Analyzer, Hunspell along with Elasticsearch Langdetect Plugin

**URL:** <https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938>\
**Category:** Elasticsearch\
**Created:** [September 24, 2014, 10:57am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938 "2014-09-24T10:57:12Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Prashant\_Pal](https://avatars.discourse-cdn.com/v4/letter/p/b77776/32.png) [@Prashant\_Pal](https://discuss.elastic.co/u/Prashant_Pal)\
**Post date:** [September 24, 2014, 10:57am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/1 "2014-09-24T10:57:12Z")

</div>

Hi All,

We are having an ES cluster which is used to index large amount of data and that too with different languages. So as of now our current settings was pointing to English analyzer, and English hunspell but how we can achieve to index multilingual data along with Multi lingual analyzer and hunspell setup for same index (as I came across like there is a plugin called Elasticsearch Langdetect Plugin ([https://github.com/jprante/elasticsearch-langdetect](https://github.com/jprante/elasticsearch-langdetect)) available from ES 1.2.1).

Current analyzer setting is like:  
index :  
analysis :  
analyzer :  
synonym :  
tokenizer : whitespace  
filter : [synonym]  
default\_index :  
type : custom  
tokenizer : whitespace  
filter : [standard, lowercase,hunspell\_US]  
default\_search :  
type : custom  
tokenizer : whitespace  
filter : [standard, lowercase, synonym,hunspell\_US]  
filter :  
synonym :  
type : synonym  
ignore\_case : true  
expand : true  
synonyms\_path : synonyms.txt  
hunspell\_US :  
type : hunspell  
locale : en\_US  
dedup : false  
ignore\_case : true

So here,

1. Can we configure multilingual analyzer, hunspell for same index and then index data by configuring lang detect plugin for specific fields. So here whether data will be indexed and analyzed as per the language analyzer mentioned? And also will it be searchable as per multiple hunspell dictionaries and synonyms configured as well:

Confirm if below settings can be ued to achieve the same:

```
  analyzer :
    synonym : 
        tokenizer : whitespace
        filter : [synonym]
    default_index :
        type : custom
        tokenizer : whitespace
        filter : [standard, lowercase,hunspell_US,hunspell_IN,hindi,english]  
    default_search :
        type : custom
        tokenizer : whitespace
        filter : [standard, lowercase, synonym,hunspell_US,hunspell_IN,hindi,english]  
  filter :
    hindi: 
      tokenizer: standard
      filter: [lowercase]
    english: 
      tokenizer: standard
      filter: [lowercase]
    synonym : 
        type : synonym
        ignore_case : true
        expand : true
        synonyms_path : synonyms.txt
    hunspell_US :
        type : hunspell
        locale : en_US 
        dedup : false
		ignore_case : true
    hunspell_IN :
        type : hunspell
        locale : hi_IN 
        dedup : false
		ignore_case : true

```

After that Say, I have configured lang detect plugin and indexed some data with different language English and hindi. So as I have configured multiple language analyzer, MultiLingual hunspell so will I be able to perform the index and search wrt different language as with different analyzer and get the data as per analyzed tokens for different languages.

Also whether synonym will also work with different languages?

~Prashant

---

<div class="post-metadata">

**Author:** ![Prashant\_Pal](https://avatars.discourse-cdn.com/v4/letter/p/b77776/32.png) [@Prashant\_Pal](https://discuss.elastic.co/u/Prashant_Pal)\
**Post date:** [September 24, 2014, 10:59am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/2 "2014-09-24T10:59:25Z")

</div>

Hi Jorg,

Can you help me out in this as I found you are the owner of lang detect plugin.

~Prashant

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 24, 2014, 3:55pm UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/3 "2014-09-24T15:55:18Z")

</div>

With langdetect plugin, there is ujst a field "lang" mapped under the  
string field that is used for detection, and in this field the languages  
codes are written. This is useful for e.g. aggregations or filtering  
documents by language.

At the moment it is not possible to use something like for example a  
dynamic "copy\_to" to duplicate the field after detection to a field with  
language-specific analyzer like a synonym analyzer.

A feature request at the issue tracker at github is much appreciated so I  
can have a look into this.

Jörg

On Wed, Sep 24, 2014 at 12:57 PM, Prashant Agrawal \<  
[prashant.agrawal@paladion.net](mailto:prashant.agrawal@paladion.net)\> wrote:

> Hi All,
> 
> We are having an ES cluster which is used to index large amount of data and  
> that too with different languages. So as of now our current settings was  
> pointing to English analyzer, and English hunspell but how we can achieve  
> to  
> index multilingual data along with Multi lingual analyzer and hunspell  
> setup  
> for same index (as I came across like there is a plugin called  
> Elasticsearch  
> Langdetect Plugin ([GitHub - jprante/elasticsearch-langdetect: A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector](https://github.com/jprante/elasticsearch-langdetect))  
> available from ES 1.2.1).
> 
> Current analyzer setting is like:  
> index :  
> analysis :  
> analyzer :  
> synonym :  
> tokenizer : whitespace  
> filter : [synonym]  
> default\_index :  
> type : custom  
> tokenizer : whitespace  
> filter : [standard, lowercase,hunspell\_US]  
> default\_search :  
> type : custom  
> tokenizer : whitespace  
> filter : [standard, lowercase, synonym,hunspell\_US]  
> filter :  
> synonym :  
> type : synonym  
> ignore\_case : true  
> expand : true  
> synonyms\_path : synonyms.txt  
> hunspell\_US :  
> type : hunspell  
> locale : en\_US  
> dedup : false  
> ignore\_case : true
> 
> So here,
> 
> 1. Can we configure multilingual analyzer, hunspell for same index and then  
> index data by configuring lang detect plugin for specific fields. So here  
> whether data will be indexed and analyzed as per the language analyzer  
> mentioned? And also will it be searchable as per multiple hunspell  
> dictionaries and synonyms configured as well:
> 
> Confirm if below settings can be ued to achieve the same:
> 
> ```
> analyzer :
> synonym :
> tokenizer : whitespace
> filter : [synonym]
> default_index :
> type : custom
> tokenizer : whitespace
> filter : [ standard,
> 
> ```
> 
> lowercase,hunspell\_US,hunspell\_IN,hindi,english]  
> default\_search :  
> type : custom  
> tokenizer : whitespace  
> filter : [standard, lowercase,  
> synonym,hunspell\_US,hunspell\_IN,hindi,english]  
> filter :  
> hindi:  
> tokenizer: standard  
> filter: [lowercase]  
> english:  
> tokenizer: standard  
> filter: [lowercase]  
> synonym :  
> type : synonym  
> ignore\_case : true  
> expand : true  
> synonyms\_path : synonyms.txt  
> hunspell\_US :  
> type : hunspell  
> locale : en\_US  
> dedup : false  
> ignore\_case : true  
> hunspell\_IN :  
> type : hunspell  
> locale : hi\_IN  
> dedup : false  
> ignore\_case : true
> 
> After that Say, I have configured lang detect plugin and indexed some data  
> with different language English and hindi. So as I have configured multiple  
> language analyzer, MultiLingual hunspell so will I be able to perform the  
> index and search wrt different language as with different analyzer and get  
> the data as per analyzed tokens for different languages.
> 
> Also whether synonym will also work with different languages?
> 
> ~Prashant
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html](http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/1411556233114-4063950.post%40n3.nabble.com](https://groups.google.com/d/msgid/elasticsearch/1411556233114-4063950.post%40n3.nabble.com)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGZDhBwBRhNdF42QZg2wiJPTwzsryCdFMLc%2BfkBesZNeA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGZDhBwBRhNdF42QZg2wiJPTwzsryCdFMLc%2BfkBesZNeA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 24, 2014, 3:56pm UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/4 "2014-09-24T15:56:38Z")

</div>

The issue tracker address:

> **[Issues · jprante/elasticsearch-langdetect](https://github.com/jprante/elasticsearch-langdetect/issues)**
>
> A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector - Issues · jprante/elasticsearch-langdetect

Jörg

On Wed, Sep 24, 2014 at 5:55 PM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<  
[joergprante@gmail.com](mailto:joergprante@gmail.com)\> wrote:

> With langdetect plugin, there is ujst a field "lang" mapped under the  
> string field that is used for detection, and in this field the languages  
> codes are written. This is useful for e.g. aggregations or filtering  
> documents by language.
> 
> At the moment it is not possible to use something like for example a  
> dynamic "copy\_to" to duplicate the field after detection to a field with  
> language-specific analyzer like a synonym analyzer.
> 
> A feature request at the issue tracker at github is much appreciated so I  
> can have a look into this.
> 
> Jörg
> 
> On Wed, Sep 24, 2014 at 12:57 PM, Prashant Agrawal \<  
> [prashant.agrawal@paladion.net](mailto:prashant.agrawal@paladion.net)\> wrote:
> 
> > Hi All,
> > 
> > We are having an ES cluster which is used to index large amount of data  
> > and  
> > that too with different languages. So as of now our current settings was  
> > pointing to English analyzer, and English hunspell but how we can achieve  
> > to  
> > index multilingual data along with Multi lingual analyzer and hunspell  
> > setup  
> > for same index (as I came across like there is a plugin called  
> > Elasticsearch  
> > Langdetect Plugin ([GitHub - jprante/elasticsearch-langdetect: A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector](https://github.com/jprante/elasticsearch-langdetect))  
> > available from ES 1.2.1).
> > 
> > Current analyzer setting is like:  
> > index :  
> > analysis :  
> > analyzer :  
> > synonym :  
> > tokenizer : whitespace  
> > filter : [synonym]  
> > default\_index :  
> > type : custom  
> > tokenizer : whitespace  
> > filter : [standard, lowercase,hunspell\_US]  
> > default\_search :  
> > type : custom  
> > tokenizer : whitespace  
> > filter : [standard, lowercase, synonym,hunspell\_US]  
> > filter :  
> > synonym :  
> > type : synonym  
> > ignore\_case : true  
> > expand : true  
> > synonyms\_path : synonyms.txt  
> > hunspell\_US :  
> > type : hunspell  
> > locale : en\_US  
> > dedup : false  
> > ignore\_case : true
> > 
> > So here,
> > 
> > 1. Can we configure multilingual analyzer, hunspell for same index and  
> > then  
> > index data by configuring lang detect plugin for specific fields. So here  
> > whether data will be indexed and analyzed as per the language analyzer  
> > mentioned? And also will it be searchable as per multiple hunspell  
> > dictionaries and synonyms configured as well:
> > 
> > Confirm if below settings can be ued to achieve the same:
> > 
> > ```
> > analyzer :
> > synonym :
> > tokenizer : whitespace
> > filter : [synonym]
> > default_index :
> > type : custom
> > tokenizer : whitespace
> > filter : [ standard,
> > 
> > ```
> > 
> > lowercase,hunspell\_US,hunspell\_IN,hindi,english]  
> > default\_search :  
> > type : custom  
> > tokenizer : whitespace  
> > filter : [standard, lowercase,  
> > synonym,hunspell\_US,hunspell\_IN,hindi,english]  
> > filter :  
> > hindi:  
> > tokenizer: standard  
> > filter: [lowercase]  
> > english:  
> > tokenizer: standard  
> > filter: [lowercase]  
> > synonym :  
> > type : synonym  
> > ignore\_case : true  
> > expand : true  
> > synonyms\_path : synonyms.txt  
> > hunspell\_US :  
> > type : hunspell  
> > locale : en\_US  
> > dedup : false  
> > ignore\_case : true  
> > hunspell\_IN :  
> > type : hunspell  
> > locale : hi\_IN  
> > dedup : false  
> > ignore\_case : true
> > 
> > After that Say, I have configured lang detect plugin and indexed some data  
> > with different language English and hindi. So as I have configured  
> > multiple  
> > language analyzer, MultiLingual hunspell so will I be able to perform the  
> > index and search wrt different language as with different analyzer and get  
> > the data as per analyzed tokens for different languages.
> > 
> > Also whether synonym will also work with different languages?
> > 
> > ~Prashant
> > 
> > --  
> > View this message in context:  
> > [http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html](http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html)  
> > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/1411556233114-4063950.post%40n3.nabble.com](https://groups.google.com/d/msgid/elasticsearch/1411556233114-4063950.post%40n3.nabble.com)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGEtcKLiA5oR6BRSvxdGGB7qmXx5iFoY-2dw2RHg%2BrX4g%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGEtcKLiA5oR6BRSvxdGGB7qmXx5iFoY-2dw2RHg%2BrX4g%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Nitin\_Maheshwari](https://avatars.discourse-cdn.com/v4/letter/n/77aa72/32.png) [@Nitin\_Maheshwari](https://discuss.elastic.co/u/Nitin_Maheshwari)\
**Post date:** [September 25, 2014, 8:20am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/5 "2014-09-25T08:20:05Z")

</div>

You can use langdetect plugin to identify the language of the document, and  
use that document path to set \_analyzer. \_analyzer can be set dynamically  
in that way, so the languages which are detected, analyzers with those  
names should be existing in the system.

"my\_index" : {  
"\_analyzer" : {  
"path" : "lang\_detect\_field.lang"  
},  
"properties" : {  
"lang\_detect\_field" : {  
"type" : "langdetect",  
"fields" : {  
"lang\_detect" : {  
"type" : "string"  
},  
"lang" : {  
"type" : "string"  
}  
}  
}  
}

On Wednesday, 24 September 2014 16:27:23 UTC+5:30, Prashy wrote:

> Hi All,
> 
> We are having an ES cluster which is used to index large amount of data  
> and  
> that too with different languages. So as of now our current settings was  
> pointing to English analyzer, and English hunspell but how we can achieve  
> to  
> index multilingual data along with Multi lingual analyzer and hunspell  
> setup  
> for same index (as I came across like there is a plugin called  
> Elasticsearch  
> Langdetect Plugin ([GitHub - jprante/elasticsearch-langdetect: A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector](https://github.com/jprante/elasticsearch-langdetect))  
> available from ES 1.2.1).
> 
> Current analyzer setting is like:  
> index :  
> analysis :  
> analyzer :  
> synonym :  
> tokenizer : whitespace  
> filter : [synonym]  
> default\_index :  
> type : custom  
> tokenizer : whitespace  
> filter : [standard, lowercase,hunspell\_US]  
> default\_search :  
> type : custom  
> tokenizer : whitespace  
> filter : [standard, lowercase, synonym,hunspell\_US]  
> filter :  
> synonym :  
> type : synonym  
> ignore\_case : true  
> expand : true  
> synonyms\_path : synonyms.txt  
> hunspell\_US :  
> type : hunspell  
> locale : en\_US  
> dedup : false  
> ignore\_case : true
> 
> So here,
> 
> 1. Can we configure multilingual analyzer, hunspell for same index and  
> then  
> index data by configuring lang detect plugin for specific fields. So here  
> whether data will be indexed and analyzed as per the language analyzer  
> mentioned? And also will it be searchable as per multiple hunspell  
> dictionaries and synonyms configured as well:
> 
> Confirm if below settings can be ued to achieve the same:
> 
> ```
> analyzer : 
> synonym : 
> tokenizer : whitespace 
> filter : [synonym] 
> default_index : 
> type : custom 
> tokenizer : whitespace 
> filter : [ standard, 
> 
> ```
> 
> lowercase,hunspell\_US,hunspell\_IN,hindi,english]  
> default\_search :  
> type : custom  
> tokenizer : whitespace  
> filter : [standard, lowercase,  
> synonym,hunspell\_US,hunspell\_IN,hindi,english]  
> filter :  
> hindi:  
> tokenizer: standard  
> filter: [lowercase]  
> english:  
> tokenizer: standard  
> filter: [lowercase]  
> synonym :  
> type : synonym  
> ignore\_case : true  
> expand : true  
> synonyms\_path : synonyms.txt  
> hunspell\_US :  
> type : hunspell  
> locale : en\_US  
> dedup : false  
> ignore\_case : true  
> hunspell\_IN :  
> type : hunspell  
> locale : hi\_IN  
> dedup : false  
> ignore\_case : true
> 
> After that Say, I have configured lang detect plugin and indexed some data  
> with different language English and hindi. So as I have configured  
> multiple  
> language analyzer, MultiLingual hunspell so will I be able to perform the  
> index and search wrt different language as with different analyzer and get  
> the data as per analyzed tokens for different languages.
> 
> Also whether synonym will also work with different languages?
> 
> ~Prashant
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html](http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 25, 2014, 8:23am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/6 "2014-09-25T08:23:47Z")

</div>

Wow, this works? Surprise....

Jörg

On Thu, Sep 25, 2014 at 10:20 AM, Nitin Maheshwari [ask4nitin@gmail.com](mailto:ask4nitin@gmail.com)  
wrote:

> You can use langdetect plugin to identify the language of the document,  
> and use that document path to set \_analyzer. \_analyzer can be set  
> dynamically in that way, so the languages which are detected, analyzers  
> with those names should be existing in the system.
> 
> "my\_index" : {  
> "\_analyzer" : {  
> "path" : "lang\_detect\_field.lang"  
> },  
> "properties" : {  
> "lang\_detect\_field" : {  
> "type" : "langdetect",  
> "fields" : {  
> "lang\_detect" : {  
> "type" : "string"  
> },  
> "lang" : {  
> "type" : "string"  
> }  
> }  
> }  
> }
> 
> On Wednesday, 24 September 2014 16:27:23 UTC+5:30, Prashy wrote:
> 
> > Hi All,
> > 
> > We are having an ES cluster which is used to index large amount of data  
> > and  
> > that too with different languages. So as of now our current settings was  
> > pointing to English analyzer, and English hunspell but how we can achieve  
> > to  
> > index multilingual data along with Multi lingual analyzer and hunspell  
> > setup  
> > for same index (as I came across like there is a plugin called  
> > Elasticsearch  
> > Langdetect Plugin ([GitHub - jprante/elasticsearch-langdetect: A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector](https://github.com/jprante/elasticsearch-langdetect))  
> > available from ES 1.2.1).
> > 
> > Current analyzer setting is like:  
> > index :  
> > analysis :  
> > analyzer :  
> > synonym :  
> > tokenizer : whitespace  
> > filter : [synonym]  
> > default\_index :  
> > type : custom  
> > tokenizer : whitespace  
> > filter : [standard, lowercase,hunspell\_US]  
> > default\_search :  
> > type : custom  
> > tokenizer : whitespace  
> > filter : [standard, lowercase, synonym,hunspell\_US]  
> > filter :  
> > synonym :  
> > type : synonym  
> > ignore\_case : true  
> > expand : true  
> > synonyms\_path : synonyms.txt  
> > hunspell\_US :  
> > type : hunspell  
> > locale : en\_US  
> > dedup : false  
> > ignore\_case : true
> > 
> > So here,
> > 
> > 1. Can we configure multilingual analyzer, hunspell for same index and  
> > then  
> > index data by configuring lang detect plugin for specific fields. So here  
> > whether data will be indexed and analyzed as per the language analyzer  
> > mentioned? And also will it be searchable as per multiple hunspell  
> > dictionaries and synonyms configured as well:
> > 
> > Confirm if below settings can be ued to achieve the same:
> > 
> > ```
> > analyzer :
> > synonym :
> > tokenizer : whitespace
> > filter : [synonym]
> > default_index :
> > type : custom
> > tokenizer : whitespace
> > filter : [ standard,
> > 
> > ```
> > 
> > lowercase,hunspell\_US,hunspell\_IN,hindi,english]  
> > default\_search :  
> > type : custom  
> > tokenizer : whitespace  
> > filter : [standard, lowercase,  
> > synonym,hunspell\_US,hunspell\_IN,hindi,english]  
> > filter :  
> > hindi:  
> > tokenizer: standard  
> > filter: [lowercase]  
> > english:  
> > tokenizer: standard  
> > filter: [lowercase]  
> > synonym :  
> > type : synonym  
> > ignore\_case : true  
> > expand : true  
> > synonyms\_path : synonyms.txt  
> > hunspell\_US :  
> > type : hunspell  
> > locale : en\_US  
> > dedup : false  
> > ignore\_case : true  
> > hunspell\_IN :  
> > type : hunspell  
> > locale : hi\_IN  
> > dedup : false  
> > ignore\_case : true
> > 
> > After that Say, I have configured lang detect plugin and indexed some  
> > data  
> > with different language English and hindi. So as I have configured  
> > multiple  
> > language analyzer, MultiLingual hunspell so will I be able to perform the  
> > index and search wrt different language as with different analyzer and  
> > get  
> > the data as per analyzed tokens for different languages.
> > 
> > Also whether synonym will also work with different languages?
> > 
> > ~Prashant
> > 
> > --  
> > View this message in context: [http://elasticsearch-users](http://elasticsearch-users).  
> > [115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-](http://115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-)  
> > Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html  
> > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Nitin\_Maheshwari](https://avatars.discourse-cdn.com/v4/letter/n/77aa72/32.png) [@Nitin\_Maheshwari](https://discuss.elastic.co/u/Nitin_Maheshwari)\
**Post date:** [September 25, 2014, 8:36am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/7 "2014-09-25T08:36:03Z")

</div>

Yes, it works... 🙂

But I am dealing with a different problem... I have a data set of around  
200,000 records, It detected wrong language for 8000 records, may be the  
text size was small. For most wrong case of lang detection it returned af.

Indexing that document is failing if the detected language based analyzer  
is not defined in the system. I am trying to find how can i set a default  
analyzer is the analyzer discovered does not exist.

On Thu, Sep 25, 2014 at 1:53 PM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<  
[joergprante@gmail.com](mailto:joergprante@gmail.com)\> wrote:

> Wow, this works? Surprise....
> 
> Jörg
> 
> On Thu, Sep 25, 2014 at 10:20 AM, Nitin Maheshwari [ask4nitin@gmail.com](mailto:ask4nitin@gmail.com)  
> wrote:
> 
> > You can use langdetect plugin to identify the language of the document,  
> > and use that document path to set \_analyzer. \_analyzer can be set  
> > dynamically in that way, so the languages which are detected, analyzers  
> > with those names should be existing in the system.
> > 
> > "my\_index" : {  
> > "\_analyzer" : {  
> > "path" : "lang\_detect\_field.lang"  
> > },  
> > "properties" : {  
> > "lang\_detect\_field" : {  
> > "type" : "langdetect",  
> > "fields" : {  
> > "lang\_detect" : {  
> > "type" : "string"  
> > },  
> > "lang" : {  
> > "type" : "string"  
> > }  
> > }  
> > }  
> > }
> > 
> > On Wednesday, 24 September 2014 16:27:23 UTC+5:30, Prashy wrote:
> > 
> > > Hi All,
> > > 
> > > We are having an ES cluster which is used to index large amount of data  
> > > and  
> > > that too with different languages. So as of now our current settings was  
> > > pointing to English analyzer, and English hunspell but how we can  
> > > achieve to  
> > > index multilingual data along with Multi lingual analyzer and hunspell  
> > > setup  
> > > for same index (as I came across like there is a plugin called  
> > > Elasticsearch  
> > > Langdetect Plugin ([GitHub - jprante/elasticsearch-langdetect: A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector](https://github.com/jprante/elasticsearch-langdetect))  
> > > available from ES 1.2.1).
> > > 
> > > Current analyzer setting is like:  
> > > index :  
> > > analysis :  
> > > analyzer :  
> > > synonym :  
> > > tokenizer : whitespace  
> > > filter : [synonym]  
> > > default\_index :  
> > > type : custom  
> > > tokenizer : whitespace  
> > > filter : [standard, lowercase,hunspell\_US]  
> > > default\_search :  
> > > type : custom  
> > > tokenizer : whitespace  
> > > filter : [standard, lowercase, synonym,hunspell\_US]  
> > > filter :  
> > > synonym :  
> > > type : synonym  
> > > ignore\_case : true  
> > > expand : true  
> > > synonyms\_path : synonyms.txt  
> > > hunspell\_US :  
> > > type : hunspell  
> > > locale : en\_US  
> > > dedup : false  
> > > ignore\_case : true
> > > 
> > > So here,
> > > 
> > > 1. Can we configure multilingual analyzer, hunspell for same index and  
> > > then  
> > > index data by configuring lang detect plugin for specific fields. So  
> > > here  
> > > whether data will be indexed and analyzed as per the language analyzer  
> > > mentioned? And also will it be searchable as per multiple hunspell  
> > > dictionaries and synonyms configured as well:
> > > 
> > > Confirm if below settings can be ued to achieve the same:
> > > 
> > > ```
> > > analyzer :
> > > synonym :
> > > tokenizer : whitespace
> > > filter : [synonym]
> > > default_index :
> > > type : custom
> > > tokenizer : whitespace
> > > filter : [ standard,
> > > 
> > > ```
> > > 
> > > lowercase,hunspell\_US,hunspell\_IN,hindi,english]  
> > > default\_search :  
> > > type : custom  
> > > tokenizer : whitespace  
> > > filter : [standard, lowercase,  
> > > synonym,hunspell\_US,hunspell\_IN,hindi,english]  
> > > filter :  
> > > hindi:  
> > > tokenizer: standard  
> > > filter: [lowercase]  
> > > english:  
> > > tokenizer: standard  
> > > filter: [lowercase]  
> > > synonym :  
> > > type : synonym  
> > > ignore\_case : true  
> > > expand : true  
> > > synonyms\_path : synonyms.txt  
> > > hunspell\_US :  
> > > type : hunspell  
> > > locale : en\_US  
> > > dedup : false  
> > > ignore\_case : true  
> > > hunspell\_IN :  
> > > type : hunspell  
> > > locale : hi\_IN  
> > > dedup : false  
> > > ignore\_case : true
> > > 
> > > After that Say, I have configured lang detect plugin and indexed some  
> > > data  
> > > with different language English and hindi. So as I have configured  
> > > multiple  
> > > language analyzer, MultiLingual hunspell so will I be able to perform  
> > > the  
> > > index and search wrt different language as with different analyzer and  
> > > get  
> > > the data as per analyzed tokens for different languages.
> > > 
> > > Also whether synonym will also work with different languages?
> > > 
> > > ~Prashant
> > > 
> > > --  
> > > View this message in context: [http://elasticsearch-users](http://elasticsearch-users).  
> > > [115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-](http://115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-)  
> > > Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html  
> > > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/BYE\_y-2ni9I/unsubscribe](https://groups.google.com/d/topic/elasticsearch/BYE_y-2ni9I/unsubscribe).  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com)  
> [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com?utm_medium=email&utm_source=footer)  
> .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Nitin (Nits)  
[http://nitinmaheshwari.in](http://nitinmaheshwari.in)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAHDjFTE1XS4DAkLvN0KegCifzR92JgL\_h2VZMHYt27FVQmJBnw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAHDjFTE1XS4DAkLvN0KegCifzR92JgL_h2VZMHYt27FVQmJBnw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 25, 2014, 9:02am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/8 "2014-09-25T09:02:19Z")

</div>

Sure. In next release, I will add a new parameters, so there will be better  
control:

- the languages that can be detected (default: all)

- if there is more than one language detected, how many language codes are  
indexed

- threshold levels for successful detection

- and, how many words must be in a field before detection is executed (or  
field length in characters). The number of words should be at least 3,  
otherwise detection is close to random.

Jörg

On Thu, Sep 25, 2014 at 10:36 AM, Nitin Maheshwari [ask4nitin@gmail.com](mailto:ask4nitin@gmail.com)  
wrote:

> Yes, it works... 🙂
> 
> But I am dealing with a different problem... I have a data set of around  
> 200,000 records, It detected wrong language for 8000 records, may be the  
> text size was small. For most wrong case of lang detection it returned af.
> 
> Indexing that document is failing if the detected language based analyzer  
> is not defined in the system. I am trying to find how can i set a default  
> analyzer is the analyzer discovered does not exist.
> 
> On Thu, Sep 25, 2014 at 1:53 PM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<  
> [joergprante@gmail.com](mailto:joergprante@gmail.com)\> wrote:
> 
> > Wow, this works? Surprise....
> > 
> > Jörg
> > 
> > On Thu, Sep 25, 2014 at 10:20 AM, Nitin Maheshwari [ask4nitin@gmail.com](mailto:ask4nitin@gmail.com)  
> > wrote:
> > 
> > > You can use langdetect plugin to identify the language of the document,  
> > > and use that document path to set \_analyzer. \_analyzer can be set  
> > > dynamically in that way, so the languages which are detected, analyzers  
> > > with those names should be existing in the system.
> > > 
> > > "my\_index" : {  
> > > "\_analyzer" : {  
> > > "path" : "lang\_detect\_field.lang"  
> > > },  
> > > "properties" : {  
> > > "lang\_detect\_field" : {  
> > > "type" : "langdetect",  
> > > "fields" : {  
> > > "lang\_detect" : {  
> > > "type" : "string"  
> > > },  
> > > "lang" : {  
> > > "type" : "string"  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > On Wednesday, 24 September 2014 16:27:23 UTC+5:30, Prashy wrote:
> > > 
> > > > Hi All,
> > > > 
> > > > We are having an ES cluster which is used to index large amount of data  
> > > > and  
> > > > that too with different languages. So as of now our current settings  
> > > > was  
> > > > pointing to English analyzer, and English hunspell but how we can  
> > > > achieve to  
> > > > index multilingual data along with Multi lingual analyzer and hunspell  
> > > > setup  
> > > > for same index (as I came across like there is a plugin called  
> > > > Elasticsearch  
> > > > Langdetect Plugin ([GitHub - jprante/elasticsearch-langdetect: A plugin for language detection in Elasticsearch using Nakatani Shuyo's language detector](https://github.com/jprante/elasticsearch-langdetect))
> > > > 
> > > > available from ES 1.2.1).
> > > > 
> > > > Current analyzer setting is like:  
> > > > index :  
> > > > analysis :  
> > > > analyzer :  
> > > > synonym :  
> > > > tokenizer : whitespace  
> > > > filter : [synonym]  
> > > > default\_index :  
> > > > type : custom  
> > > > tokenizer : whitespace  
> > > > filter : [standard, lowercase,hunspell\_US]  
> > > > default\_search :  
> > > > type : custom  
> > > > tokenizer : whitespace  
> > > > filter : [standard, lowercase, synonym,hunspell\_US]  
> > > > filter :  
> > > > synonym :  
> > > > type : synonym  
> > > > ignore\_case : true  
> > > > expand : true  
> > > > synonyms\_path : synonyms.txt  
> > > > hunspell\_US :  
> > > > type : hunspell  
> > > > locale : en\_US  
> > > > dedup : false  
> > > > ignore\_case : true
> > > > 
> > > > So here,
> > > > 
> > > > 1. Can we configure multilingual analyzer, hunspell for same index and  
> > > > then  
> > > > index data by configuring lang detect plugin for specific fields. So  
> > > > here  
> > > > whether data will be indexed and analyzed as per the language analyzer  
> > > > mentioned? And also will it be searchable as per multiple hunspell  
> > > > dictionaries and synonyms configured as well:
> > > > 
> > > > Confirm if below settings can be ued to achieve the same:
> > > > 
> > > > ```
> > > > analyzer :
> > > > synonym :
> > > > tokenizer : whitespace
> > > > filter : [synonym]
> > > > default_index :
> > > > type : custom
> > > > tokenizer : whitespace
> > > > filter : [ standard,
> > > > 
> > > > ```
> > > > 
> > > > lowercase,hunspell\_US,hunspell\_IN,hindi,english]  
> > > > default\_search :  
> > > > type : custom  
> > > > tokenizer : whitespace  
> > > > filter : [standard, lowercase,  
> > > > synonym,hunspell\_US,hunspell\_IN,hindi,english]  
> > > > filter :  
> > > > hindi:  
> > > > tokenizer: standard  
> > > > filter: [lowercase]  
> > > > english:  
> > > > tokenizer: standard  
> > > > filter: [lowercase]  
> > > > synonym :  
> > > > type : synonym  
> > > > ignore\_case : true  
> > > > expand : true  
> > > > synonyms\_path : synonyms.txt  
> > > > hunspell\_US :  
> > > > type : hunspell  
> > > > locale : en\_US  
> > > > dedup : false  
> > > > ignore\_case : true  
> > > > hunspell\_IN :  
> > > > type : hunspell  
> > > > locale : hi\_IN  
> > > > dedup : false  
> > > > ignore\_case : true
> > > > 
> > > > After that Say, I have configured lang detect plugin and indexed some  
> > > > data  
> > > > with different language English and hindi. So as I have configured  
> > > > multiple  
> > > > language analyzer, MultiLingual hunspell so will I be able to perform  
> > > > the  
> > > > index and search wrt different language as with different analyzer and  
> > > > get  
> > > > the data as per analyzed tokens for different languages.
> > > > 
> > > > Also whether synonym will also work with different languages?
> > > > 
> > > > ~Prashant
> > > > 
> > > > --  
> > > > View this message in context: [http://elasticsearch-users](http://elasticsearch-users).  
> > > > [115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-](http://115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-)  
> > > > Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950.html  
> > > > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > To view this discussion on the web visit  
> > > [https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ba265cf2-aaad-48b3-b454-930f9ac8a2cf%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .
> > > 
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > 
> > --  
> > You received this message because you are subscribed to a topic in the  
> > Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit  
> > [https://groups.google.com/d/topic/elasticsearch/BYE\_y-2ni9I/unsubscribe](https://groups.google.com/d/topic/elasticsearch/BYE_y-2ni9I/unsubscribe).  
> > To unsubscribe from this group and all its topics, send an email to  
> > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFCYHDPF%2BKMa67HUJ5DzWh%3D%2BpUcVAgQat%2B2AEWPb0QECw%40mail.gmail.com?utm_medium=email&utm_source=footer)  
> > .
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> Nitin (Nits)  
> [http://nitinmaheshwari.in](http://nitinmaheshwari.in)
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CAHDjFTE1XS4DAkLvN0KegCifzR92JgL\_h2VZMHYt27FVQmJBnw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAHDjFTE1XS4DAkLvN0KegCifzR92JgL_h2VZMHYt27FVQmJBnw%40mail.gmail.com)  
> [https://groups.google.com/d/msgid/elasticsearch/CAHDjFTE1XS4DAkLvN0KegCifzR92JgL\_h2VZMHYt27FVQmJBnw%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CAHDjFTE1XS4DAkLvN0KegCifzR92JgL_h2VZMHYt27FVQmJBnw%40mail.gmail.com?utm_medium=email&utm_source=footer)  
> .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGza763fsvTbmXgY71DyzQEdJPO4nAyHqnuL-QEKXPosw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGza763fsvTbmXgY71DyzQEdJPO4nAyHqnuL-QEKXPosw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Prashant\_Pal](https://avatars.discourse-cdn.com/v4/letter/p/b77776/32.png) [@Prashant\_Pal](https://discuss.elastic.co/u/Prashant_Pal)\
**Post date:** [September 25, 2014, 9:11am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/9 "2014-09-25T09:11:57Z")

</div>

Hi Jörg/Nitin,

If I am not wrong then you are setting me to set the analyzer dynamically which I have setup Manually like below:  
analyzer :  
synonym :  
tokenizer : whitespace  
filter : [synonym]  
default\_index :  
\_analyzer :  
path : lang\_detect\_field.lang  
tokenizer : whitespace  
filter : [standard, lowercase,hunspell\_US,hunspell\_IN]  
default\_search :  
\_analyzer :  
path : lang\_detect\_field.lang  
tokenizer : whitespace  
filter : [standard, lowercase, synonym,hunspell\_US,hunspell\_IN]  
filter :  
synonym :  
type : synonym  
ignore\_case : true  
expand : true  
synonyms\_path : synonyms.txt  
hunspell\_US :  
type : hunspell  
locale : en\_US  
dedup : false  
ignore\_case : true  
hunspell\_IN :  
type : hunspell  
locale : hi\_IN  
dedup : false  
ignore\_case : true .

Considering I have lang\_detect\_field in my mapping.

Correct me if I am wrong.

1. Also what if my content is an attachment type, can I still use the same. As the content for indexing will be send as base64 encoded format? If yes how we can configure that as type will be an attachment for the same.

2. If we can not achieve it using langdetect plugin then also can we have multiple language analyzer in our config as mentioned in first post, and will ES be able to recognise the same and perform indexing?

3. @Jorg, as you are owner of hunspell plugin as well, So can you let me know if I can use multiple language hunspell configured for my index setting?

~Prashant

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 25, 2014, 9:29am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/10 "2014-09-25T09:29:45Z")

</div>

You can not use "\_analyzer", which is a root mapping property, in an  
analyzer definition.

My Hunspell plugin is kind of stalled, since there is hunspell support in  
the core code:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

To be honest, it was quite an age ago when I was busy with hunspell, so I  
can not answer your question instantly. I remember the results of hunspell  
dictionaries used for stems were not satisfying. This was also due to a  
poor hunspell dictionary reader. I hope the ES core code works better.

Because hunspell stemming is a token filter, you'd have to create a bunch  
of custom analyzers with a hunspell token filter per language, and address  
them via the "\_analyzer" path method as shown above.

Jörg

On Thu, Sep 25, 2014 at 11:11 AM, Prashant Agrawal \<  
[prashant.agrawal@paladion.net](mailto:prashant.agrawal@paladion.net)\> wrote:

> Hi Jörg/Nitin,
> 
> If I am not wrong then you are setting me to set the analyzer dynamically  
> which I have setup Manually like below:  
> analyzer :  
> synonym :  
> tokenizer : whitespace  
> filter : [synonym]  
> default\_index :  
> \_analyzer :  
> path : lang\_detect\_field.lang  
> tokenizer : whitespace  
> filter : [standard, lowercase,hunspell\_US,hunspell\_IN]  
> default\_search :  
> \_analyzer :  
> path : lang\_detect\_field.lang  
> tokenizer : whitespace  
> filter : [standard, lowercase, synonym,hunspell\_US,hunspell\_IN]  
> filter :  
> synonym :  
> type : synonym  
> ignore\_case : true  
> expand : true  
> synonyms\_path : synonyms.txt  
> hunspell\_US :  
> type : hunspell  
> locale : en\_US  
> dedup : false  
> ignore\_case : true  
> hunspell\_IN :  
> type : hunspell  
> locale : hi\_IN  
> dedup : false  
> ignore\_case : true .
> 
> Considering I have lang\_detect\_field in my mapping.
> 
> Correct me if I am wrong.
> 
> 1. Also what if my content is an attachment type, can I still use the same.  
> As the content for indexing will be send as base64 encoded format? If yes  
> how we can configure that as type will be an attachment for the same.
> 
> 2. If we can not achieve it using langdetect plugin then also can we have  
> multiple language analyzer in our config as mentioned in first post, and  
> will ES be able to recognise the same and perform indexing?
> 
> 3. @Jorg, as you are owner of hunspell plugin as well, So can you let me  
> know if I can use multiple language hunspell configured for my index  
> setting?
> 
> ~Prashant
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950p4064005.html](http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950p4064005.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/1411636317742-4064005.post%40n3.nabble.com](https://groups.google.com/d/msgid/elasticsearch/1411636317742-4064005.post%40n3.nabble.com)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoE2wne8VQx0oLJOojqdcxZVmk3\_axe5BHcDuPe2%2Bj6NrA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoE2wne8VQx0oLJOojqdcxZVmk3_axe5BHcDuPe2%2Bj6NrA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Prashant\_Pal](https://avatars.discourse-cdn.com/v4/letter/p/b77776/32.png) [@Prashant\_Pal](https://discuss.elastic.co/u/Prashant_Pal)\
**Post date:** [September 25, 2014, 9:41am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/11 "2014-09-25T09:41:29Z")

</div>

Hi Jorg,

What about this

1. Also what if my content is an attachment type, can I still use the same. As the content for indexing will be send as base64 encoded format? If yes how we can configure that as type will be an attachment for the same.

2. If we can not achieve it using langdetect plugin then also can we have multiple language analyzer in our config as mentioned in first post, and will ES be able to recognise the same and perform indexing?

3. I hope the ES core code works better.  
Here what do you mean by ES core code, is there any specific settings which can be used for grammar based search ?

Also as I am not getting the clear picture (it got mixed somewhere with \_analyzer) for analyzer stuff like how we can create multiple (different language) analyzer for same index. It would be great if you can give me a little demonstration which can be overwritten with analyzer setting in my first post (where I have used default\_analyzer and default\_search) to replace with two analyzer dynamically.

~Prashant

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [September 25, 2014, 10:01am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/12 "2014-09-25T10:01:18Z")

</div>

Langdetect works with binary content, e.g. from attachment mapper.

The other questions can not be answered quick. If I find time, I can post  
something at [gist.github.com](http://gist.github.com)

Jörg

On Thu, Sep 25, 2014 at 11:41 AM, Prashant Agrawal \<  
[prashant.agrawal@paladion.net](mailto:prashant.agrawal@paladion.net)\> wrote:

> Hi Jorg,
> 
> What about this
> 
> 1. Also what if my content is an attachment type, can I still use the same.  
> As the content for indexing will be send as base64 encoded format? If yes  
> how we can configure that as type will be an attachment for the same.
> 
> 2. If we can not achieve it using langdetect plugin then also can we have  
> multiple language analyzer in our config as mentioned in first post, and  
> will ES be able to recognise the same and perform indexing?
> 
> 3. I hope the ES core code works better.  
> Here what do you mean by ES core code, is there any specific settings which  
> can be used for grammar based search ?
> 
> Also as I am not getting the clear picture (it got mixed somewhere with  
> \_analyzer) for analyzer stuff like how we can create multiple (different  
> language) analyzer for same index. It would be great if you can give me a  
> little demonstration which can be overwritten with analyzer setting in my  
> first post (where I have used default\_analyzer and default\_search) to  
> replace with two analyzer dynamically.
> 
> ~Prashant
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950p4064009.html](http://elasticsearch-users.115913.n3.nabble.com/Use-case-of-multiple-Language-Analyzer-Hunspell-along-with-Elasticsearch-Langdetect-Plugin-tp4063950p4064009.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/1411638089417-4064009.post%40n3.nabble.com](https://groups.google.com/d/msgid/elasticsearch/1411638089417-4064009.post%40n3.nabble.com)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFLgwg2AwQiuiXGD27L3DddhnzGcn5-1Z5mkXJp9zE\_6A%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoFLgwg2AwQiuiXGD27L3DddhnzGcn5-1Z5mkXJp9zE_6A%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Prashant\_Pal](https://avatars.discourse-cdn.com/v4/letter/p/b77776/32.png) [@Prashant\_Pal](https://discuss.elastic.co/u/Prashant_Pal)\
**Post date:** [September 25, 2014, 10:18am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/13 "2014-09-25T10:18:40Z")

</div>

Ok no problem.

so can you let me know if you post to any of query in [gist.github.com](http://gist.github.com).

Also I hope by using langdetect plugin we can analyze the different language content as well (using \_analyzer) so a feature request for "a dynamic "copy\_to" to duplicate the field after detection to a field with language-specific analyzer like a synonym analyzer." is not required now right ?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:00am UTC](https://discuss.elastic.co/t/use-case-of-multiple-language-analyzer-hunspell-along-with-elasticsearch-langdetect-plugin/19938/14 "2017-07-06T01:00:03Z")

</div>


