# Combo analyzer - Issue with English and Japanese text being stored in same fields

**URL:** https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595
**Category:** Elasticsearch
**Created:** [February 1, 2013, 6:07pm UTC](https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595 "2013-02-01T18:07:42Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![hemant\_pahilwani](https://avatars.discourse-cdn.com/v4/letter/h/f08c70/32.png) [@hemant\_pahilwani](https://discuss.elastic.co/u/hemant_pahilwani)
#### Post date: [February 1, 2013, 6:07pm UTC](https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595/1 "2013-02-01T18:07:42Z")

</div>

I am trying to add multilingual support to elastic search and part of the  
requirement is to allow same field to store either english and japanese  
text. While research i stumbled upon combo analyzer plugin where we can  
give two analyzers on single field. Following is the configuration. I am  
using standard analyzer for english and kuromoji for japanese:

index :  
analysis :  
analyzer :  
my\_combo :  
type : combo  
sub\_analyzers : [standard, kuromoji]  
deduplication : true

Evaluating japanese text yields correct results with kuromoji: curl -XGET  
'localhost:9200/myindex/\_analyze?analyzer=kuromoji&pretty=true' -d '最近どうですか'

But when analyzing with my\_combo, it also applies standard analyzer to  
japanese text which results in tokens being created for each japanese  
character (behaviour of standard analyzer) as well as tokens created using  
kuromoji .  
curl -XGET 'localhost:9200/myindex/\_analyze?analyzer= my\_combo&pretty=true'  
-d '最近どうですか'

Is there anyway in which elastic search can detect japanese language and  
apply only kuromoji analyzer to japanese text? The other option that i was  
considering was to use multi field type and store japanese text in  
different field altogether but was wondering if there is easy way defined  
in elastic search to do handle such scenarios?

Thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)
#### Post date: [February 2, 2013, 8:18pm UTC](https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595/2 "2013-02-02T20:18:21Z")

</div>

Hey,

I'm not sure if this is sufficient for you but I would build my own  
analyzer in the settings based on the japanese tokenizer and the  
tokenfilters you need and base the tokenization on the JapaneseTokenizer.  
It should tokenize the english text only on the whitespaces and then for  
further processing I'd add lowercase, word-delimiter etc. as part of the  
filter chain and only work with one analyzer.

does this make sense?

simon

On Friday, February 1, 2013 7:07:42 PM UTC+1, hemant pahilwani wrote:

> I am trying to add multilingual support to Elasticsearch and part of the  
> requirement is to allow same field to store either english and japanese  
> text. While research i stumbled upon combo analyzer plugin where we can  
> give two analyzers on single field. Following is the configuration. I am  
> using standard analyzer for english and kuromoji for japanese:
> 
> index :  
> analysis :  
> analyzer :  
> my\_combo :  
> type : combo  
> sub\_analyzers : [standard, kuromoji]  
> deduplication : true
> 
> Evaluating japanese text yields correct results with kuromoji: curl  
> -XGET 'localhost:9200/myindex/\_analyze?analyzer=kuromoji&pretty=true' -d  
> '最近どうですか'
> 
> But when analyzing with my\_combo, it also applies standard analyzer to  
> japanese text which results in tokens being created for each japanese  
> character (behaviour of standard analyzer) as well as tokens created using  
> kuromoji .  
> curl -XGET  
> 'localhost:9200/myindex/\_analyze?analyzer= my\_combo&pretty=true' -d  
> '最近どうですか'
> 
> Is there anyway in which Elasticsearch can detect japanese language and  
> apply only kuromoji analyzer to japanese text? The other option that i was  
> considering was to use multi field type and store japanese text in  
> different field altogether but was wondering if there is easy way defined  
> in Elasticsearch to do handle such scenarios?
> 
> Thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![hemant\_pahilwani](https://avatars.discourse-cdn.com/v4/letter/h/f08c70/32.png) [@hemant\_pahilwani](https://discuss.elastic.co/u/hemant_pahilwani)
#### Post date: [February 7, 2013, 5:17am UTC](https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595/3 "2013-02-07T05:17:29Z")

</div>

Hey...i tried out the suggested approach but ran into issue with how wild  
card search works. I have been using standard analyzer and its default  
filters to do tokenization as i require regular prefix wild card search on  
english words. For example if i have "takanori" stored in index then  
searching for tak\* gives me document with takanori. Sample query below:

{query:{"query\_string" : {"query" : "tak\*","fields" : [  
"firstName"],"use\_dis\_max" : true,"analyze\_wildcard" : true}}}

Based on the suggestion above i created the custom analyzer with japanese  
tokenizer and added standard filter to it and used mapped firstName field  
to use it:

```
        my_default_analyzer : 
            type : custom
            tokenizer : kuromoji_tokenizer
            filter : [kuromoji_baseform, kuromoji_part_of_speech, 

```

kuromoji\_readingform, kuromoji\_stemmer, standard]

But running the same query having wild card search doesn't give me back any  
results. It does give me back results if i search for while string  
"takanori"

My suspicion is the type that is getting associated - "word" in case of  
my\_default\_analyzer and "" in case of standard analyzer.

$ curl -XGET  
'localhost:9200/myindex/\_analyze?analyzer=my\_default\_analyzer&pretty=true'  
-d 'Takanori'  
{  
"tokens" : [ {  
"token" : "takanori",  
"start\_offset" : 0,  
"end\_offset" : 8,  
"type" : "word",  
"position" : 1  
} ]

$ curl -XGET  
'localhost:9200/myindex/\_analyze?analyzer=standard&pretty=true' -d  
'Takanori'  
{  
"tokens" : [ {  
"token" : "takanori",  
"start\_offset" : 0,  
"end\_offset" : 8,  
"type" : "",  
"position" : 1  
} ]  
}

Any suggestion on how to make wild card search work in this case? Or may  
be i am missing something in configuration?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)
#### Post date: [February 7, 2013, 9:13am UTC](https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595/4 "2013-02-07T09:13:35Z")

</div>

Maybe I don't fully understand. But "Takanori" is written in Rōmaji. The  
Kuromoji analyzer is for Kanji.

Best regards,

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![hemant\_pahilwani](https://avatars.discourse-cdn.com/v4/letter/h/f08c70/32.png) [@hemant\_pahilwani](https://discuss.elastic.co/u/hemant_pahilwani)
#### Post date: [February 8, 2013, 5:31pm UTC](https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595/5 "2013-02-08T17:31:07Z")

</div>

Same behavior is observed if i use english word:

$curl -XGET  
'localhost:9200/myindex/\_analyze?analyzer=my\_default\_analyzer&pretty=true'  
-d 'something'  
{  
"tokens" : [ {  
"token" : "something",  
"start\_offset" : 0,  
"end\_offset" : 9,  
"type" : "word",  
"position" : 1  
} ]

}

$ curl -XGET  
'localhost:9200/myindex/\_analyze?analyzer=standard&pretty=true' -d  
'something'  
{  
"tokens" : [ {  
"token" : "something",  
"start\_offset" : 0,  
"end\_offset" : 9,  
"type" : "",  
"position" : 1  
} ]  
}

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 2:52am UTC](https://discuss.elastic.co/t/combo-analyzer-issue-with-english-and-japanese-text-being-stored-in-same-fields/10595/6 "2017-07-06T02:52:21Z")

</div>


