# Cjk and thai analyzer customization

**URL:** <https://discuss.elastic.co/t/cjk-and-thai-analyzer-customization/10566>\
**Category:** Elasticsearch\
**Created:** [January 31, 2013, 8:01am UTC](https://discuss.elastic.co/t/cjk-and-thai-analyzer-customization/10566 "2013-01-31T08:01:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![lukas\_stanek](https://avatars.discourse-cdn.com/v4/letter/l/cc9497/32.png) [@lukas\_stanek](https://discuss.elastic.co/u/lukas_stanek)\
**Post date:** [January 31, 2013, 8:01am UTC](https://discuss.elastic.co/t/cjk-and-thai-analyzer-customization/10566/1 "2013-01-31T08:01:51Z")

</div>

Hello,

we use elasticsearch 0.20 to index short texts in many languages. We have  
configured custom analyzer - whitespace tokenizer and pattern filter in  
index settings for most languages.  
But there is a problem with Chinese, Japanese and Thai, cjk and thai  
analyzer in ES is not suitable for our needs - they contain standard  
tokenizer, which removes symbols and punctuation marks, we want to replace  
standard tokenizer with whitespace tokenizer.  
Please, can you give me an advice?  
How can cjk and thai analyzer be customized in ES?  
Is it possible to configure custom analyzer built from CJKBigramFilter or  
ThaiWordFilter in index settings, or do we have to prepare a plugin or are  
there other possibilities?

Thanks you.

Lukas

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 31, 2013, 3:22pm UTC](https://discuss.elastic.co/t/cjk-and-thai-analyzer-customization/10566/2 "2013-01-31T15:22:30Z")

</div>

If you look at the source of the CJKAnalyzer, you would notice that it  
basically is a CJKTokenizer followed by a StopFilter. The heart of the  
analyzer is the CJKTokenizer, not the standard tokenizer, so it simply  
cannot be replaced. You can modify the source and build your own plugin. I  
am assuming that most language analyzers are the same.

--  
Ivan

On Thu, Jan 31, 2013 at 12:01 AM, [lukas.stanek@memsource.com](mailto:lukas.stanek@memsource.com) wrote:

> Hello,
> 
> we use elasticsearch 0.20 to index short texts in many languages. We have  
> configured custom analyzer - whitespace tokenizer and pattern filter in  
> index settings for most languages.  
> But there is a problem with Chinese, Japanese and Thai, cjk and thai  
> analyzer in ES is not suitable for our needs - they contain standard  
> tokenizer, which removes symbols and punctuation marks, we want to replace  
> standard tokenizer with whitespace tokenizer.  
> Please, can you give me an advice?  
> How can cjk and thai analyzer be customized in ES?  
> Is it possible to configure custom analyzer built from CJKBigramFilter or  
> ThaiWordFilter in index settings, or do we have to prepare a plugin or are  
> there other possibilities?
> 
> Thanks you.
> 
> Lukas
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![lukas\_stanek](https://avatars.discourse-cdn.com/v4/letter/l/cc9497/32.png) [@lukas\_stanek](https://discuss.elastic.co/u/lukas_stanek)\
**Post date:** [January 31, 2013, 3:50pm UTC](https://discuss.elastic.co/t/cjk-and-thai-analyzer-customization/10566/3 "2013-01-31T15:50:19Z")

</div>

Hello Ivan,

thank you for reply.  
From Lucene version 3.6 CJKAnalyzer is composed of StandardTokenizer,  
CJKWidthFilter, LowerCaseFilter, CJKBigramFilter and StopFilter  
I would like to replace StandardTokenizer with WhitespaceTokenizer and  
remove StopFilter.  
In index settings like:  
{  
"analysis":{  
"analyzer":{  
"cjk\_cust":{  
"filter":[  
"cjk\_width", "lowercase", "cjk\_bigram"  
],  
"type":"custom",  
"tokenizer": "whitespace"  
}  
}  
}  
}

Can it be achieved in a simpler way than developing a new plugin with  
custom analyzer for cjk and thai?

Lukas

On Thursday, January 31, 2013 4:22:30 PM UTC+1, Ivan Brusic wrote:

> If you look at the source of the CJKAnalyzer, you would notice that it  
> basically is a CJKTokenizer followed by a StopFilter. The heart of the  
> analyzer is the CJKTokenizer, not the standard tokenizer, so it simply  
> cannot be replaced. You can modify the source and build your own plugin. I  
> am assuming that most language analyzers are the same.
> 
> --  
> Ivan
> 
> On Thu, Jan 31, 2013 at 12:01 AM, \<[lukas....@memsource.com](mailto:lukas....@memsource.com) \<javascript:\>\>wrote:
> 
> > Hello,
> > 
> > we use elasticsearch 0.20 to index short texts in many languages. We have  
> > configured custom analyzer - whitespace tokenizer and pattern filter in  
> > index settings for most languages.  
> > But there is a problem with Chinese, Japanese and Thai, cjk and thai  
> > analyzer in ES is not suitable for our needs - they contain standard  
> > tokenizer, which removes symbols and punctuation marks, we want to replace  
> > standard tokenizer with whitespace tokenizer.  
> > Please, can you give me an advice?  
> > How can cjk and thai analyzer be customized in ES?  
> > Is it possible to configure custom analyzer built from CJKBigramFilter or  
> > ThaiWordFilter in index settings, or do we have to prepare a plugin or are  
> > there other possibilities?
> > 
> > Thanks you.
> > 
> > Lukas
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [January 31, 2013, 4:17pm UTC](https://discuss.elastic.co/t/cjk-and-thai-analyzer-customization/10566/4 "2013-01-31T16:17:25Z")

</div>

Hi Lukas,

Sorry about the confusion. The CJKAnalyzer did in fact change with the 3.6  
release. I still have 3.5 in my classpath. Interesting, the old analyzer is  
now deprecated.

Your solution looks correct. I would swap the positioning of the lowercase  
and cjk\_width filters to be consistent with the original analyzer. You  
might want to look at the pattern tokenizer if the whitespace tokenizer is  
too lenient with word boundaries.

Cheers,

Ivan

On Thu, Jan 31, 2013 at 7:50 AM, [lukas.stanek@memsource.com](mailto:lukas.stanek@memsource.com) wrote:

> Hello Ivan,
> 
> thank you for reply.  
> From Lucene version 3.6 CJKAnalyzer is composed of StandardTokenizer,  
> CJKWidthFilter, LowerCaseFilter, CJKBigramFilter and StopFilter  
> I would like to replace StandardTokenizer with WhitespaceTokenizer and  
> remove StopFilter.  
> In index settings like:  
> {  
> "analysis":{  
> "analyzer":{  
> "cjk\_cust":{  
> "filter":[  
> "cjk\_width", "lowercase", "cjk\_bigram"  
> ],  
> "type":"custom",  
> "tokenizer": "whitespace"  
> }  
> }  
> }  
> }
> 
> Can it be achieved in a simpler way than developing a new plugin with  
> custom analyzer for cjk and thai?
> 
> Lukas
> 
> On Thursday, January 31, 2013 4:22:30 PM UTC+1, Ivan Brusic wrote:
> 
> > If you look at the source of the CJKAnalyzer, you would notice that it  
> > basically is a CJKTokenizer followed by a StopFilter. The heart of the  
> > analyzer is the CJKTokenizer, not the standard tokenizer, so it simply  
> > cannot be replaced. You can modify the source and build your own plugin. I  
> > am assuming that most language analyzers are the same.
> > 
> > --  
> > Ivan
> > 
> > On Thu, Jan 31, 2013 at 12:01 AM, [lukas....@memsource.com](mailto:lukas....@memsource.com) wrote:
> > 
> > > Hello,
> > > 
> > > we use elasticsearch 0.20 to index short texts in many languages. We  
> > > have configured custom analyzer - whitespace tokenizer and pattern filter  
> > > in index settings for most languages.  
> > > But there is a problem with Chinese, Japanese and Thai, cjk and thai  
> > > analyzer in ES is not suitable for our needs - they contain standard  
> > > tokenizer, which removes symbols and punctuation marks, we want to replace  
> > > standard tokenizer with whitespace tokenizer.  
> > > Please, can you give me an advice?  
> > > How can cjk and thai analyzer be customized in ES?  
> > > Is it possible to configure custom analyzer built from CJKBigramFilter  
> > > or ThaiWordFilter in index settings, or do we have to prepare a plugin or  
> > > are there other possibilities?
> > > 
> > > Thanks you.
> > > 
> > > Lukas
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > 
> > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > .
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:53am UTC](https://discuss.elastic.co/t/cjk-and-thai-analyzer-customization/10566/5 "2017-07-06T02:53:45Z")

</div>


