# Smartcn analyzer with Chinese punctuation

**URL:** <https://discuss.elastic.co/t/smartcn-analyzer-with-chinese-punctuation/92008>\
**Category:** Elasticsearch\
**Created:** [July 6, 2017, 3:08am UTC](https://discuss.elastic.co/t/smartcn-analyzer-with-chinese-punctuation/92008 "2017-07-06T03:08:54Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Patrick\_Lam](https://avatars.discourse-cdn.com/v4/letter/p/ed8c4c/32.png) [@Patrick\_Lam](https://discuss.elastic.co/u/Patrick_Lam)\
**Post date:** [July 6, 2017, 3:08am UTC](https://discuss.elastic.co/t/smartcn-analyzer-with-chinese-punctuation/92008/1 "2017-07-06T03:08:54Z")

</div>

Hi,

I'm wondering if anybody has any insight on how the smartcn analyzer/tokenizer handles (or doesn't handle) Chinese punctuation?

Example:

`负债表` tokenizes into one single word, but if you had a Chinese comma to the end with `负债表，`, the tokenizer now tokenizes this into two separate words: `负债` and `表`. This seems incorrect and it feels like the tokenizer should first filter out any Chinese punctuation marks?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 3, 2017, 3:09am UTC](https://discuss.elastic.co/t/smartcn-analyzer-with-chinese-punctuation/92008/2 "2017-08-03T03:09:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
