# Confused with the Smart Chinese Analysis plugin

**URL:** <https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782>\
**Category:** Elasticsearch\
**Created:** [April 8, 2016, 8:52am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782 "2016-04-08T08:52:47Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [April 8, 2016, 8:52am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/1 "2016-04-08T08:52:47Z")

</div>

I set smartcn as default analyzer in config file

```auto
index.analysis.analyzer.default.type : "smartcn"

```

And I create a type `item` which had a string field `name`, like this:

```json
"name": "开关/插座 -代金券-西门子"

```

You can see it as `A/B - C-D` if you are confused with Chinese.  
I tried the analyze api

```json
GET /newmall/_analyze
{
  "text":"开关/插座 -代金券-西门子"
}

```

The text successfully cut into “A”, "B", "C", "D". All the stopwords and punctuation were deprecated.  
When I searched “B", "C", "D" separately, it all returned the right doc.  
However, when I searched “A”, which is "开关", it hit zero.

```json
GET /newmall/item/_search
{
  "query": {
    "match": {
      "name": "开关"
    }
  }
}

```

And the funny thing is that when I tried with the term “A/”, which is "开关/". It returned the right doc.

```json
GET /newmall/item/_search
{
  "query": {
    "match": {
      "name": "开关/"
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [April 12, 2016, 8:14am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/2 "2016-04-12T08:14:50Z")

</div>

Hi,

I have to admit that I have not the slightest clue about Chinese so bear with me. 🙂

Based on what I see, the analysis process does not seem to emit the token 开关 correctly. Could you please post the response of:

```auto
GET /newmall/_analyze
{
  "text":"开关/插座 -代金券-西门子"
}

```

Daniel

---

<div class="post-metadata">

**Author:** ![forloop](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/forloop/32/9021_2.png) [@forloop](https://discuss.elastic.co/u/forloop)\
**Post date:** [April 12, 2016, 8:56am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/3 "2016-04-12T08:56:11Z")

</div>

I've just tried this on Elasticsearch 2.3.0 and get the same issue. Here are the steps to reproduce

1.install smart chinese analysis plugin

```auto
sudo bin/plugin install analysis-smartcn

```

2.create an index

```auto
curl -XPUT "http://localhost:9200/newmall" -d'
{
  "mappings": {
    "mall":{
      "properties": {
        "text": {
          "type": "string",
          "analyzer": "smartcn"
        }
      }
    }
  }
}'

```

3.Test the analysis of "开关/插座 -代金券-西门子"

```auto
curl -XGET "http://localhost:9200/newmall/_analyze?analyzer=smartcn" -d'
{
  "text":"开关/插座 -代金券-西门子"
}'

```

yields the correct tokens

```auto
{
  "tokens": [
    {
      "token": "开关",
      "start_offset": 0,
      "end_offset": 2,
      "type": "word",
      "position": 0
    },
    {
      "token": "插座",
      "start_offset": 3,
      "end_offset": 5,
      "type": "word",
      "position": 2
    },
    {
      "token": "代金",
      "start_offset": 7,
      "end_offset": 9,
      "type": "word",
      "position": 4
    },
    {
      "token": "券",
      "start_offset": 9,
      "end_offset": 10,
      "type": "word",
      "position": 5
    },
    {
      "token": "西门子",
      "start_offset": 11,
      "end_offset": 14,
      "type": "word",
      "position": 7
    }
  ]
}

```

4.Index a document with text we just analyzed

```auto
curl -XPOST "http://localhost:9200/newmall/mall/1" -d'
{
  "text": "开关/插座 -代金券-西门子"
}'

```

5.Perform `match` query on "开关"

```auto
curl -XGET "http://localhost:9200/newmall/mall/_search?explain" -d'
{
  "query": {
    "match": {
      "text": "开关"
    }
  }
}'

```

yields **no** results

6.But performing `match` query on "开关/"

```auto
curl -XGET "http://localhost:9200/newmall/mall/_search?explain" -d'
{
  "query": {
    "match": {
      "text": "开关/"
    }
  }
}'

```

yields results (with explanation)

```auto
{
  "took": 2,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "failed": 0
  },
  "hits": {
    "total": 1,
    "max_score": 0.13424811,
    "hits": [
      {
        "_shard": 3,
        "_node": "F4_VMt-7Qpi11ACD3y4QYw",
        "_index": "newmall",
        "_type": "mall",
        "_id": "1",
        "_score": 0.13424811,
        "_source": {
          "text": "开关/插座 -代金券-西门子"
        },
        "_explanation": {
          "value": 0.13424811,
          "description": "sum of:",
          "details": [
            {
              "value": 0.13424811,
              "description": "weight(text:开关 in 0) [PerFieldSimilarity], result of:",
              "details": [
                {
                  "value": 0.13424811,
                  "description": "fieldWeight in 0, product of:",
                  "details": [
                    {
                      "value": 1,
                      "description": "tf(freq=1.0), with freq of:",
                      "details": [
                        {
                          "value": 1,
                          "description": "termFreq=1.0",
                          "details": []
                        }
                      ]
                    },
                    {
                      "value": 0.30685282,
                      "description": "idf(docFreq=1, maxDocs=1)",
                      "details": []
                    },
                    {
                      "value": 0.4375,
                      "description": "fieldNorm(doc=0)",
                      "details": []
                    }
                  ]
                }
              ]
            },
            {
              "value": 0,
              "description": "match on required clause, product of:",
              "details": [
                {
                  "value": 0,
                  "description": "# clause",
                  "details": []
                },
                {
                  "value": 3.2588913,
                  "description": "_type:mall, product of:",
                  "details": [
                    {
                      "value": 1,
                      "description": "boost",
                      "details": []
                    },
                    {
                      "value": 3.2588913,
                      "description": "queryNorm",
                      "details": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      }
    ]
  }
}

```

---

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [April 12, 2016, 9:09am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/4 "2016-04-12T09:09:34Z")

</div>

@forloop had reproduced the issue below

---

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [April 12, 2016, 9:14am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/5 "2016-04-12T09:14:33Z")

</div>

That's exactly the issue what I had met.  
Thank you for you reproducing. My poor English may not explain the issue clearly.

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [April 12, 2016, 10:57am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/6 "2016-04-12T10:57:50Z")

</div>

I tried two more things:

```auto
GET /newmall/_analyze?analyzer=smartcn
{
    "text": "开关"
}

```

This returns:

```auto
{
   "tokens": [
      {
         "token": "开",
         "start_offset": 0,
         "end_offset": 1,
         "type": "word",
         "position": 0
      },
      {
         "token": "关",
         "start_offset": 1,
         "end_offset": 2,
         "type": "word",
         "position": 1
      }
   ]
}

```

Whereas:

```auto
GET /newmall/_analyze?analyzer=smartcn
{
    "text": "开关/"
}

```

returns

```auto
{
   "tokens": [
      {
         "token": "开关",
         "start_offset": 0,
         "end_offset": 2,
         "type": "word",
         "position": 0
      }
   ]
}

```

So it appears to me that the problem is in the search phase. For a Western language I'd say this is weird but I don't know whether it makes sense to tokenize "开关" as two separate tokens or they always belong together.

As an alternative you could try if you get better results with the [icu-analysis plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/analysis-icu.html).

Daniel

---

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [April 12, 2016, 12:26pm UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/7 "2016-04-12T12:26:27Z")

</div>

Hi, Daniel  
开 means switch on, while 关 means switch off. Both are verbs.  
If combined together, 开关 means noun switch.

So it's the problem of the words segmentation algorithm?  
When indexing, the smartcn analyzer cut the text into ["开关", "插座", "代金券", "西门子"].  
When searching, the smartcn analyzer cut the query "开关" into ["开", "关"].

But why cannot the term (or called token?) "开" and “关” match "开关"?  
As contrast, could "turn" or "over" match "turnover"?

---

<div class="post-metadata">

**Author:** ![thn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thn/32/8061_2.png) [@thn](https://discuss.elastic.co/u/thn)\
**Post date:** [April 12, 2016, 12:33pm UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/8 "2016-04-12T12:33:47Z")

</div>

> [@Morriaty](#):
>
> "开

Can you try "开\*"?

> [@Morriaty](#):
>
> "开关"

Can you try "开关\*"?

What you are seeing is the analyzer that analyzes the search query string, it splits "开关" into two parts, not the analyzer that was used or specified for the "text" field. I think you can look up for a way to specify the analyzer that you prefer to use with your query string, otherwise it will use the default. Please check.

---

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [April 13, 2016, 1:19am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/9 "2016-04-13T01:19:56Z")

</div>

> [@Morriaty](#):
>
> I set smartcn as default analyzer in config file
> 
> index.analysis.analyzer.default.type : "smartcn"

Please pay attention to the first line of this topic. I have set smartcn as the default analyzer.

---

<div class="post-metadata">

**Author:** ![medcl.net](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/medcl.net/32/4414_2.png) [@medcl.net](https://discuss.elastic.co/u/medcl.net)\
**Post date:** [April 13, 2016, 2:32am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/10 "2016-04-13T02:32:03Z")

</div>

> [@forloop](#):
>
> 开关/插座 -代金券-西门子

Hey @Morriaty

There issue is with the tokenizer, you know the SmartCN will do tokenization against the sentence, and SmartCN is using HMM algorithm to compute how to segment the text, and context matters, so long text and short text may have different tokenization result, maybe you can try this analyzer：

> **[GitHub - medcl/elasticsearch-analysis-ik: The IK Analysis plugin integrates...](https://github.com/medcl/elasticsearch-analysis-ik)**
>
> The IK Analysis plugin integrates Lucene IK analyzer into elasticsearch, support customized dictionary. - GitHub - medcl/elasticsearch-analysis-ik: The IK Analysis plugin integrates Lucene IK analy...

and try to use `ik_max_word ` .

---

<div class="post-metadata">

**Author:** ![Morriaty](https://avatars.discourse-cdn.com/v4/letter/m/8e8cbc/32.png) [@Morriaty](https://discuss.elastic.co/u/Morriaty)\
**Post date:** [April 13, 2016, 2:52am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/11 "2016-04-13T02:52:22Z")

</div>

Hi @medcl1 medcl, I suppose we could talk in Chinese.😂  
事实上，我已经用过ik了，结果跟上面的一样。这是我现在用的template。还是说我应该在template中只写`ik_max_word`？

```auto
{
  "template": "newmall",
  "settings": {
    "number_of_shards": 3,
    "number_of_replicas": 1,
    "analysis": {
      "analyzer": {
        "ik": {
          "type": "ik"
        },
        "ik_max_word": {
          "type": "ik",
          "use_smart": false
        },
        "ik_smart": {
          "type": "ik",
          "use_smart": true
        }
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![medcl.net](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/medcl.net/32/4414_2.png) [@medcl.net](https://discuss.elastic.co/u/medcl.net)\
**Post date:** [April 13, 2016, 3:00am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/12 "2016-04-13T03:00:59Z")

</div>

```auto
PUT /newmall
{
  "mappings": {
    "mall":{
      "properties": {
        "text": {
          "type": "string",
          "analyzer": "ik"
        }
      }
    }
  }
}

POST /newmall/mall/1
{
  "text": "开关/插座 -代金券-西门子"
}
GET /newmall/mall/_search?explain
{
  "query": {
    "match": {
      "text": "开关"
    }
  }
}

```

you don't need to custom the analyzer, they are ready to use directly, try the example show above.  
if you are using template, you must recreate the index.

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [April 13, 2016, 6:48am UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/13 "2016-04-13T06:48:29Z")

</div>

@medcl1: Thanks for chiming in! 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:59pm UTC](https://discuss.elastic.co/t/confused-with-the-smart-chinese-analysis-plugin/46782/14 "2017-07-05T22:59:44Z")

</div>


