# Duplicated fields in analyze response object when analyzing korean with nori\_tokenizer

**URL:** <https://discuss.elastic.co/t/duplicated-fields-in-analyze-response-object-when-analyzing-korean-with-nori-tokenizer/305943>\
**Category:** Elasticsearch\
**Created:** [May 30, 2022, 12:05pm UTC](https://discuss.elastic.co/t/duplicated-fields-in-analyze-response-object-when-analyzing-korean-with-nori-tokenizer/305943 "2022-05-30T12:05:49Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![maruoovv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maruoovv/32/106201_2.png) [@maruoovv](https://discuss.elastic.co/u/maruoovv)\
**Post date:** [May 30, 2022, 12:05pm UTC](https://discuss.elastic.co/t/duplicated-fields-in-analyze-response-object-when-analyzing-korean-with-nori-tokenizer/305943/1 "2022-05-30T12:05:49Z")

</div>

Hi, i'm analyzing korean keyword with **nori tokenizer** and **synonym\_graph** , i found some field is duplicated in analyze response object.

Elasticsearch version : 7.1.1

index settings

```auto
PUT /test-index
{
  "settings" : {
  "analysis": {
          "filter": {
            "search_synonym": {
              "type": "synonym_graph",
              "synonyms": ["ㅇㅇㅇ => ㅇㅇㅇ,ㅁㅇㄹ,ㅇㄹㅇ"]
            }
          },
          "analyzer": {
            "search_synonym": {
              "filter": [
                "search_synonym"
              ],
              "type": "custom",
              "tokenizer": "nori_tokenizer"
            }
          }
        }
  }
}

```

Analyze Request

```auto
GET /test-index/_analyze
{
    "text" : "ㅇㅇㅇ",
    "analyzer" : "search_synonym",
    "explain" : true
}

```

Analyze Response

```auto
{
    "detail": {
        "custom_analyzer": true,
        "tokenizer": {
            "name": "nori_tokenizer",
            "tokens": [
                {
                    "token": "ㅇㅇㅇ",
                    "start_offset": 0,
                    "end_offset": 3,
                    "type": "word",
                    "position": 0,
                    "bytes": "[e3 85 87 e3 85 87 e3 85 87]",
                    "leftPOS": "UNKNOWN(Unknown)",
                    "morphemes": null,
                    "posType": "MORPHEME",
                    "positionLength": 1,
                    "reading": null,
                    "rightPOS": "UNKNOWN(Unknown)",
                    "termFrequency": 1
                }
            ]
        },
        "tokenfilters": [
            {
                "name": "search_synonym",
                "tokens": [
                    {
                        "token": "ㅇㅇㅇ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 0,
                        "positionLength": 2,
                        "bytes": "[e3 85 87 e3 85 87 e3 85 87]",
                        "leftPOS": null,
                        "morphemes": null,
                        "posType": null,
                        "positionLength": 2,
                        "reading": null,
                        "rightPOS": null,
                        "termFrequency": 1
                    },
                    {
                        "token": "ㅁ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 0,
                        "bytes": "[e3 85 81]",
                        "leftPOS": null,
                        "morphemes": null,
                        "posType": null,
                        "positionLength": 1,
                        "reading": null,
                        "rightPOS": null,
                        "termFrequency": 1
                    },
                    {
                        "token": "ㅇㄹㅇ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 0,
                        "positionLength": 2,
                        "bytes": "[e3 85 87 e3 84 b9 e3 85 87]",
                        "leftPOS": null,
                        "morphemes": null,
                        "posType": null,
                        "positionLength": 2,
                        "reading": null,
                        "rightPOS": null,
                        "termFrequency": 1
                    },
                    {
                        "token": "ㅇㄹ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 1,
                        "bytes": "[e3 85 87 e3 84 b9]",
                        "leftPOS": null,
                        "morphemes": null,
                        "posType": null,
                        "positionLength": 1,
                        "reading": null,
                        "rightPOS": null,
                        "termFrequency": 1
                    }
                ]
            }
        ]
    }
}

```

detail.tokenFilters.tokens[0] has duplicated field : positionLength.  
But if i added attributes other field in analyze request, that field removed.

Analyze Request

```auto
GET /test-index/_analyze
{
    "text" : "ㅇㅇㅇ",
    "analyzer" : "search_synonym",
    "explain" : true,
    "attributes" : ["rightPOS"] // adds attributes other field
}

```

Response

```auto
{
    "detail": {
        "custom_analyzer": true,
        "tokenizer": {
            "name": "nori_tokenizer",
            "tokens": [
                {
                    "token": "ㅇㅇㅇ",
                    "start_offset": 0,
                    "end_offset": 3,
                    "type": "word",
                    "position": 0,
                    "rightPOS": "UNKNOWN(Unknown)"
                }
            ]
        },
        "tokenfilters": [
            {
                "name": "search_synonym",
                "tokens": [
                    {
                        "token": "ㅇㅇㅇ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 0,
                        "positionLength": 2,
                        "rightPOS": null
                    },
                    {
                        "token": "ㅁ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 0,
                        "rightPOS": null
                    },
                    {
                        "token": "ㅇㄹㅇ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 0,
                        "positionLength": 2,
                        "rightPOS": null
                    },
                    {
                        "token": "ㅇㄹ",
                        "start_offset": 0,
                        "end_offset": 3,
                        "type": "SYNONYM",
                        "position": 1,
                        "rightPOS": null
                    }
                ]
            }
        ]
    }
}

```

Although the JSON doesn't saying "Not allow duplicated field", But recommended thing.  
I'm using Elasticsearch-rest-high-level-client, and an error occurs when request analyze in this situation.

com.fasterxml.jackson.core.JsonParseException: Duplicate field 'positionLength'

Is this normal behavior? or nori\_tokenizer's bug?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 27, 2022, 12:06pm UTC](https://discuss.elastic.co/t/duplicated-fields-in-analyze-response-object-when-analyzing-korean-with-nori-tokenizer/305943/2 "2022-06-27T12:06:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
