# Remove duplicate tokens after edge\_ngram on array

**URL:** <https://discuss.elastic.co/t/remove-duplicate-tokens-after-edge-ngram-on-array/294945>\
**Category:** Elasticsearch\
**Created:** [January 20, 2022, 11:30am UTC](https://discuss.elastic.co/t/remove-duplicate-tokens-after-edge-ngram-on-array/294945 "2022-01-20T11:30:55Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![timon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timon/32/100619_2.png) [@timon](https://discuss.elastic.co/u/timon)\
**Post date:** [January 20, 2022, 11:30am UTC](https://discuss.elastic.co/t/remove-duplicate-tokens-after-edge-ngram-on-array/294945/1 "2022-01-20T11:30:55Z")

</div>

I have a field where multiple values are stored:

`field: ["testOneTwo", "testThreeFour"]`

I would like to analyze this field with an `edge_ngram` filter, but also remove duplicate tokens. I tried the `unique` and `remove_duplicates` filter.  
Example Settings:

```auto
{
        "settings": {
            "analysis": {
                "filter": {
                    "edgengram_filter": {
                        "type": "edgeNGram",
                        "min_gram": 1,
                        "max_gram": 24
                    }
                },
                "tokenizer": {
                    "edgengram": {
                        "type": "edge_ngram",
                        "min_gram": 1,
                        "max_gram": 24,
                        "token_chars": [
                            "letter",
                            "digit"
                        ]
                    }
                },
                "analyzer": {
                    "testunique": {
                        "type": "custom",
                         "tokenizer": "standard",
                        "filter": [
                            "edgengram_filter",
                            "unique"
                        ]
                    },
                    "testremove": {
                        "type": "custom",
                        "tokenizer": "standard",
                        "filter": [
                            "edgengram_filter",
                            "remove_duplicates"
                        ]
                    },
                    "testedgeunique": {
                        "tokenizer": "edgengram",
                        "filter": [
                            "unique"
                        ]
                    },
                    "testedgeremove": {
                        "tokenizer": "edgengram",
                        "filter": [
                            "remove_duplicates"
                        ]
                    }
                }
            }
        }
    }

```

For each analyzer the `_analyze` API shows the tokens `t`, `te`, `tes`, `test` two times.

When having a field where both values are stored in a single value like `"testOneTwo testThreeFour"` it works. But this is not a solution for me as I use `copy_to` and `edge_ngram` as a token filter.

Any way to enforce this? Thanks!

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 20, 2022, 11:56am UTC](https://discuss.elastic.co/t/remove-duplicate-tokens-after-edge-ngram-on-array/294945/2 "2022-01-20T11:56:00Z")

</div>

How about using [join processor](https://www.elastic.co/guide/en/elasticsearch/reference/current/join-processor.html) of ingest pipeline? You can concatenate the strings and it will work as `"testOneTwo testThreeFour"`.

---

<div class="post-metadata">

**Author:** ![timon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timon/32/100619_2.png) [@timon](https://discuss.elastic.co/u/timon)\
**Post date:** [January 20, 2022, 3:31pm UTC](https://discuss.elastic.co/t/remove-duplicate-tokens-after-edge-ngram-on-array/294945/3 "2022-01-20T15:31:41Z")

</div>

Sadly, this wont be working when I use `copy_to`. So I guess I have to join them before indexing...

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 17, 2022, 3:32pm UTC](https://discuss.elastic.co/t/remove-duplicate-tokens-after-edge-ngram-on-array/294945/4 "2022-02-17T15:32:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
