# ICU transform filters slowing down indexing: how avoid duplicate transliterations?

**URL:** <https://discuss.elastic.co/t/icu-transform-filters-slowing-down-indexing-how-avoid-duplicate-transliterations/234547>\
**Category:** Elasticsearch\
**Created:** [May 27, 2020, 1:28pm UTC](https://discuss.elastic.co/t/icu-transform-filters-slowing-down-indexing-how-avoid-duplicate-transliterations/234547 "2020-05-27T13:28:24Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pyppe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pyppe/32/3531_2.png) [@Pyppe](https://discuss.elastic.co/u/Pyppe)\
**Post date:** [May 27, 2020, 1:28pm UTC](https://discuss.elastic.co/t/icu-transform-filters-slowing-down-indexing-how-avoid-duplicate-transliterations/234547/1 "2020-05-27T13:28:24Z")

</div>

I was investigating slow Bulk API indexing as discussed in [Bulk index slowing down as index size increases](https://discuss.elastic.co/t/bulk-index-slowing-down-as-index-size-increases/234288).

It turns out the root cause seems to be slow ICU transform filters, when we start to index Chinese data.

We have mappings such as this:

```json
{
  "settings" : {
    "analysis" : {
      "analyzer" : {
        "ascii_keyword": {
          "tokenizer": "keyword",
          "char_filter": ["multi_space_char_filter", "apostrophe_remove"],
          "filter": ["lowercase", "no_accent_latin", "trim"]
        },
        "ascii_keyword_reverse": {
          "tokenizer": "keyword",
          "char_filter": ["multi_space_char_filter", "apostrophe_remove"],
          "filter": ["lowercase", "no_accent_latin", "trim", "reverse"]
        },
        "standard_ascii": {
          "tokenizer": "standard",
          "char_filter": ["multi_space_char_filter", "apostrophe_remove"],
          "filter": ["lowercase", "no_accent_latin"]
        },
        "standard_ascii_reverse": {
          "tokenizer": "standard",
          "char_filter": ["multi_space_char_filter", "apostrophe_remove"],
          "filter": ["lowercase", "no_accent_latin", "reverse"]
        }
      },
      "normalizer": {
        "lowercase_ascii": {
          "type": "custom",
          "filter": ["lowercase", "no_accent_latin"]
        }
      },
      "char_filter": {
        "multi_space_char_filter": ...,
        "apostrophe_remove": ...
      },
      "filter" : {
        "no_accent_latin" : {
          "type" : "icu_transform",
          "id" : "Any-Latin; NFD; [:Nonspacing Mark:] Remove; NFC"
        }
      }
    }
  },
  "mappings" : {
    "dynamic" : "strict",
    "properties" : {
      "searchableName" : {
        "type" : "text",
        "fields": {
          "ascii_keyword": { "type": "text", "analyzer": "ascii_keyword" },
          "ascii_keyword_reverse": { "type": "text", "analyzer": "ascii_keyword_reverse" },
          "standard_ascii": { "type": "text", "analyzer": "standard_ascii" },
          "standard_ascii_reverse": { "type": "text", "analyzer": "standard_ascii_reverse" },
          "sorted_latin": { "type": "keyword", "normalizer": "lowercase_ascii" },
          ...
        }
      }
    }
  }
}

```

Am I correct to assume that because `searchableName` fields, in the example above, use 5 different analyzers/normalized that all utilize `icu_transform`, we do the transliteration for the same text 5 different times? Can we somehow optimize this? Can filters somehow utilize intermediate results of other filters?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 24, 2020, 1:28pm UTC](https://discuss.elastic.co/t/icu-transform-filters-slowing-down-indexing-how-avoid-duplicate-transliterations/234547/2 "2020-06-24T13:28:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
