# Enrich Processor is slow on multi nodes

**URL:** <https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245>\
**Category:** Elasticsearch\
**Tags:** ingest-pipeline\
**Created:** [January 5, 2021, 5:03pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245 "2021-01-05T17:03:57Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mike\_Connor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike_connor/32/81788_2.png) [@Mike\_Connor](https://discuss.elastic.co/u/Mike_Connor)\
**Post date:** [January 5, 2021, 5:03pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/1 "2021-01-05T17:03:57Z")

</div>

Hello,

I am helping with a project that is using an ingest pipeline, specifically the enrich processor. The enrich processor is pulling from a source index that is fairly small (300k records, ~25MB). The documents that are being "enriched" are coming in at a rate of 2500/s. When running on a single node, it works fine. When we add another node, the throughput starts to decline until it stops working. We do not have a deep enough understanding of what is happening under the hood to troubleshoot this. Looking for any help/suggestions to get this working across multiple nodes.

Cheers,

---

<div class="post-metadata">

**Author:** ![ylasri](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ylasri/32/86120_2.png) [@ylasri](https://discuss.elastic.co/u/ylasri)\
**Post date:** [January 5, 2021, 7:19pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/2 "2021-01-05T19:19:00Z")

</div>

What's the setting of the "source index" ? number of replicas ? that may speed up when having multiple nodes

---

<div class="post-metadata">

**Author:** ![Mike\_Connor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike_connor/32/81788_2.png) [@Mike\_Connor](https://discuss.elastic.co/u/Mike_Connor)\
**Post date:** [January 5, 2021, 7:33pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/3 "2021-01-05T19:33:21Z")

</div>

Here is the settings for the source index:

```auto
{
  "settings": {
    "index": {
      "creation_date": "1608237270261",
      "number_of_shards": "1",
      "number_of_replicas": "1",
      "uuid": "4wuW43OyQLql9UvXvsXLTQ",
      "version": {
        "created": "7090399"
      },
      "provided_name": "vess-000001"
    }
  },
  "defaults": {
    "index": {
      "flush_after_merge": "512mb",
      "final_pipeline": "_none",
      "max_inner_result_window": "100",
      "unassigned": {
        "node_left": {
          "delayed_timeout": "1m"
        }
      },
      "max_terms_count": "65536",
      "lifecycle": {
        "name": "",
        "parse_origination_date": "false",
        "indexing_complete": "false",
        "rollover_alias": "",
        "origination_date": "-1"
      },
      "routing_partition_size": "1",
      "force_memory_term_dictionary": "false",
      "max_docvalue_fields_search": "100",
      "merge": {
        "scheduler": {
          "max_thread_count": "4",
          "auto_throttle": "true",
          "max_merge_count": "9"
        },
        "policy": {
          "reclaim_deletes_weight": "2.0",
          "floor_segment": "2mb",
          "max_merge_at_once_explicit": "30",
          "max_merge_at_once": "10",
          "max_merged_segment": "5gb",
          "expunge_deletes_allowed": "10.0",
          "segments_per_tier": "10.0",
          "deletes_pct_allowed": "33.0"
        }
      },
      "max_refresh_listeners": "1000",
      "max_regex_length": "1000",
      "load_fixed_bitset_filters_eagerly": "true",
      "number_of_routing_shards": "1",
      "write": {
        "wait_for_active_shards": "1"
      },
      "verified_before_close": "false",
      "mapping": {
        "coerce": "false",
        "nested_fields": {
          "limit": "50"
        },
        "depth": {
          "limit": "20"
        },
        "field_name_length": {
          "limit": "9223372036854775807"
        },
        "total_fields": {
          "limit": "1000"
        },
        "nested_objects": {
          "limit": "10000"
        },
        "ignore_malformed": "false"
      },
      "source_only": "false",
      "soft_deletes": {
        "enabled": "false",
        "retention": {
          "operations": "0"
        },
        "retention_lease": {
          "period": "12h"
        }
      },
      "max_script_fields": "32",
      "query": {
        "default_field": [
          "*"
        ],
        "parse": {
          "allow_unmapped_fields": "true"
        }
      },
      "format": "0",
      "frozen": "false",
      "sort": {
        "missing": [],
        "mode": [],
        "field": [],
        "order": []
      },
      "priority": "1",
      "codec": "default",
      "max_rescore_window": "10000",
      "max_adjacency_matrix_filters": "100",
      "analyze": {
        "max_token_count": "10000"
      },
      "gc_deletes": "60s",
      "top_metrics_max_size": "10",
      "optimize_auto_generated_id": "true",
      "max_ngram_diff": "1",
      "hidden": "false",
      "translog": {
        "generation_threshold_size": "64mb",
        "flush_threshold_size": "512mb",
        "sync_interval": "5s",
        "retention": {
          "size": "512MB",
          "age": "12h"
        },
        "durability": "REQUEST"
      },
      "auto_expand_replicas": "false",
      "mapper": {
        "dynamic": "true"
      },
      "recovery": {
        "type": ""
      },
      "requests": {
        "cache": {
          "enable": "true"
        }
      },
      "data_path": "",
      "highlight": {
        "max_analyzed_offset": "1000000"
      },
      "routing": {
        "rebalance": {
          "enable": "all"
        },
        "allocation": {
          "enable": "all",
          "total_shards_per_node": "-1"
        }
      },
      "search": {
        "slowlog": {
          "level": "TRACE",
          "threshold": {
            "fetch": {
              "warn": "-1",
              "trace": "-1",
              "debug": "-1",
              "info": "-1"
            },
            "query": {
              "warn": "-1",
              "trace": "-1",
              "debug": "-1",
              "info": "-1"
            }
          }
        },
        "idle": {
          "after": "30s"
        },
        "throttled": "false"
      },
      "fielddata": {
        "cache": "node"
      },
      "default_pipeline": "_none",
      "max_slices_per_scroll": "1024",
      "shard": {
        "check_on_startup": "false"
      },
      "xpack": {
        "watcher": {
          "template": {
            "version": ""
          }
        },
        "version": "",
        "ccr": {
          "following_index": "false"
        }
      },
      "percolator": {
        "map_unmapped_fields_as_text": "false"
      },
      "allocation": {
        "max_retries": "5",
        "existing_shards_allocator": "gateway_allocator"
      },
      "refresh_interval": "1s",
      "indexing": {
        "slowlog": {
          "reformat": "true",
          "threshold": {
            "index": {
              "warn": "-1",
              "trace": "-1",
              "debug": "-1",
              "info": "-1"
            }
          },
          "source": "1000",
          "level": "TRACE"
        }
      },
      "compound_format": "0.1",
      "blocks": {
        "metadata": "false",
        "read": "false",
        "read_only_allow_delete": "false",
        "read_only": "false",
        "write": "false"
      },
      "max_result_window": "10000",
      "store": {
        "stats_refresh_interval": "10s",
        "type": "",
        "fs": {
          "fs_lock": "native"
        },
        "preload": []
      },
      "queries": {
        "cache": {
          "enabled": "true"
        }
      },
      "warmer": {
        "enabled": "true"
      },
      "max_shingle_diff": "3",
      "query_string": {
        "lenient": "false"
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![Mike\_Connor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike_connor/32/81788_2.png) [@Mike\_Connor](https://discuss.elastic.co/u/Mike_Connor)\
**Post date:** [January 5, 2021, 11:34pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/4 "2021-01-05T23:34:40Z")

</div>

Is more replicas going to increase or decrease complexity for the enrich processor?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 5, 2021, 11:36pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/5 "2021-01-05T23:36:36Z")

</div>

Shouldn't really matter as that happens after the pipeline.

---

<div class="post-metadata">

**Author:** ![Mike\_Connor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike_connor/32/81788_2.png) [@Mike\_Connor](https://discuss.elastic.co/u/Mike_Connor)\
**Post date:** [January 5, 2021, 11:41pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/6 "2021-01-05T23:41:25Z")

</div>

Any ideas what could be slowing down the pipeline on multiple nodes? How often is the enrich index created/updated? Could the `force_merge`, when the enrich index is created, across multiple nodes be slowing the pipeline?

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [January 5, 2021, 11:56pm UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/7 "2021-01-05T23:56:48Z")

</div>

I suspect there is something unusual happened (captain obvious here)

So yes how often is the lookup index created, updated, deleted etc is it constant or is it fairly static.

Why are you force merging it it's a tiny ... How often are you doing that?

Without an the enrich processor does ingest work / scale on multiple nodes?

Are we even sure it's the enrich processor?

---

<div class="post-metadata">

**Author:** ![Mike\_Connor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike_connor/32/81788_2.png) [@Mike\_Connor](https://discuss.elastic.co/u/Mike_Connor)\
**Post date:** [January 6, 2021, 12:10am UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/8 "2021-01-06T00:10:02Z")

</div>

I was reading the docs for the enrich processor. It says the enrich index (which is managed by the system) is created using `force_merge` ([https://www.elastic.co/guide/en/elasticsearch/reference/7.9/ingest-enriching-data.html](https://www.elastic.co/guide/en/elasticsearch/reference/7.9/ingest-enriching-data.html)). The docs do not say how often the enrich index is created/updated. The creation date on the index says midnight, and it has been running for 2 weeks. So I would guess daily.

The ingest does work with the enrich processor, but only on a single node.

---

<div class="post-metadata">

**Author:** ![Alain\_St\_Pierre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alain_st_pierre/32/81083_2.png) [@Alain\_St\_Pierre](https://discuss.elastic.co/u/Alain_St_Pierre)\
**Post date:** [January 6, 2021, 12:24am UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/9 "2021-01-06T00:24:50Z")

</div>

Hey guys... I am working with Mike on this...

We have a Lambda function executing the enrich policy daily. This keeps the enrich index fresh as the source index is constantly being updated. When I remove the enrich processor, we are able to ingest successfully with multiple nodes (even across AZ's).

This issue is related to a thread I started here: [ES Cloud - Running ingest pipelines on warm nodes?](https://discuss.elastic.co/t/es-cloud-running-ingest-pipelines-on-warm-nodes/259054/4)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 6, 2021, 2:21am UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/10 "2021-01-06T02:21:57Z")

</div>

It's going to be easier if we can keep a single thread with all the info. Let's carry on here please 🙂

> [@ES Cloud - Running ingest pipelines on warm nodes?](https://discuss.elastic.co/t/es-cloud-running-ingest-pipelines-on-warm-nodes/259054/5):
>
> Hi Gents, Can we keep this in 1 thread please, thx. @Alain_St_Pierre @Mike_Connor@warkolm can you merge these 2 threads? or close the other one. @Alain_St_Pierre thanks for answering on the other thread that help me understand a few things, it sounded like someone was executing the force merge on their own etc.. I am not sure whether I can help you or not but if you are up for it I will ask some basic questions perhaps we can get to the bottom of this, but I am a volunteer here / I have a…

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 6, 2021, 2:22am UTC](https://discuss.elastic.co/t/enrich-processor-is-slow-on-multi-nodes/260245/11 "2021-01-06T02:22:00Z")

</div>


