# Error calling pipeline when setting the document ID in beats

**URL:** <https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [May 10, 2021, 9:26am UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560 "2021-05-10T09:26:25Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Francisco\_Peralta\_Gu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/francisco_peralta_gu/32/83575_2.png) [@Francisco\_Peralta\_Gu](https://discuss.elastic.co/u/Francisco_Peralta_Gu)\
**Post date:** [May 10, 2021, 9:26am UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/1 "2021-05-10T09:26:25Z")

</div>

Hi.  
I'm facing issues when I try to set the document id in filebeat to avoid duplicates:  
I have a filebeat configuration as following:

```auto
filebeat.inputs:  
    - type: log
      paths:
        - /dumps-json/wsdumps/**
      multiline.pattern: '(?s)|^}'
      multiline.negate: false 
      multiline.match: after
      processors:
      - add_id: ~
      - add_fields:
          fields:
            index: dumps
....

 output.elasticsearch:
      protocol: https
      hosts: ['${ELASTICSEARCH_HOST:elasticsearch}:${ELASTICSEARCH_PORT:9200}']
      username: ${ELASTICSEARCH_USERNAME}
      password: ${ELASTICSEARCH_PASSWORD}
      pipelines:
        - pipeline: "%{[fields.index]}"
          mappings:
            dumps: "xmltojson"

```

With the configuration above, the Elastic pipeline xmltojson is not been launched. Notice the - add\_id: ~ processor.

But if I delete that processor, the ingest pipeline is working propertly.

---

<div class="post-metadata">

**Author:** ![Marius\_Iversen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marius_iversen/32/68988_2.png) [@Marius\_Iversen](https://discuss.elastic.co/u/Marius_Iversen)\
**Post date:** [May 10, 2021, 11:52am UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/2 "2021-05-10T11:52:05Z")

</div>

The ID's field should be autogenerated when a new document is created in Elasticsearch, there shouldn't be any need to set the `add_id` processor anymore.

Do you see any duplications on ES?

---

<div class="post-metadata">

**Author:** ![Francisco\_Peralta\_Gu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/francisco_peralta_gu/32/83575_2.png) [@Francisco\_Peralta\_Gu](https://discuss.elastic.co/u/Francisco_Peralta_Gu)\
**Post date:** [May 10, 2021, 12:49pm UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/3 "2021-05-10T12:49:24Z")

</div>

Yes I can see duplicated documents as explained in this section:

> **[Deduplicate data | Filebeat Reference \[7.12\] | Elastic](https://www.elastic.co/guide/en/beats/filebeat/current/filebeat-deduplication.html)**

---

<div class="post-metadata">

**Author:** ![Marius\_Iversen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marius_iversen/32/68988_2.png) [@Marius\_Iversen](https://discuss.elastic.co/u/Marius_Iversen)\
**Post date:** [May 10, 2021, 12:54pm UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/4 "2021-05-10T12:54:03Z")

</div>

Ah okay gotcha! Which errors are you getting when you try to start it up? Or does it just not generate an ID?

You can stop the service and run it with some extra debugging logging and in the foreground if you want, with `filebeat -e -d "*"`.

The best places to look for errors is at the start when its loading the config, and when it is trying to send files to ES.

If you don't want to run it in the foreground, you can always then just grep the logs for ERR and WARN.

---

<div class="post-metadata">

**Author:** ![Francisco\_Peralta\_Gu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/francisco_peralta_gu/32/83575_2.png) [@Francisco\_Peralta\_Gu](https://discuss.elastic.co/u/Francisco_Peralta_Gu)\
**Post date:** [May 10, 2021, 12:59pm UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/5 "2021-05-10T12:59:03Z")

</div>

I'm not getting any errors. The documents are been populated into Elasticsearch but once I try to stablish the document\_id in Filebeat , the ingest pipeline "xmltojson" is no longer executed. When delete the add\_id processor, the pipeline is launched and the transformation included in it is properly applied.

---

<div class="post-metadata">

**Author:** ![Marius\_Iversen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marius_iversen/32/68988_2.png) [@Marius\_Iversen](https://discuss.elastic.co/u/Marius_Iversen)\
**Post date:** [May 10, 2021, 1:23pm UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/6 "2021-05-10T13:23:59Z")

</div>

I am a bit unsure if I understand what you mean, your pipeline is not set to "xmltojson", it is set to another type of value:

```
  pipelines:
    - pipeline: "%{[fields.index]}"
      mappings:
        dumps: "xmltojson"

```

This configuration won't execute the a pipeline named xmltojson.

We also have a `decode_xml` processor, if the purpose is to only convert XML to JSON, might that be of interest?  
Ref: [Decode XML | Filebeat Reference [7.12] | Elastic](https://www.elastic.co/guide/en/beats/filebeat/current/decode_xml.html)

---

<div class="post-metadata">

**Author:** ![Francisco\_Peralta\_Gu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/francisco_peralta_gu/32/83575_2.png) [@Francisco\_Peralta\_Gu](https://discuss.elastic.co/u/Francisco_Peralta_Gu)\
**Post date:** [May 10, 2021, 1:30pm UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/7 "2021-05-10T13:30:29Z")

</div>

Really the pipeline is not called.

I have tried that configuration and it worked depending on the fields.index value stablished by the add\_fields processor above :

```auto
 - add_fields:
          fields:
            index: dumps

```

That configuration works without the add\_id processor.

I cannot use decode\_xml due to a bug that will be resolved in 7.13 version

> [@Filebeat decode\_xml processor missing](https://discuss.elastic.co/t/filebeat-decode-xml-processor-missing/268915/6):
>
> Hi. Please we need information about if the configuration is applied right or if this is an issue of the module decode\_xml. Thank you!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 7, 2021, 3:30pm UTC](https://discuss.elastic.co/t/error-calling-pipeline-when-setting-the-document-id-in-beats/272560/8 "2021-06-07T15:30:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
