# Filebeat set add\_id: ~ does not take effect

**URL:** <https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [June 25, 2023, 3:08am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825 "2023-06-25T03:08:26Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [June 25, 2023, 3:08am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/1 "2023-06-25T03:08:26Z")

</div>

I have a filebeat 7.16.2 to extract messages in kafka 3.4. After setting `add_id: ~`, restarting filebeat will repeatedly send data to elasticsearch.

filebeat.yml:

```auto
filebeat.yml: |
  filebeat.inputs:
  - type: kafka
    hosts:
      - 10.43.182.158:9092
    topics: ["yh_standardadapter_api_pro"]
    group_id: "yh_standardadapter_api_pro_filebeat"
    tags: ["yh_standardadapter_api_pro"]
    parsers:
    - ndjson:
      keys_under_root: true
      add_error_key: true
      overwrite_keys: true
    processors:
    - add_id: ~
  processors:
  - drop_fields:
      fields: ["service_id","kafka.partition","input.type","kafka.topic","agent.ephemeral_id","agent.hostname","agent.id","agent.name","agent.type","agent.version","ecs.version","host.name","kafka.key","kafka.offset"]
  output.elasticsearch:
    hosts: 'http://10.43.100.17:9200'
    username: "elastic"
    password: "elastic"
    indices:
      - index: "yh_standardadapter_api_console_pro_%{+yyyy.MM.dd}"
        when.and:
          - equals:
              fields.LoggerType: "Console"
          - contains:
              tags: "yh_standardadapter_api_pro"
      - index: "yh_standardadapter_api_requestlog_pro_%{+yyyy.MM.dd}"
        when.and:
          - equals:
              fields.LoggerType: "RequestLog"
          - contains:
              tags: "yh_standardadapter_api_pro"
      - index: "yh_standardadapter_api_messagetypelog_pro_%{+yyyy.MM.dd}"
        when.and:
          - equals:
              fields.LoggerType: "MessageTypeLog"
          - contains:
              tags: "yh_standardadapter_api_pro"

```

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 14, 2023, 3:09am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/2 "2023-07-14T03:09:10Z")

</div>

Does anyone know why this is

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 14, 2023, 11:33pm UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/3 "2023-07-14T23:33:19Z")

</div>

> [@yt\_h](#):
>
> restarting filebeat will repeatedly send data to elasticsearch.

What do you mean by repeatedly send data? Send the same messages over and over?

Does it do it if you take that add\_id out?

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 27, 2023, 7:46am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/4 "2023-07-27T07:46:48Z")

</div>

When I produce a json message like this in kafka

```auto
{
    "@timestamp":"2023-07-27T14:43:35.778+08:00",
    "message":"[820535e7-22a1-4797-b5df-f646aab2c1b2] Ack server push request, request = NotifySubscriberRequest, requestId = 179",
    "level":"INFO"
}

```

filebeat will be correctly consumed in elasticsearch, and when I restart filebeat, it will send the message to elasticsearch again, and there will be an extra record for this, why does `add_id: ~` not take effect here

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 27, 2023, 1:48pm UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/5 "2023-07-27T13:48:34Z")

</div>

Can you show the 2 duplicate docs in elasticsearch please

This sounds more like filebeat re-reading the Kafka queue

Perhaps @leandrojmp might have so insight as I am not a Kafka expert

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [July 27, 2023, 2:18pm UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/6 "2023-07-27T14:18:28Z")

</div>

Can you share the duplicates message in Kibana as @stephenb asked?

I'm not a Kafka expert, but looking at your config it doesn't seem that filebeat would consume already consumed messages as the `group_id` doesn't change.

Could be the producer sending the same message twice to kafka?

Also, what you want to achieve with the `add_id` processor? This processor just adds a random id to the event, but if it receives the same message one or more times, the id of those messages will be different.

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 1:33am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/7 "2023-07-28T01:33:24Z")

</div>

Yes, I can show two duplicate documents in elasticsearch, filebeat double consumes messages from kafka

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 1:38am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/8 "2023-07-28T01:38:23Z")

</div>

This is in my test environment. I use the producer to send only one message to Kafka. There is always one message in Kafka. When filebeat is running, the message will be delivered to elasticsearch once. When I restart filebeat, filebeat will deliver the message once again. Elasticsearch, during the restart of filebeat, there is no new message in kafka. I want to use add\_id to generate a unique id for each message, so as to avoid repeated consumption of data to Elasticsearch when filebeat restarts

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [July 28, 2023, 1:38am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/9 "2023-07-28T01:38:58Z")

</div>

> [@yt\_h](#):
>
> Yes, I can show two duplicate documents in elasticsearch, filebeat double consumes messages from kafka

If possible please edit your filebeat and remove this field `"kafka.offset"` from the list of fields that you are dropping to check that it is indeed the same message from kafka.

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 1:39am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/10 "2023-07-28T01:39:52Z")

</div>

It is described in the document that add\_id processors is to generate a unique id for time, is my understanding wrong?  
`https://www.elastic.co/guide/en/beats/filebeat/7.16/add-id.html`

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 1:41am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/11 "2023-07-28T01:41:13Z")

</div>

> [@leandrojmp](#):
>
> If possible please edit your filebeat and remove this field `"kafka.offset"` from the list of fields that you are dropping to check that it is indeed the same message from kafka.

In my test environment, it can be confirmed that the message is from kafka

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [July 28, 2023, 1:41am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/12 "2023-07-28T01:41:59Z")

</div>

> [@yt\_h](#):
>
> In my test environment, it can be confirmed that the message is from kafka

Yes, but share the duplicated messages that you are seeing in kibana and include the kafka offset field.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [July 28, 2023, 2:03am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/13 "2023-07-28T02:03:28Z")

</div>

> [@yt\_h](#):
>
> It is described in the document that add\_id processors is to generate a unique id for time, is my understanding wrong?

If I'm not wrong it is unique in the sense that each event will have an unique id, but if filebeat process the same message again for some reason the id will not be the same.

For example, if you have a log file with the following lines:

```auto
first message
first message
second message

```

Each one of those lines are one event, and each one will have a unique id, the first two lines are the same, but the add generated by the `add_id` processor will be diferent.

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 2:07am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/14 "2023-07-28T02:07:10Z")

</div>

![image](https://us1.discourse-cdn.com/elastic/original/3X/4/b/4bc142bc93b2071bf17b725d9d1f9130aea2ecbd.png)  
Hi, this is what I just simulated. Except for the filebeat **agent.hostname** field, the other fields are exactly the same. Is it because the filebeat **agent.hostname** field is not the same? Is it not a repeated event?

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 2:11am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/15 "2023-07-28T02:11:23Z")

</div>

It means that as long as filebeat is restarted, each document will be recorded repeatedly. Does the add\_id processor have no practical effect?

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 2:21am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/16 "2023-07-28T02:21:07Z")

</div>

Oh no, it failed, I fixed agent.hostname as filebeat-0

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/e/bed99a24e8e285d0f483422fa170646704961791.png)

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [July 28, 2023, 2:26am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/17 "2023-07-28T02:26:52Z")

</div>

> [@yt\_h](#):
>
> Does the add\_id processor have no practical effect?

It just adds a unique ad, it will not help with duplicates, what you want is the [fingerprint processor](https://www.elastic.co/guide/en/beats/filebeat/current/fingerprint.html#fingerprint).

This will generate a unique id based on some field of your document, like the `message` field, so if the message is the same, the id generated would be the same.

But this is not the issue here, the issue is that your kafka input is reading the same offset twice, this should not happen if the group\_id is the same, but I'm not sure if the issue is on filebeat or on Kafka.

---

<div class="post-metadata">

**Author:** ![yt\_h](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yt_h/32/79484_2.png) [@yt\_h](https://discuss.elastic.co/u/yt_h)\
**Post date:** [July 28, 2023, 3:19am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/18 "2023-07-28T03:19:49Z")

</div>

Here is my kafka manifest file if you are interested  
[https://github.com/huangyutongs/hyt](https://github.com/huangyutongs/hyt)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 25, 2023, 5:20am UTC](https://discuss.elastic.co/t/filebeat-set-add-id-does-not-take-effect/336825/19 "2023-08-25T05:20:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
