# How to deal with duplicate data

**URL:** <https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229>\
**Category:** Logstash\
**Created:** [February 28, 2025, 5:10pm UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229 "2025-02-28T17:10:37Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ksobon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ksobon/32/103774_2.png) [@ksobon](https://discuss.elastic.co/u/ksobon)\
**Post date:** [February 28, 2025, 5:10pm UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229/1 "2025-02-28T17:10:37Z")

</div>

I have a Filebeat pipeline that is ingesting data from an end-user machine that might be stored there for 30 days. My Logstash pipeline has the following settings:

```auto
document_id => "%{[@metadata][newId]}"
action => "create"

```

Because I wanted to make sure that the same log is never written into the database twice I set the action to create and I created my unique document\_id. That setup works fine for me but I had a situation recently where Filebeat was uninstalled and the registry folder for it was wiped clean on that machine. Now, when we got it re-installed it tries to re-ingest 30 days worth of logs and write them again.

That results in errors because the action "create" doesn't allow for overrides. I can change that to "update" but another issue that pops up is that some of the indexes that this tries to update are already in the Warm tier, and are not flagged as "write" indexes.

Any idea how I should handle this situation? Logstash is returning a lot of errors because "create" fails to deal with existing documents, and that slows it down to a grind.

I was thinking that could change the action to "update" and add the `doc_as_upsert` to `true` to make sure that I can override existing documents. Then I would need to make older indexes somehow writable again. Should I reindex that older data into the Hot tier for now, and then after logstash overrides all of the documents, I can just move it back to warm again? Does that sound like a reasonable thing to do? Any other ways to deal with this issue?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 28, 2025, 6:19pm UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229/2 "2025-02-28T18:19:28Z")

</div>

> [@ksobon](#):
>
> That results in errors because the action "create" doesn't allow for overrides.

But is this an issue? This is expected as the document already exists, so it will be reject, you can just ignore those errors.

> [@ksobon](#):
>
> I can change that to "update" but another issue that pops up is that some of the indexes that this tries to update are already in the Warm tier, and are not flagged as "write" indexes.

What the rest of your output looks like? Are you using data streams? Normal indices? Are your indices time based?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 28, 2025, 6:39pm UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229/3 "2025-02-28T18:39:26Z")

</div>

> [@ksobon](#):
>
> Logstash is returning a lot of errors because "create" fails to deal with existing documents, and that slows it down to a grind

Just saw this, the performance impact could be related to the amount of logs being written, I had a past issue related to this.

One quick solution would be to simple change the Logstash loglevel for the Elasticsearch output, to only logs on `EROR` or `FATAL` logs for example, not sure in which level the current `create` errors are logged.

Something like this:

```auto
curl -XPUT 'localhost:9600/_node/logging?pretty' -H 'Content-Type: application/json' -d'{ "logger.logstash.outputs.elasticsearch" : "FATAL"}'

```

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [March 1, 2025, 4:26am UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229/5 "2025-03-01T04:26:13Z")

</div>

> [@RainTown](#):
>
> Elastic support many integrations, vast majority I dont know in any detail, but do any of them self-generate the \_id in any default/OOTB config?

There are a lot of native Elastic Agent integrations that relies on using a custom `_id` to avoid duplicate events.

This id normally comes from a field in the original message or a fingerprint of one or more fields from the original message.

Using a custom `_id` value is how you can avoid duplicate data in Elasticsearch, it is a common approach when you want to avoid duplicated events.

---

<div class="post-metadata">

**Author:** ![strawgate](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/strawgate/32/131008_2.png) [@strawgate](https://discuss.elastic.co/u/strawgate)\
**Post date:** [March 1, 2025, 4:29am UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229/6 "2025-03-01T04:29:14Z")

</div>

If these are spread across many files you can set `ignore_older` to the number of days ago the registry was reset

If the file is older than ignore\_older, Filebeat will add the file to its registry with the offset set to the end of the file and then you can simply revert back or remove the ignore older setting.

Similarly, you could just add a processor to Filebeat or to Logstash that drops events older than the last message processed from the device prior to the registry reset.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [March 1, 2025, 7:51am UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229/7 "2025-03-01T07:51:44Z")

</div>

> [@leandrojmp](#):
>
> There are a lot of native Elastic Agent integrations that relies on using a custom `_id` to avoid duplicate events.

Obviously I stand corrected, thank you. On reflection, I have deleted the above comment as a) it was not helpful and b) it was also factually wrong. Apologies.

---

<div class="post-metadata">

**Author:** ![ksobon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ksobon/32/103774_2.png) [@ksobon](https://discuss.elastic.co/u/ksobon)\
**Post date:** [March 31, 2025, 2:24pm UTC](https://discuss.elastic.co/t/how-to-deal-with-duplicate-data/375229/8 "2025-03-31T14:24:17Z")

</div>

@leandrojmp yeah, as it turns out it's just a lot of warnings being logged into the file, and that has performance implications. I saw Logstash run out of memory and crash. When it reboots it will use up CPU. I also saw Java eat up a lot of CPU on that machine, probably due to heap memory being all used up, and it constantly performing garbage collection, etc. I tried using `failure_type_logging_whitelist` to exclude 409 from logging but that doesn't seem to work. I can change the log level to Error and that will solve the issue. I think at the moment the best way for me to go is to make sure that I keep the "registry" of Filebeat safely tucked away so that it doesn't get deleted causing this issue in the first place. Other than that I cranked up heap memory allocation to minimize out of memory errors and ease up garbage collection. That seems to be helping.
