# How to use document id to avoid duplication of logs?

**URL:** <https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644>\
**Category:** Logstash\
**Created:** [May 28, 2020, 3:50am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644 "2020-05-28T03:50:00Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nani\_20](https://avatars.discourse-cdn.com/v4/letter/n/4491bb/32.png) [@Nani\_20](https://discuss.elastic.co/u/Nani_20)\
**Post date:** [May 28, 2020, 3:50am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/1 "2020-05-28T03:50:01Z")

</div>

Hi,

I'm using logstash **file input** to index various log files available in the input location. These logfiles are generated by log4j in various systems and are fetched using a cron job to the input location.

1. Is there any way I can use filebeats to fetch logs created by log4j in these systems? As I mentioned above, currently I am using a batch script cron job to fetch these from the user systems to the input location.

By default, log4j creates the backup of a log file after a size limit, so the logstash receives duplicate logs from time to time.

1. Is there any way to use an ID to avoid indexing duplicate data? Currently, I'm using **custom document\_id** with a combination of @timestamp and ID field(see below my **output filter** ). But, this seems to be overwriting the indexed data(correct me if I am wrong here). Instead, I would like to avoid indexing if it is a duplicate.

**My Output filter**

```auto
output{
   elasticsearch{
      hosts => ["http://localhost:9200"]
      index => "test"
      document_id => "%{@timestamp}_%{ID}"
   }
}

```

Any help here is appreciated. Thanks in advance

---

<div class="post-metadata">

**Author:** ![ptamba](https://avatars.discourse-cdn.com/v4/letter/p/7feea3/32.png) [@ptamba](https://discuss.elastic.co/u/ptamba)\
**Post date:** [May 28, 2020, 4:26am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/2 "2020-05-28T04:26:18Z")

</div>

> [@Nani\_20](#):
>
> Is there any way I can use filebeats to fetch logs created by log4j in these systems? As I mentioned above, currently I am using a batch script cron job to fetch these from the user systems to the input location.

filebeat has a [logstash module](https://www.elastic.co/guide/en/beats/filebeat/current/filebeat-module-logstash.html)

> [@Nani\_20](#):
>
> Is there any way to use an ID to avoid indexing duplicate data? Currently, I'm using **custom document\_id** with a combination of @timestamp and ID field(see below my **output filter** ). But, this seems to be overwriting the indexed data(correct me if I am wrong here). Instead, I would like to avoid indexing if it is a duplicate.

[doc\_as\_upsert](https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-doc_as_upsert) directives allows creation of new document if the document\_id does not exist in ES. ES will overwrite (update) document if the same document\_id exists.

---

<div class="post-metadata">

**Author:** ![Nani\_20](https://avatars.discourse-cdn.com/v4/letter/n/4491bb/32.png) [@Nani\_20](https://discuss.elastic.co/u/Nani_20)\
**Post date:** [May 28, 2020, 8:54am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/3 "2020-05-28T08:54:55Z")

</div>

@ptamba thanks for the reply.  
For the second question, I don't ES to update the file. Instead, i want logstash to ignore the duplicate. Is there any way i can do that?

---

<div class="post-metadata">

**Author:** ![ptamba](https://avatars.discourse-cdn.com/v4/letter/p/7feea3/32.png) [@ptamba](https://discuss.elastic.co/u/ptamba)\
**Post date:** [May 28, 2020, 9:29am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/4 "2020-05-28T09:29:34Z")

</div>

i haven't tried this but specifying [action create](https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-action) appears to avoid overwriting an existing document

> " \* create: indexes a document, fails if a document by that id already exists in the index. "

---

<div class="post-metadata">

**Author:** ![Nani\_20](https://avatars.discourse-cdn.com/v4/letter/n/4491bb/32.png) [@Nani\_20](https://discuss.elastic.co/u/Nani_20)\
**Post date:** [May 28, 2020, 10:28am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/5 "2020-05-28T10:28:54Z")

</div>

> action =\> "create"

Yes, the above **action** option is avoiding the duplicate to index. But, since this is in output plugin, I can see the duplicates passing through all my **filter plugin** operations which is kind of redundant.  
Is there any way to identify the duplicate before the **filter plugin** and avoid it?

---

<div class="post-metadata">

**Author:** ![ptamba](https://avatars.discourse-cdn.com/v4/letter/p/7feea3/32.png) [@ptamba](https://discuss.elastic.co/u/ptamba)\
**Post date:** [May 28, 2020, 11:48am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/6 "2020-05-28T11:48:42Z")

</div>

not that i’m aware of.

---

<div class="post-metadata">

**Author:** ![Nani\_20](https://avatars.discourse-cdn.com/v4/letter/n/4491bb/32.png) [@Nani\_20](https://discuss.elastic.co/u/Nani_20)\
**Post date:** [May 28, 2020, 6:28pm UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/7 "2020-05-28T18:28:33Z")

</div>

> [@ptamba](#):
>
> filebeat has a [logstash module](https://www.elastic.co/guide/en/beats/filebeat/current/filebeat-module-logstash.html)

@ptamba Yes, I can use filebeat agent and point the output to my Logstash instance.  
But, in my case there are 100s of users. Do I need to manually install filebeat in each system. Is there any simpler way to do this?  
Thank You

---

<div class="post-metadata">

**Author:** ![ptamba](https://avatars.discourse-cdn.com/v4/letter/p/7feea3/32.png) [@ptamba](https://discuss.elastic.co/u/ptamba)\
**Post date:** [May 29, 2020, 2:56am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/8 "2020-05-29T02:56:41Z")

</div>

if you want to use filebeat to collect logs on those systems then yes you have to install it on every system.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 26, 2020, 2:56am UTC](https://discuss.elastic.co/t/how-to-use-document-id-to-avoid-duplication-of-logs/234644/9 "2020-06-26T02:56:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
