# Detect filebeat retries to remove duplicates in the server side

**URL:** <https://discuss.elastic.co/t/detect-filebeat-retries-to-remove-duplicates-in-the-server-side/47754>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [April 19, 2016, 7:17am UTC](https://discuss.elastic.co/t/detect-filebeat-retries-to-remove-duplicates-in-the-server-side/47754 "2016-04-19T07:17:03Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![palmerabollo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/palmerabollo/32/44772_2.png) [@palmerabollo](https://discuss.elastic.co/u/palmerabollo)\
**Post date:** [April 19, 2016, 7:17am UTC](https://discuss.elastic.co/t/detect-filebeat-retries-to-remove-duplicates-in-the-server-side/47754/1 "2016-04-19T07:17:03Z")

</div>

I'm using filebeat to send logs to a remote logstash endpoint.

Is there a flag or any way to detect if a log has already been sent? That is, detect if it is a retry.

I think filebeat adds a @timestamp field (I could use that in combination with a hash(log)) but it changes in every retry, if I'm not wrong.

Thank you.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 19, 2016, 7:38am UTC](https://discuss.elastic.co/t/detect-filebeat-retries-to-remove-duplicates-in-the-server-side/47754/2 "2016-04-19T07:38:54Z")

</div>

If it retries, it is generally not clear whether the initial record reached the destination or not. Filebeat can provide additional metadata around the event, e.g. filename and offset in file, that you could use rather than the timestamp.

---

<div class="post-metadata">

**Author:** ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)\
**Post date:** [April 19, 2016, 12:35pm UTC](https://discuss.elastic.co/t/detect-filebeat-retries-to-remove-duplicates-in-the-server-side/47754/3 "2016-04-19T12:35:37Z")

</div>

filename, offset + beat (shipper name) are a good source for deduplication. One trick (when sending to logstash) is to build an `id` based on these fields and just re-index the document. I think re-indexing in elasticsearch will mark the old entry deleted and create a new one (right, takes some disk space + CPU usage, but on compaction deleted entries are finally removed from disk). It's a very simple trick to implement deduplication.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 9:53pm UTC](https://discuss.elastic.co/t/detect-filebeat-retries-to-remove-duplicates-in-the-server-side/47754/4 "2017-07-05T21:53:01Z")

</div>


