# Removing ongoing duplicates

**URL:** <https://discuss.elastic.co/t/removing-ongoing-duplicates/259157>\
**Category:** Logstash\
**Created:** [December 19, 2020, 11:37am UTC](https://discuss.elastic.co/t/removing-ongoing-duplicates/259157 "2020-12-19T11:37:49Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![cyberzlo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cyberzlo/32/65490_2.png) [@cyberzlo](https://discuss.elastic.co/u/cyberzlo)\
**Post date:** [December 19, 2020, 11:37am UTC](https://discuss.elastic.co/t/removing-ongoing-duplicates/259157/1 "2020-12-19T11:37:50Z")

</div>

> **[How to Find and Remove Duplicate Documents in Elasticsearch](https://www.elastic.co/blog/how-to-find-and-remove-duplicate-documents-in-elasticsearch)**
>
> Learn how to detect and remove duplicate documents from Elasticsearch using Logstash or a custom Python script.

This is what I need however I would like do it on input due parsing jsons. What I mean this tutorial show how do it post-factum, but I would like do it every time I recieve new data basing for example on data from last 1 minute or so. Is this possible and will not generate high load? When I recieve data to logstash they are duplicates in pairs, like system logs when same line will be 10 times but in successive so it is no needed to check like whole index, just last records, even like last 10 records etc. What I need is removing ongoing duplicates.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [December 19, 2020, 3:36pm UTC](https://discuss.elastic.co/t/removing-ongoing-duplicates/259157/2 "2020-12-19T15:36:14Z")

</div>

You could use a ruby filter. Keep the most recent messages (or a hash of them) in an array and test whether the array .include? the current message. Then .shift to remove the first entry and .push to add the current message as the last entry. The cost grows with the number of entries in the array, since Ruby will iterate over each entry in .include?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 16, 2021, 3:36pm UTC](https://discuss.elastic.co/t/removing-ongoing-duplicates/259157/3 "2021-01-16T15:36:18Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
