# Keeping duplicates

**URL:** <https://discuss.elastic.co/t/keeping-duplicates/196337>\
**Category:** Logstash\
**Created:** [August 22, 2019, 1:52pm UTC](https://discuss.elastic.co/t/keeping-duplicates/196337 "2019-08-22T13:52:54Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![edster](https://avatars.discourse-cdn.com/v4/letter/e/da6949/32.png) [@edster](https://discuss.elastic.co/u/edster)\
**Post date:** [August 22, 2019, 1:52pm UTC](https://discuss.elastic.co/t/keeping-duplicates/196337/1 "2019-08-22T13:52:55Z")

</div>

Hi,

Is there anyway to keep duplicate rows as they are and maybe include a number which differentiates them. The duplicate records aren't necessarily duplicates so much as they are multiple instances of the same item. They are all needed. This is for CSV files.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 22, 2019, 6:11pm UTC](https://discuss.elastic.co/t/keeping-duplicates/196337/2 "2019-08-22T18:11:27Z")

</div>

logstash will retain duplicate rows. elasticsearch will too unless you force them to have the same document id.

---

<div class="post-metadata">

**Author:** ![edster](https://avatars.discourse-cdn.com/v4/letter/e/da6949/32.png) [@edster](https://discuss.elastic.co/u/edster)\
**Post date:** [August 22, 2019, 7:04pm UTC](https://discuss.elastic.co/t/keeping-duplicates/196337/3 "2019-08-22T19:04:31Z")

</div>

Well the thing is, if no document\_id is set then it's generated by elastic, however, the next time those logstash config files are run it will just duplicate the data by adding another new id to each row. However, the duplicates already exists and are valid because they are just multiple instances of the same object. There is no field/column that differentiates it though. So would there be a way of including I guess a field to the id that would separate all of the multiple instances and then change those if next time there are more or less of those objects.

---

<div class="post-metadata">

**Author:** ![edster](https://avatars.discourse-cdn.com/v4/letter/e/da6949/32.png) [@edster](https://discuss.elastic.co/u/edster)\
**Post date:** [August 22, 2019, 8:08pm UTC](https://discuss.elastic.co/t/keeping-duplicates/196337/4 "2019-08-22T20:08:12Z")

</div>

Or what if i could add a field to take into account the number of occurrences for those duplicates. Then when the logstash file runs again would it just be possible to overwrite that number should the number of duplicates increase or lower?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 22, 2019, 8:35pm UTC](https://discuss.elastic.co/t/keeping-duplicates/196337/5 "2019-08-22T20:35:13Z")

</div>

I do not see a solution in that case.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 19, 2019, 8:35pm UTC](https://discuss.elastic.co/t/keeping-duplicates/196337/6 "2019-09-19T20:35:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
