# Duplicate Deletion in Elasticsearch 2.X

**URL:** <https://discuss.elastic.co/t/duplicate-deletion-in-elasticsearch-2-x/90934>\
**Category:** Elasticsearch\
**Created:** [June 27, 2017, 9:20am UTC](https://discuss.elastic.co/t/duplicate-deletion-in-elasticsearch-2-x/90934 "2017-06-27T09:20:05Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![saksham\_bathla](https://avatars.discourse-cdn.com/v4/letter/s/f1d935/32.png) [@saksham\_bathla](https://discuss.elastic.co/u/saksham_bathla)\
**Post date:** [June 27, 2017, 9:20am UTC](https://discuss.elastic.co/t/duplicate-deletion-in-elasticsearch-2-x/90934/1 "2017-06-27T09:20:05Z")

</div>

I have an Elasticsearch Cluster with close to a billion records (around 300 GB). The \_id field was not set at the time of ingestion. I recently discovered that there are some duplicates in my data. How can I delete those. (Their \_id is different but all the other fields are same). Combining four of the attributes gives me a unique identifier and that is what i would be using as the id in future when ingesting the data. Is there some practical way of deleting the duplicates without reindexing.

---

<div class="post-metadata">

**Author:** ![Julien](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/julien/32/19688_2.png) [@Julien](https://discuss.elastic.co/u/Julien)\
**Post date:** [June 27, 2017, 2:16pm UTC](https://discuss.elastic.co/t/duplicate-deletion-in-elasticsearch-2-x/90934/2 "2017-06-27T14:16:35Z")

</div>

There isn't any tool specifically to help with deduplication of data  
You may need to push an update to the old versions of the duplicated documents to mark them as old and do aggregations filtering those out.

To delete, you would use the delete API and probably use bulk api to save resources ([https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-delete.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-delete.html))

You can use min\_doc\_count to find duplicates which I guess you have already achieved ([https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html))... Depending on your requirement, you could check if data is different outside of the 4 fields you use to find duplicates and then decide what to do.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 25, 2017, 2:16pm UTC](https://discuss.elastic.co/t/duplicate-deletion-in-elasticsearch-2-x/90934/3 "2017-07-25T14:16:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
