# How to identify and remove duplicates in Elasticsearch index

**URL:** <https://discuss.elastic.co/t/how-to-identify-and-remove-duplicates-in-elasticsearch-index/307866>\
**Category:** Elasticsearch\
**Created:** [June 22, 2022, 10:40am UTC](https://discuss.elastic.co/t/how-to-identify-and-remove-duplicates-in-elasticsearch-index/307866 "2022-06-22T10:40:05Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sowmya1](https://avatars.discourse-cdn.com/v4/letter/s/58f4c7/32.png) [@Sowmya1](https://discuss.elastic.co/u/Sowmya1)\
**Post date:** [June 22, 2022, 10:40am UTC](https://discuss.elastic.co/t/how-to-identify-and-remove-duplicates-in-elasticsearch-index/307866/1 "2022-06-22T10:40:05Z")

</div>

we are using elasticsearch 7.11.1 recently we observed an issue and below are the points for it.

1. we store data for every 15 mins interval and we get time stamp from our input file (ex: 05:00, 23:15, 20:30, 11:45 )
2. recently we observed our input file at 23:15 has 1890 records, but index has 3533 records.
3. now we want to delete 1643 duplicate records from index, with out disturbing 1890 records.

We need API query for that.

for example

input file

name product sale id  
sai pen 100 1  
kumar car 30 2  
sai pen 100 1  
sai pen 100 1  
ram bike 288 3  
kumar car 30 2

After deleting duplicates my index should look like below,

name product sale id  
sai pen 100 1  
ram bike 288 3  
kumar car 30 2

I need help with

1. query to find only duplicates at 23:15
2. query to delete duplicates

Can you please share the API query for the above issue.

---

<div class="post-metadata">

**Author:** ![Sandeep\_Raju](https://avatars.discourse-cdn.com/v4/letter/s/8797f3/32.png) [@Sandeep\_Raju](https://discuss.elastic.co/u/Sandeep_Raju)\
**Post date:** [June 22, 2022, 11:37am UTC](https://discuss.elastic.co/t/how-to-identify-and-remove-duplicates-in-elasticsearch-index/307866/2 "2022-06-22T11:37:50Z")

</div>

One way we can do this is be concatenating all 4 fields into one field and then if that field count \> 1 , then use delete by query to delete that duplicate.  
If you are using some time field like `timestamp` or `updated` , I believe we can use `delete by query` for this.

---

<div class="post-metadata">

**Author:** ![Sowmya1](https://avatars.discourse-cdn.com/v4/letter/s/58f4c7/32.png) [@Sowmya1](https://discuss.elastic.co/u/Sowmya1)\
**Post date:** [June 22, 2022, 1:25pm UTC](https://discuss.elastic.co/t/how-to-identify-and-remove-duplicates-in-elasticsearch-index/307866/3 "2022-06-22T13:25:40Z")

</div>

Please help us with the query to delete the duplicate records at time stamp of 23:15

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 22, 2022, 1:35pm UTC](https://discuss.elastic.co/t/how-to-identify-and-remove-duplicates-in-elasticsearch-index/307866/4 "2022-06-22T13:35:56Z")

</div>

Maybe [this blog post](https://www.elastic.co/blog/how-to-find-and-remove-duplicate-documents-in-elasticsearch) might be useful? I am not sure there is a way to reliably create a query to use with delete by query to handle this, so the approach described in the blog post may be safer.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 20, 2022, 1:36pm UTC](https://discuss.elastic.co/t/how-to-identify-and-remove-duplicates-in-elasticsearch-index/307866/5 "2022-07-20T13:36:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
