# Find and delete duplicate documents

**URL:** <https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575>\
**Category:** Elasticsearch\
**Created:** [July 30, 2015, 2:14pm UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575 "2015-07-30T14:14:40Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![noamt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/noamt/32/660_2.png) [@noamt](https://discuss.elastic.co/u/noamt)\
**Post date:** [July 30, 2015, 2:14pm UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/1 "2015-07-30T14:14:40Z")

</div>

Sorry if this has already been asked; I've mostly seen questions of how to deal with duplicate documents in the result set, but not how to actually locate and remove them from the index.

We have a type within an index that contains ~7 million documents.  
Because this data was migrated from an earlier version, there's a subset of this type that is duplicated; that is, the type contains an unknown number of documents with same data and the same ID so that using the REST-API:

1. Getting the document by ID returns a single document.
2. Searching by document ID returns multiple documents.
3. Searching by term returns multiple documents.

Assuming I've got no preliminary information about the duplicate documents other than their type, is there any way I can find and delete the duplicates while keeping only one copy?

The only solution I've got that seems viable is to dump the data from the external datasource and query each ID to check for duplicates, but I was hoping there might be a more straightforward way.

Cheers,  
Noam

---

<div class="post-metadata">

**Author:** ![PatrickKik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/patrickkik/32/619_2.png) [@PatrickKik](https://discuss.elastic.co/u/PatrickKik)\
**Post date:** [August 3, 2015, 5:05am UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/2 "2015-08-03T05:05:02Z")

</div>

There is no straightforward solution.

You could scroll over all documents and query for each and every document id if there are any duplicates.

---

<div class="post-metadata">

**Author:** ![Dominik\_Stadler](https://avatars.discourse-cdn.com/v4/letter/d/8e8cbc/32.png) [@Dominik\_Stadler](https://discuss.elastic.co/u/Dominik_Stadler)\
**Post date:** [June 15, 2016, 11:49am UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/3 "2016-06-15T11:49:03Z")

</div>

See the description at [https://qbox.io/blog/minimizing-document-duplication-in-elasticsearch](https://qbox.io/blog/minimizing-document-duplication-in-elasticsearch), it suggests something like

```
curl -XGET 'http://localhost:9200/employeeid/info/_search?pretty=true' -d '{
  "size": 0,
  "aggs": {
    "duplicateCount": {
      "terms": {
      "field": "name",
        "min_doc_count": 2
      },
      "aggs": {
        "duplicateDocuments": {
          "top_hits": {}
        }
      }
    }
  }
}'

```

which should do what you are looking for.

---

<div class="post-metadata">

**Author:** ![buxticka](https://avatars.discourse-cdn.com/v4/letter/b/f6c823/32.png) [@buxticka](https://discuss.elastic.co/u/buxticka)\
**Post date:** [October 4, 2016, 10:51am UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/4 "2016-10-04T10:51:55Z")

</div>

Well done. But pls. tell me how to delete these documents. I have query that gives me all duplicity records I want to delete:  
GET /sube-2016.08.18/commercial\_act/\_search  
{  
"query": {  
"term": {  
"commercial\_act\_type\_name": "New subscription - Porting into the OSK"  
}  
},  
"size": 0,  
"aggs": {  
"duplicated": {  
"terms": {  
"field": "id",  
"min\_doc\_count": 2  
},  
"aggs": {  
"duplicateDocuments": {  
"top\_hits": {  
"sort": [  
{  
"@timestamp": {  
"order": "desc"  
}  
}  
],  
"fields": [  
"\_id",  
"@timestamp"  
],  
"size": 1  
}  
}  
}  
}  
}  
}

but I don't know how to send this query to delete by query.  
Thanx.

---

<div class="post-metadata">

**Author:** ![stefws](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stefws/32/6442_2.png) [@stefws](https://discuss.elastic.co/u/stefws)\
**Post date:** [December 13, 2016, 5:16pm UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/5 "2016-12-13T17:16:08Z")

</div>

Depending on the number of your duplicate, search duplicate \_id and their index and then loop through them and do DELETE on the doc id as it appear only to delete one of the duplicate.

---

<div class="post-metadata">

**Author:** ![buxticka](https://avatars.discourse-cdn.com/v4/letter/b/f6c823/32.png) [@buxticka](https://discuss.elastic.co/u/buxticka)\
**Post date:** [December 13, 2016, 7:54pm UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/6 "2016-12-13T19:54:58Z")

</div>

Thank You. I expected something like DELETE\_BY\_QUERY DSL example but there's no possibility to put aggregation buckets as a query result.

---

<div class="post-metadata">

**Author:** ![mulaninasrin](https://avatars.discourse-cdn.com/v4/letter/m/b19c9b/32.png) [@mulaninasrin](https://discuss.elastic.co/u/mulaninasrin)\
**Post date:** [May 26, 2017, 10:17am UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/7 "2017-05-26T10:17:33Z")

</div>

I have also the same problem, please give some solution.  
Also is there any functionality like "unique index" provided in mongoDB for maintaining unique data?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:00pm UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/8 "2017-07-05T22:00:00Z")

</div>



---

<div class="post-metadata">

**Author:** ![Alex\_Marquardt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alex_marquardt/32/42925_2.png) [@Alex\_Marquardt](https://discuss.elastic.co/u/Alex_Marquardt)\
**Post date:** [July 27, 2018, 5:52pm UTC](https://discuss.elastic.co/t/find-and-delete-duplicate-documents/26575/9 "2018-07-27T17:52:34Z")

</div>

I have written a blog post that describes how to deduplicate documents from Elasticsearch using either Logstash or with a custom Python script. This can be found at the following URL: [https://alexmarquardt.com/2018/07/23/deduplicating-documents-in-elasticsearch/](https://alexmarquardt.com/2018/07/23/deduplicating-documents-in-elasticsearch/)
