# How to prevent duplicates in ElasticSearch 2.X

**URL:** <https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068>\
**Category:** Elasticsearch\
**Created:** [February 17, 2016, 8:17pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068 "2016-02-17T20:17:32Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![acabrol](https://avatars.discourse-cdn.com/v4/letter/a/9de0a6/32.png) [@acabrol](https://discuss.elastic.co/u/acabrol)\
**Post date:** [February 17, 2016, 8:17pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068/1 "2016-02-17T20:17:32Z")

</div>

Dear all,

I've just upgraded to ES 2.2 from 1.x and discovered that custom settings on "\_id" is now deprecated more specifically "path" field.  
(see [https://www.elastic.co/guide/en/elasticsearch/reference/1.7/mapping-id-field.html](https://www.elastic.co/guide/en/elasticsearch/reference/1.7/mapping-id-field.html)).

I've found this blog which propose workaround to prevent duplicates during insert of docs:

> **[Eliminating Duplicate Documents in Elasticsearch](https://qbox.io/blog/minimizing-document-duplication-in-elasticsearch)**
>
> Avoiding duplication in your Elasticsearch indexes is always a good thing. But you can gain other benefits by eliminating duplicates: save disk space, improve search accuracy, improve the efficiency of hardware resource management. Perhaps most...

From my point of view they are several risk with those solutions:

1. the duplicate check is not stored in index settings so in case of multiple clients which are not managed you can have duplicates
2. preprocessing increase load during insertion where path to unique field solved the duplication issue easily
3. the duplication prevention became much more complex so implementation failure risk is important

Could you help me to find an easy ES settings side way to prevent duplicate insert in Elasticsearch 2.0 or 2.2?

Regards,  
Alexandre.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [February 19, 2016, 5:29am UTC](https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068/2 "2016-02-19T05:29:43Z")

</div>

> [@acabrol](#):
>
> Could you help me to find an easy ES settings side way to prevent duplicate insert in Elasticsearch 2.0 or 2.2?

There isn't a settings side way. The only way to prevent duplication is to manage the \_id in your ingesting application. Your point number 1 is valid but once you have multiple clients you can't trust you have bigger problems then \_id management. I don't buy point number 2 or 3 because any application that can build the document in the first place has the data readily at hand to build the \_id.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [February 21, 2016, 3:37pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068/3 "2016-02-21T15:37:58Z")

</div>

I raised some additional serious issues with that article in their comments section. Unfortunately they are no longer there.

---

<div class="post-metadata">

**Author:** ![acabrol](https://avatars.discourse-cdn.com/v4/letter/a/9de0a6/32.png) [@acabrol](https://discuss.elastic.co/u/acabrol)\
**Post date:** [February 25, 2016, 7:45pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068/4 "2016-02-25T19:45:10Z")

</div>

Thank you for your answers.

I take note.

---

<div class="post-metadata">

**Author:** ![thn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thn/32/8061_2.png) [@thn](https://discuss.elastic.co/u/thn)\
**Post date:** [February 26, 2016, 12:48pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068/5 "2016-02-26T12:48:06Z")

</div>

You can use the MD5 hash (or something equivalent) of the document as \_id, this will prevent the dup. The way it works is when ES sees the same \_id in the index, I think it replaces the current document in the index with the incoming document (kind of like an update to an existing document) Solr works the same way.

Note: some may argue that MD5 hash does have a potential collision but the probability is low. If you are not comfortable with the hash, then look at your document... if there is something from the document that you can use as a unique value that can be assigned to \_id, then use that value.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:13pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicates-in-elasticsearch-2-x/42068/6 "2017-07-05T23:13:20Z")

</div>


