# Duplicate documents get inserted when moving data from hive using elasticsearch-hadoop plugin to elasticsearch

**URL:** <https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [December 1, 2017, 8:48am UTC](https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895 "2017-12-01T08:48:31Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Harbeer\_Kadian](https://avatars.discourse-cdn.com/v4/letter/h/45deac/32.png) [@Harbeer\_Kadian](https://discuss.elastic.co/u/Harbeer_Kadian)\
**Post date:** [December 1, 2017, 8:48am UTC](https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895/1 "2017-12-01T08:48:31Z")

</div>

My need is to insert records from hive to elasticsearch which was going fine for me. From past few days we are observing that few of the records get duplicated while inserted into elasticsearch.

I browsed about this problem and found out that one reason could be speculative execution present in hadoop. So I set following flags to false to disable that.

Changes in mapred-site.xml file  
mapred.reduce.tasks.speculative.execution false  
mapred.map.tasks.speculative.execution false

Changes in hive-site.xml  
hive.mapred.reduce.tasks.speculative.execution false  
But even after doing this, i am still getting document duplicacy issue. Also my elasticsearch insertion query is very straight.

insert into select \* from   
As per my understanding only mapper will be involved in this query. In short ES has more document count than record count in hive table.

I am using AWS EMR cluster for hive and AWS ElasticSearch Service for elasticsearch.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [December 13, 2017, 7:35pm UTC](https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895/2 "2017-12-13T19:35:54Z")

</div>

Do you see any failures on the executors for writing to Elasticsearch? In the event of an executor failing, all data from that task is retried, which can lead to duplicates. If the data you are ingesting is sensitive to duplication, you could specify an ID for each record to ensure that it does not duplicate data if tasks must be retried.

---

<div class="post-metadata">

**Author:** ![Harbeer\_Kadian](https://avatars.discourse-cdn.com/v4/letter/h/45deac/32.png) [@Harbeer\_Kadian](https://discuss.elastic.co/u/Harbeer_Kadian)\
**Post date:** [December 14, 2017, 8:39am UTC](https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895/3 "2017-12-14T08:39:04Z")

</div>

Actually I am using AWS Elasticsearch, It throws gateway timeout error (504). If the operation takes more than 60 seconds to complete. So I am feeling some times this bulk insert takes more than 60 seconds, and in that case not all records are inserted. Since the elasticsearch-hadoop plugin always retries all the records again, it creates duplicates.  
This issue is very particular to AWS ElasticSearch as you can not increase the gateway timeout from 60 seconds.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [December 14, 2017, 7:37pm UTC](https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895/4 "2017-12-14T19:37:03Z")

</div>

Yeah, in this situation your best bet is to ensure a unique ID is accompanied to each document you are indexing to ensure that duplicate writes are collapsed on retry.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 11, 2018, 7:37pm UTC](https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895/5 "2018-01-11T19:37:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
