# Error when moving data from hive using elasticsearch-hadoop plugin to elasticsearch

**URL:** <https://discuss.elastic.co/t/error-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/131124>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 9, 2018, 8:19am UTC](https://discuss.elastic.co/t/error-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/131124 "2018-05-09T08:19:14Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Harbeer\_Kadian](https://avatars.discourse-cdn.com/v4/letter/h/45deac/32.png) [@Harbeer\_Kadian](https://discuss.elastic.co/u/Harbeer_Kadian)\
**Post date:** [May 9, 2018, 8:19am UTC](https://discuss.elastic.co/t/error-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/131124/1 "2018-05-09T08:19:14Z")

</div>

Hi all,

I raised following issue in the past.

> [@Duplicate documents get inserted when moving data from hive using elasticsearch-hadoop plugin to elasticsearch](https://discuss.elastic.co/t/duplicate-documents-get-inserted-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/109895):
>
> My need is to insert records from hive to elasticsearch which was going fine for me. From past few days we are observing that few of the records get duplicated while inserted into elasticsearch. I browsed about this problem and found out that one reason could be speculative execution present in hadoop. So I set following flags to false to disable that. Changes in mapred-site.xml file mapred.reduce.tasks.speculative.execution false mapred.map.tasks.speculative.execution false Changes in h…

Here the advised solution was to generate unique id for each record to prevent duplicates from getting inserted.  
I am creating my primary key using md5 function. I take all the columns which creates uniqueness for the record and create its md5 as primary key.

Before doing this primary key fix, record count used to be higher in elasticsearch because of duplicate insertion  
Now after the fix, record count is less in elasticsearch.  
I am not able to find the reason.  
Here is my ES Table properties.

TBLPROPERTIES('es.mapping.id' = 'id', 'es.nodes' = '%ES\_NODES%',  
'es.port' = '%ES\_PORT%', 'es.index.auto.create' = 'false', 'es.batch.size.bytes' = '1mb', 'es.batch.size.entries' = '500', 'es.batch.write.retry.count' = '100',  
'es.batch.write.retry.wait' = '60s', 'es.batch.write.refresh' = 'false','es.nodes.discovery' = 'false',  
'es.nodes.client.only' = 'false', 'es.resource' = '%ES\_RESOURCE%', 'es.query' = '?q=\*', 'es.nodes.wan.only' = 'true')

Am i missing some property?

Harbeer

---

<div class="post-metadata">

**Author:** ![Harbeer\_Kadian](https://avatars.discourse-cdn.com/v4/letter/h/45deac/32.png) [@Harbeer\_Kadian](https://discuss.elastic.co/u/Harbeer_Kadian)\
**Post date:** [May 12, 2018, 10:48am UTC](https://discuss.elastic.co/t/error-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/131124/2 "2018-05-12T10:48:24Z")

</div>

Also here is the error, which i see in logs.

[HISTORY][DAG:dag\_1525939456125\_0003\_1][Event:TASK\_ATTEMPT\_FINISHED]: vertexName=Map 1, taskAttemptId=atte mpt\_1525939456125\_0003\_1\_00\_000002\_0, creationTime=1525941957539, allocationTime=1525941958909, startTime=1525941963683, finishTime=1525942544544, timeTaken=580861, status=FAILED, taskFailureType=NO N\_FATAL, errorEnum=FRAMEWORK\_ERROR, diagnostics=Error: Error while running task ( failure ) : attempt\_1525939456125\_0003\_1\_00\_000002\_0:java.lang.RuntimeException: java.lang.RuntimeException: org.apa che.hadoop.hive.ql.metadata.HiveException: Hive Runtime Error while processing row

at org.apache.hadoop.hive.ql.exec.tez.TezProcessor.initializeAndRunProcessor(TezProcessor.java:211)  
at org.apache.hadoop.hive.ql.exec.tez.TezProcessor.run(TezProcessor.java:168)  
at org.apache.tez.runtime.LogicalIOProcessorRuntimeTask.run(LogicalIOProcessorRuntimeTask.java:370)  
at org.apache.tez.runtime.task.TaskRunner2Callable$1.run(TaskRunner2Callable.java:73)  
at org.apache.tez.runtime.task.TaskRunner2Callable$1.run(TaskRunner2Callable.java:61)  
at java.security.AccessController.doPrivileged(Native Method)  
at javax.security.auth.Subject.doAs(Subject.java:422)  
at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1698)  
at org.apache.tez.runtime.task.TaskRunner2Callable.callInternal(TaskRunner2Callable.java:61)  
at org.apache.tez.runtime.task.TaskRunner2Callable.callInternal(TaskRunner2Callable.java:37)  
at org.apache.tez.common.CallableWithNdc.call(CallableWithNdc.java:36)  
at java.util.concurrent.FutureTask.run(FutureTask.java:266)  
at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)  
at java.lang.Thread.run(Thread.java:748)

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [May 30, 2018, 7:33pm UTC](https://discuss.elastic.co/t/error-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/131124/3 "2018-05-30T19:33:30Z")

</div>

Unfortunately, that error message does not help too much. Can you check the job task logs to see if there's anything else that might highlight a problem?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 27, 2018, 7:37pm UTC](https://discuss.elastic.co/t/error-when-moving-data-from-hive-using-elasticsearch-hadoop-plugin-to-elasticsearch/131124/4 "2018-06-27T19:37:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
