# Data duplicated in Elasticsearch when added from Hive - RESOLVED

**URL:** <https://discuss.elastic.co/t/data-duplicated-in-elasticsearch-when-added-from-hive-resolved/140400>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [July 17, 2018, 9:47pm UTC](https://discuss.elastic.co/t/data-duplicated-in-elasticsearch-when-added-from-hive-resolved/140400 "2018-07-17T21:47:57Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![hsk04](https://avatars.discourse-cdn.com/v4/letter/h/c57346/32.png) [@hsk04](https://discuss.elastic.co/u/hsk04)\
**Post date:** [July 17, 2018, 9:47pm UTC](https://discuss.elastic.co/t/data-duplicated-in-elasticsearch-when-added-from-hive-resolved/140400/1 "2018-07-17T21:47:58Z")

</div>

Hi all,

I have a table in Hive and it has: 1,412,444 records  
with following fields:  
timestamp TIMESTAMP, clientid STRING, kitnumber STRING, sessionid STRING, sessionduration INT, countofsession INT, dayssincelastsession INT, totalengegedtime INT, pagevisible INT, pagehidden INT, eventvalue INT, totalevents INT, hits INT, eventcategory STRING, eventaction STRING, eventlabel STRING, vimeouploaddate STRING, usertype STRING, page STRING, pagedepth INT, pagetitle STRING, screenname STRING, landingpage STRING, landingscreenname STRING, secondpage STRING, exitpage STRING, regionisocode STRING, city STRING, latitude DOUBLE, longitude DOUBLE, serviceprovider STRING, devicecategory STRING, javaenabled STRING, operatingsystem STRING, operatingsystemversion STRING, screencolors STRING, screenresolution STRING, browser STRING, browsersize STRING, browserversion STRING, mobiledeviceinfo STRING, mobileinputselector STRING

When I upload the same data to an external Hive table linked to Elasticsearch,  
The record count becomes: 2,290,171 and many duplicates are created.

Can anyone help me figure-out why is this happening?  
Thanks in advance,  
Kishore.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [July 25, 2018, 2:50pm UTC](https://discuss.elastic.co/t/data-duplicated-in-elasticsearch-when-added-from-hive-resolved/140400/2 "2018-07-25T14:50:30Z")

</div>

Are you specifying one of the fields to be an ID field when you write? Have you had any failed tasks during the write operations? It's possible that rescheduled tasks, or re-run jobs can add duplicate data to Elasticsearch since without an ID field, one will be generated for each record sent.

---

<div class="post-metadata">

**Author:** ![hsk04](https://avatars.discourse-cdn.com/v4/letter/h/c57346/32.png) [@hsk04](https://discuss.elastic.co/u/hsk04)\
**Post date:** [July 26, 2018, 10:44pm UTC](https://discuss.elastic.co/t/data-duplicated-in-elasticsearch-when-added-from-hive-resolved/140400/3 "2018-07-26T22:44:58Z")

</div>

Hey James,  
Thanks for the reply. Sorry to update here. The issue got resolved.  
Yes, there was no ID field.  
How I fixed it:  
\*\*\* use this command to load data into elasticsearch, 'order by 1' avoids duplicates in the elasticsearch \*\*\*

$ insert overwrite table elk select \* from act order by 1;

This forces the HIVE to use only 1 reducer whereas in previous case 4 reducers were running and data was getting duplicated.

Thanks again James,  
Kishore.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 23, 2018, 10:44pm UTC](https://discuss.elastic.co/t/data-duplicated-in-elasticsearch-when-added-from-hive-resolved/140400/4 "2018-08-23T22:44:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
