# Duplicate data on hadoop

**URL:** <https://discuss.elastic.co/t/duplicate-data-on-hadoop/20825>\
**Category:** Elasticsearch\
**Created:** [November 19, 2014, 2:57am UTC](https://discuss.elastic.co/t/duplicate-data-on-hadoop/20825 "2014-11-19T02:57:26Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jimmy\_Carter](https://avatars.discourse-cdn.com/v4/letter/j/e19b73/32.png) [@Jimmy\_Carter](https://discuss.elastic.co/u/Jimmy_Carter)\
**Post date:** [November 19, 2014, 2:57am UTC](https://discuss.elastic.co/t/duplicate-data-on-hadoop/20825/1 "2014-11-19T02:57:26Z")

</div>

Dear All,

i have problem with import data from es to hadoop ( querying using hive),  
the problem is i have 3 recod in my es, but when i select in hive table it  
return result 21, when i check each record in es duplicate 6 times,

here the query table creation in hadoop :  
CREATE EXTERNAL TABLE conversion (  
ip string,  
user\_id string,  
tracker\_id string,  
time\_log bigint,  
session\_id string,  
spot\_id string,  
traker\_type\_id bigint,

)  
STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'  
TBLPROPERTIES('es.resource' = 'event-conversion\*/conversion',  
'es.mapping.names' = 'ip:ip, user\_id:user\_id,  
tracker\_id:tracker\_id, time\_log:datetime\_log,  
session\_id:session\_id,spot\_id:spot\_id, tracker\_type\_id:tracker\_type\_id');

so is there something that wrong in my create table query?

thank you

jimmy

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b446dd4f-4619-40f7-a083-72ae0fbbf163%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b446dd4f-4619-40f7-a083-72ae0fbbf163%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Mungeol\_Heo](https://avatars.discourse-cdn.com/v4/letter/m/dec6dc/32.png) [@Mungeol\_Heo](https://discuss.elastic.co/u/Mungeol_Heo)\
**Post date:** [December 17, 2014, 1:08am UTC](https://discuss.elastic.co/t/duplicate-data-on-hadoop/20825/2 "2014-12-17T01:08:16Z")

</div>

Try to use one index instead of multiple indexes.  
For instance, change 'es.resource' = 'event-conversion\*/conversion' to  
'es.resource' = 'event-conversion-01/conversion'.  
I think es-hadoop does not support multiple indexes setting for now.

On Wed, Nov 19, 2014 at 11:57 AM, Jimmy Carter [jimsproject@gmail.com](mailto:jimsproject@gmail.com) wrote:

> Dear All,
> 
> i have problem with import data from es to hadoop ( querying using hive),  
> the problem is i have 3 recod in my es, but when i select in hive table it  
> return result 21, when i check each record in es duplicate 6 times,
> 
> here the query table creation in hadoop :  
> CREATE EXTERNAL TABLE conversion (  
> ip string,  
> user\_id string,  
> tracker\_id string,  
> time\_log bigint,  
> session\_id string,  
> spot\_id string,  
> traker\_type\_id bigint,
> 
> )  
> STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'  
> TBLPROPERTIES('es.resource' = 'event-conversion\*/conversion',  
> 'es.mapping.names' = 'ip:ip, user\_id:user\_id,  
> tracker\_id:tracker\_id, time\_log:datetime\_log,  
> session\_id:session\_id,spot\_id:spot\_id, tracker\_type\_id:tracker\_type\_id');
> 
> so is there something that wrong in my create table query?
> 
> thank you
> 
> jimmy
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/b446dd4f-4619-40f7-a083-72ae0fbbf163%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b446dd4f-4619-40f7-a083-72ae0fbbf163%40googlegroups.com).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CADQPeWz7J6-M%3DDVp8Nj0MYkikjRyUzOg6t72dU7oawqyrRn81Q%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CADQPeWz7J6-M%3DDVp8Nj0MYkikjRyUzOg6t72dU7oawqyrRn81Q%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:43am UTC](https://discuss.elastic.co/t/duplicate-data-on-hadoop/20825/3 "2017-07-06T00:43:22Z")

</div>


