# Elasticsearch Hadoop

**URL:** <https://discuss.elastic.co/t/elasticsearch-hadoop/15160>\
**Category:** Elasticsearch\
**Created:** [January 9, 2014, 7:49am UTC](https://discuss.elastic.co/t/elasticsearch-hadoop/15160 "2014-01-09T07:49:41Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Badal\_Mohapatra](https://avatars.discourse-cdn.com/v4/letter/b/f0a364/32.png) [@Badal\_Mohapatra](https://discuss.elastic.co/u/Badal_Mohapatra)\
**Post date:** [January 9, 2014, 7:49am UTC](https://discuss.elastic.co/t/elasticsearch-hadoop/15160/1 "2014-01-09T07:49:41Z")

</div>

Hi,

To index Hadoop data into elasticsearch as I understand,  
We create an external table with essstorage handler and then copy the data  
from another internal hive table doesn't it duplicate the data in HDFS?  
Is there any way to use the hive internal tables directly to index instead  
of having two tables with same data?

Kind Regards,  
Badal

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ed08fd38-05e4-437a-a8e2-3295f2195e2a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed08fd38-05e4-437a-a8e2-3295f2195e2a%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [February 3, 2014, 10:44am UTC](https://discuss.elastic.co/t/elasticsearch-hadoop/15160/2 "2014-02-03T10:44:31Z")

</div>

There is no duplication per-se in HDFS. Hive tables are just 'views' of data - one sits unindexed, in raw format in HDFS  
the other one is indexed and analyzed in Elasticsearch.

You can't combine the two since they are completely different things - one is a file-system, the other one is a search  
and analytics engine.

On 09/01/2014 9:49 AM, Badal Mohapatra wrote:

> Hi,
> 
> ```
> To index Hadoop data into elasticsearch as I understand,
> 
> ```
> 
> We create an external table with essstorage handler and then copy the data from another internal hive table doesn't it  
> duplicate the data in HDFS?  
> Is there any way to use the hive internal tables directly to index instead of having two tables with same data?
> 
> Kind Regards,  
> Badal
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/ed08fd38-05e4-437a-a8e2-3295f2195e2a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed08fd38-05e4-437a-a8e2-3295f2195e2a%40googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/52EF730F.4060508%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/52EF730F.4060508%40gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:53am UTC](https://discuss.elastic.co/t/elasticsearch-hadoop/15160/3 "2017-07-06T01:53:10Z")

</div>


