# Ingesting data from HDFS to ElasticSearch

**URL:** <https://discuss.elastic.co/t/ingesting-data-from-hdfs-to-elasticsearch/71507>\
**Category:** Elasticsearch\
**Created:** [January 13, 2017, 10:21am UTC](https://discuss.elastic.co/t/ingesting-data-from-hdfs-to-elasticsearch/71507 "2017-01-13T10:21:53Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![mkarthikswamy](https://avatars.discourse-cdn.com/v4/letter/m/ea666f/32.png) [@mkarthikswamy](https://discuss.elastic.co/u/mkarthikswamy)\
**Post date:** [January 13, 2017, 10:21am UTC](https://discuss.elastic.co/t/ingesting-data-from-hdfs-to-elasticsearch/71507/1 "2017-01-13T10:21:53Z")

</div>

Hi,  
We have a Hadoop cluster in which a Spark job creates about 3 GB of data every hour in hour partitioned directories in HDFS. This data needs to be moved to ElasticSearch in the most efficient way. The following are the options we are considering:

1. In the existing Spark job that writes the data to HDFS, add another stage to write the Data Frame to ElasticSearch - using ES-Hadoop connector. This bypasses one extra disk i/o if we write to HDFS & read from it again to store in ES. On the other hand, it increases the execution time of the hourly Spark job
2. Have another Spark job get triggered after the first Spark job gets completed, to push the data from HDFS to ES using ES-Hadoop connector
3. Is there a way to use logstash here? Hourly batch ingest using Spark ES-Hadoop looks most optimal from performance point of view, but does logstash support a hdfs input plugin, I didn't see that in the Elastic site? Even if it does, can it scale well enough?
4. In order to reduce the data transfer over the wire, is it better to overlay the ElasticSearch cluster on the HDFS/Spark cluster?  
Any suggestions, pointers will be very helpful, Thanks a lot!  
MK

---

<div class="post-metadata">

**Author:** ![mainec](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mainec/32/5557_2.png) [@mainec](https://discuss.elastic.co/u/mainec)\
**Post date:** [January 18, 2017, 11:39am UTC](https://discuss.elastic.co/t/ingesting-data-from-hdfs-to-elasticsearch/71507/2 "2017-01-18T11:39:51Z")

</div>

I'd suggest taking a look at the spark-elasticsearch connector described here:

[https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html)

though @costin probably has more information for you.

Hope this helps,  
Isabel

---

<div class="post-metadata">

**Author:** ![mkarthikswamy](https://avatars.discourse-cdn.com/v4/letter/m/ea666f/32.png) [@mkarthikswamy](https://discuss.elastic.co/u/mkarthikswamy)\
**Post date:** [January 18, 2017, 3:44pm UTC](https://discuss.elastic.co/t/ingesting-data-from-hdfs-to-elasticsearch/71507/3 "2017-01-18T15:44:47Z")

</div>

@mainec, thanks for your response.  
Yes, es-hadoop is the connector we are using in Spark

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 15, 2017, 3:44pm UTC](https://discuss.elastic.co/t/ingesting-data-from-hdfs-to-elasticsearch/71507/4 "2017-02-15T15:44:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
