# Load data into HDFS using ES-Spark

**URL:** <https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 6, 2015, 3:03pm UTC](https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297 "2015-05-06T15:03:31Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lucas\_Weissert](https://avatars.discourse-cdn.com/v4/letter/l/22d042/32.png) [@Lucas\_Weissert](https://discuss.elastic.co/u/Lucas_Weissert)\
**Post date:** [May 6, 2015, 3:03pm UTC](https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297/1 "2015-05-06T15:03:31Z")

</div>

Hello,

I am reading data from elasticsearch using spark (ES-Spark). After i get the data using sc.esRDD(".../...") I want to store everything in HDFS so I use the saveAsTextFile method but it is very slow ...  
Am I doing the right things ? It takes 15min to save (and it is saving 11Go)

Es-Hadoop can be used to store data in HDFS or it is just used to write into ES or read data from ES and display some queries ?

Best regards

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [May 14, 2015, 6:22am UTC](https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297/2 "2015-05-14T06:22:55Z")

</div>

There might be various reasons why the `saveAsTextFile` takes a long time - typically it might be because the parallelism is small (there's only one task handling it) or because the there's a large number of values (sometimes all) under the same key.  
What does you RDD looks like - any information on Spark during the wait and what it is doing? What's your hardware?

As for es-hadoop, in a nutshell it's a _connector_ between Elasticsearch and Hadoop so it likely fits the latter description.  
es-hadoop itself doesn't store any state, rather it helps data move between Elastic and Hadoop.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:28pm UTC](https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297/3 "2017-07-06T13:28:25Z")

</div>


