# Load data into HDFS using ES-Spark

**URL:** <https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 6, 2015, 3:03pm UTC](https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297 "2015-05-06T15:03:31Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [May 14, 2015, 6:22am UTC](https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297/2 "2015-05-14T06:22:55Z")

</div>

There might be various reasons why the `saveAsTextFile` takes a long time - typically it might be because the parallelism is small (there's only one task handling it) or because the there's a large number of values (sometimes all) under the same key.  
What does you RDD looks like - any information on Spark during the wait and what it is doing? What's your hardware?

As for es-hadoop, in a nutshell it's a _connector_ between Elasticsearch and Hadoop so it likely fits the latter description.  
es-hadoop itself doesn't store any state, rather it helps data move between Elastic and Hadoop.

---

_[View the full topic](https://discuss.elastic.co/t/load-data-into-hdfs-using-es-spark/297)._
