# Reading from Elasticsearch to Spark is very slow

**URL:** <https://discuss.elastic.co/t/reading-from-elasticsearch-to-spark-is-very-slow/188307>\
**Category:** Elasticsearch\
**Created:** [July 1, 2019, 12:01pm UTC](https://discuss.elastic.co/t/reading-from-elasticsearch-to-spark-is-very-slow/188307 "2019-07-01T12:01:11Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Rami\_Batal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rami_batal/32/43696_2.png) [@Rami\_Batal](https://discuss.elastic.co/u/Rami_Batal)\
**Post date:** [July 1, 2019, 12:01pm UTC](https://discuss.elastic.co/t/reading-from-elasticsearch-to-spark-is-very-slow/188307/1 "2019-07-01T12:01:11Z")

</div>

Hi, I am using **Spark 2.4.0** , **Elasticsearch 6.6.2** , and **elasticsearch-spark-20\_2.11-6.8.1.jar** as connector.

I am running Spark local mode and configured the memory to 8GB.  
I have an Elasticsearch index with 14 million documents.

I want to load the whole index to a Spark DataFrame, so I am doing:

```
import org.apache.spark.sql.SQLContext        
import org.elasticsearch.spark.sql._

val sql = new SQLContext(sc)
val myDF = sql.esDF("my-index/my-type").cache()

println(myDF.count())

```

I can see the memory being filled bit by bit, which is expected because of the `cache()`, the memory is large enough to hold the entire data, but the process is extremely slow (over 2 hours).

Any hint is highly appreciated.

Thanks

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 29, 2019, 12:04pm UTC](https://discuss.elastic.co/t/reading-from-elasticsearch-to-spark-is-very-slow/188307/2 "2019-07-29T12:04:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
