# Cant get Spark to actually retrieve data

**URL:** <https://discuss.elastic.co/t/cant-get-spark-to-actually-retrieve-data/90310>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [June 21, 2017, 2:26pm UTC](https://discuss.elastic.co/t/cant-get-spark-to-actually-retrieve-data/90310 "2017-06-21T14:26:46Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![gzz](https://avatars.discourse-cdn.com/v4/letter/g/b77776/32.png) [@gzz](https://discuss.elastic.co/u/gzz)\
**Post date:** [June 21, 2017, 2:26pm UTC](https://discuss.elastic.co/t/cant-get-spark-to-actually-retrieve-data/90310/1 "2017-06-21T14:26:47Z")

</div>

Hi,  
I've setup hadoop+spark 1.6 via CDH and added the latest Zeppelin. Here I've included the latest es-hadoop binding and am now trying to just load some data from my ES cluster. While it retrieves the mapping and also issues the query, it immediately deletes the scroll id after the query without ever getting any data. Consequently, I end up having the schema but not data in Zeppelin.

I'm really out of ideas here, can anyone help?! Thank you!

Here are some of my queries:

> %spark  
> import org.apache.spark.sql.SQLContext  
> import org.elasticsearch.spark.sql.\_

> var sql = new org.apache.spark.sql.SQLContext(sc)  
> sql.esDF("logstash-2017.06.21/logs",Map(  
> "es.nodes" -\> "my.host.name",  
> "es.read.field.include" -\> "host")).registerTempTable("logs")  
> z.show(sql.sql("select count(host) from logs"))

returns 0

Or even simpler:

> %spark  
> sqlContext.read.format("es").load("logstash-\*/logs").limit(10).show()

Returns a nice table with a bunch of columns, but empty

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [June 21, 2017, 2:35pm UTC](https://discuss.elastic.co/t/cant-get-spark-to-actually-retrieve-data/90310/2 "2017-06-21T14:35:20Z")

</div>

Could you increase the logging of the `org.elasticsearch.hadoop.rest.commonshttp` package to `TRACE` and post all the relevant logs here? That should help us track down what the connector is receiving.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 19, 2017, 2:35pm UTC](https://discuss.elastic.co/t/cant-get-spark-to-actually-retrieve-data/90310/3 "2017-07-19T14:35:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
