# Spark, read data from ES, how to specify fields?

**URL:** <https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [December 17, 2015, 8:59am UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451 "2015-12-17T08:59:02Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [December 17, 2015, 8:59am UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/1 "2015-12-17T08:59:02Z")

</div>

I am trying to get some data from Spark to a remote elasticsearch cluster:

JavaPairRDD\<String, Map\<String, Object\>\> esRDD = JavaEsSpark.esRDD(jsc, "es\_articles", "?fields=title");

It doesnt seem to be implemented like "q" 😕

Any way to select fields to return? (network performance and workaround to fix no ISO datetime.)

Thanks,

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [December 18, 2015, 1:04pm UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/2 "2015-12-18T13:04:18Z")

</div>

You cannot apply projection since `fields` is internally used as well. For fine grained control over the mapping, consider using `DataFrame`s which are basically `RDD`s plus schema.

---

<div class="post-metadata">

**Author:** ![Ashish\_Belokar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashish_belokar/32/7064_2.png) [@Ashish\_Belokar](https://discuss.elastic.co/u/Ashish_Belokar)\
**Post date:** [January 8, 2016, 9:06am UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/3 "2016-01-08T09:06:03Z")

</div>

Here is what I did to get specific fields:

`QueryBuilder query = QueryBuilders.matchAllQuery(); List<String> fields = new ArrayList<String>(); fields.add("field1"); fields.add("field2"); JavaPairRDD<String, Map<String, Object>> esRDD = JavaEsSpark.esRDD(jsc, "index/type", (new SearchSourceBuilder().query(query) .fields(fields).toString()) ); System.out.println(esRDD.take(1));`

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [January 9, 2016, 2:06pm UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/4 "2016-01-09T14:06:04Z")

</div>

Using Elasticsearch to create such a basic query (to select 1-2 fields) is just wasteful. Simply add "fields" to the query as indicated [here](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-fields.html).

I'll reiterate my point though, an `RDD` with a schema is a Spark `DataFrame`. That provides not just fine control over the underlying structure but also pushed down operations - that is, the connector translating the SQL to an actual ES query.  
This documentation [section](https://www.elastic.co/guide/en/elasticsearch/hadoop/master/spark.html#spark-sql) provides more information.

Using an `RDD` while trying to select the fields and such, will not only reinvent parts of Spark SQL that are already available, but also provide only a subset and ignore all the other optimizations available.

---

<div class="post-metadata">

**Author:** ![Ashish\_Belokar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ashish_belokar/32/7064_2.png) [@Ashish\_Belokar](https://discuss.elastic.co/u/Ashish_Belokar)\
**Post date:** [January 10, 2016, 12:02pm UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/5 "2016-01-10T12:02:58Z")

</div>

Sounds like a great explanation for preferring Data Frames over RDDs. But, i already have the ES queries created. I guess in such cases the `es.mapping.include` and `es.mapping.exclude` properties in the [configuration](https://www.elastic.co/guide/en/elasticsearch/hadoop/master/configuration.html) must be used. However, this makes the configuration object specific to a particular index. I think I will have to move to DataFrames eventually!

---

<div class="post-metadata">

**Author:** ![mindon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mindon/32/13033_2.png) [@mindon](https://discuss.elastic.co/u/mindon)\
**Post date:** [November 8, 2016, 6:34am UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/6 "2016-11-08T06:34:17Z")

</div>

in scala i do like this

```
sc.esRDD("somedoc/sometype", "?q=something", Map[String, String]("es.read.field.include"->"field1,field2,..."))
```

---

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [April 12, 2017, 11:01am UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/7 "2017-04-12T11:01:52Z")

</div>

Yeah but it seems there is a bug with nested field:

I do:

```
scala> val df = sqlContext.read.format("org.elasticsearch.spark.sql").options(Map("es.scroll.limit" -> "100000", "es.read.field.include" -> "client.hash,client.token,name")).load("events-prod2/events")
df: org.apache.spark.sql.DataFrame = [client: struct<hash:string,token:string>, name: string]

```

Then get:

```
scala> df.first
res3: org.apache.spark.sql.Row = [[null,null],search]
```

---

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [April 12, 2017, 11:39am UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/8 "2017-04-12T11:39:45Z")

</div>

Running:

> sqlContext.read.format("org.elasticsearch.spark.sql").options(Map("es.scroll.limit" -\> "100", "es.query" -\> """{"fields" : ["client.hash"], "query" : {"match\_all" : {}}}""")).load("events-prod2/events")

Gives me all the fields (because es4Hadoop sends all fields in \_source query parameter, seen via an HTTP proxy)

With version 2.4.0 I got this error:

> Field 'client.hash' is backed by an array but the associated Spark Schema does not reflect this

---

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [April 12, 2017, 11:56am UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/9 "2017-04-12T11:56:34Z")

</div>

Thanks you, this is the only way I managed to use!

@costin I believe this is very important to select fields, because this save band-width, isnit ?

And in my case, I use it to perform terms aggregations and get ALL results (+ some join).

So I download from ES via RDD (bcoz I can select fields), transform it to DataFrame, and run SQL query at the end!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:22pm UTC](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451/10 "2017-07-06T13:22:13Z")

</div>


