# Spark code to get select firelds from ES

**URL:** <https://discuss.elastic.co/t/spark-code-to-get-select-firelds-from-es/101179>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [September 20, 2017, 2:00pm UTC](https://discuss.elastic.co/t/spark-code-to-get-select-firelds-from-es/101179 "2017-09-20T14:00:07Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![kedarsdixit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kedarsdixit/32/20502_2.png) [@kedarsdixit](https://discuss.elastic.co/u/kedarsdixit)\
**Post date:** [September 20, 2017, 2:00pm UTC](https://discuss.elastic.co/t/spark-code-to-get-select-firelds-from-es/101179/1 "2017-09-20T14:00:07Z")

</div>

Hi,

I want to get only select fields from ES using Spark ES connector.

I have done some code which is fetching all the documents matching given index as below:

JavaPairRDD\<String, Map\<String, Object\>\> esRDD = JavaEsSpark.esRDD(jsc, searchIndex);

However, is there a way to only get specific fields from documents for every index in ES than getting everything ?

Example: Let's say, I have many fields in the documents as below and I have @timestamp which is also a field in the response { .............., @timestamp=Fri Jul 07 01:36:00 IST 2017, ..............}, Here how can I get the only field @timestamp for all my indexes ?

I could see something [here](https://discuss.elastic.co/t/spark-read-data-from-es-how-to-specify-fields/37451) but unable to correlate. can someone help me please ?

Many Thanks!  
~KD

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [October 4, 2017, 3:38am UTC](https://discuss.elastic.co/t/spark-code-to-get-select-firelds-from-es/101179/2 "2017-10-04T03:38:02Z")

</div>

@kedarsdixit If you are using Spark SQL - We provide a native integration with Spark SQL that allows you to push down predicate filters and field projections directly to Elasticsearch (i.e. if you `SELECT timestamp FROM ...` then the connector will recognize that this field is the only one needed and will only return the `timestamp` field to the executors processing the data.)

Alternatively, if you are using vanilla Spark RDDs that do not support query planning and schema optimizations like Spark SQL does, we provide a configuration that you can set with the names of the fields you would like to return from the cluster (see `es.read.source.filter` in the [docs](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/configuration.html#configuration-options-index).

Hopefully that helps!

---

<div class="post-metadata">

**Author:** ![kedarsdixit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kedarsdixit/32/20502_2.png) [@kedarsdixit](https://discuss.elastic.co/u/kedarsdixit)\
**Post date:** [October 4, 2017, 4:13pm UTC](https://discuss.elastic.co/t/spark-code-to-get-select-firelds-from-es/101179/3 "2017-10-04T16:13:21Z")

</div>

thanks @james.baiera well I am using the SparkEs connector and I could figure out the way to select the specific fields. Many Thanks! ~Kedar

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 1, 2017, 4:13pm UTC](https://discuss.elastic.co/t/spark-code-to-get-select-firelds-from-es/101179/4 "2017-11-01T16:13:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
