# Multiple Elastic Queries per spark job

**URL:** <https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 17, 2017, 1:51pm UTC](https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115 "2017-05-17T13:51:42Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![mattmattmatt1](https://avatars.discourse-cdn.com/v4/letter/m/258eb7/32.png) [@mattmattmatt1](https://discuss.elastic.co/u/mattmattmatt1)\
**Post date:** [May 17, 2017, 1:51pm UTC](https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115/1 "2017-05-17T13:51:42Z")

</div>

HI All,

I've got a spark job that starts with a single elastic query. I want to use the result of this in a new query within the same spark job to pull down more data, based on what the first one brings back.

I've attempted to do this using the following code;

```
  val conf = new Conf()
  conf.set("es.query", query)

  val sc = new SparkContext(conf)

 // SPARK STUFF TO DETERMINE SECOND ELASTIC QUERY

 val newElasticQuery = ...
 val newConf = new Conf()
 newConf.set("es.query", newElasticQuery)

 val newSparkContext = SparkContext.getOrCreate(newConf)

```

The issue I have is this does not use the new query but the original conf, so the same data is pulled.

How would you go about doing this?

Cheers!

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [May 17, 2017, 7:28pm UTC](https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115/2 "2017-05-17T19:28:15Z")

</div>

You can only have one SparkContext active per JVM, so your last call to SparkContext.getOrCreate is getting the previously created spark context. You will need to specify these settings on an RDD by RDD basis. You should be able to pass those settings to the RDD create call via a Map.

---

<div class="post-metadata">

**Author:** ![mattmattmatt1](https://avatars.discourse-cdn.com/v4/letter/m/258eb7/32.png) [@mattmattmatt1](https://discuss.elastic.co/u/mattmattmatt1)\
**Post date:** [May 18, 2017, 9:46am UTC](https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115/3 "2017-05-18T09:46:30Z")

</div>

thanks for you reply! That makes sense, I misunderstood what getorCreate was actually doing.....

I'm still having trouble, can you give me an example of how you would apply this via a Map?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [May 18, 2017, 1:43pm UTC](https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115/4 "2017-05-18T13:43:04Z")

</div>

The last example in the scala portion of [this section](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html#spark-write) in the docs has an example of the syntax:

```auto
EsSpark.saveToEs(rdd, "index/type", Map("setting" -> "value"))

```

---

<div class="post-metadata">

**Author:** ![mattmattmatt1](https://avatars.discourse-cdn.com/v4/letter/m/258eb7/32.png) [@mattmattmatt1](https://discuss.elastic.co/u/mattmattmatt1)\
**Post date:** [May 19, 2017, 9:46am UTC](https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115/5 "2017-05-19T09:46:35Z")

</div>

Got it working, thank you very much!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 16, 2017, 9:46am UTC](https://discuss.elastic.co/t/multiple-elastic-queries-per-spark-job/86115/6 "2017-06-16T09:46:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
