# How do I build results dinamically from a Dataframe? (Apache Spark)

**URL:** <https://discuss.elastic.co/t/how-do-i-build-results-dinamically-from-a-dataframe-apache-spark/162366>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [December 28, 2018, 4:10pm UTC](https://discuss.elastic.co/t/how-do-i-build-results-dinamically-from-a-dataframe-apache-spark/162366 "2018-12-28T16:10:35Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![yeikel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yeikel/32/125434_2.png) [@yeikel](https://discuss.elastic.co/u/yeikel)\
**Post date:** [December 28, 2018, 4:10pm UTC](https://discuss.elastic.co/t/how-do-i-build-results-dinamically-from-a-dataframe-apache-spark/162366/1 "2018-12-28T16:10:35Z")

</div>

I have a Dataframe containing a list of cities as following :

```auto
val cities = sc.parallelize(Seq("New York")).toDF()

```

Now , for each city , I would like to query Elastic and build a set of results similar to the following logic :

```auto
val cities = sc.parallelize(Seq("New York")).toDF()
   cities.foreach(r => {
    val city = r.getString(0)
    val dfs = sqlContext.esDF("cities/docs", "?q=" + city) //returns a DataFrame which triggers the exception 
    })

```

Problem is that Spark does not allow nested operations that return dataframes. What options do I have to iterate a dataframe and get the results?

---

<div class="post-metadata">

**Author:** ![yeikel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yeikel/32/125434_2.png) [@yeikel](https://discuss.elastic.co/u/yeikel)\
**Post date:** [December 30, 2018, 4:33am UTC](https://discuss.elastic.co/t/how-do-i-build-results-dinamically-from-a-dataframe-apache-spark/162366/2 "2018-12-30T04:33:58Z")

</div>

Another option , is it possible to get a regular data structure that is not a dataframe using this connector?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [January 2, 2019, 4:48pm UTC](https://discuss.elastic.co/t/how-do-i-build-results-dinamically-from-a-dataframe-apache-spark/162366/3 "2019-01-02T16:48:57Z")

</div>

This seems like something better suited to using the regular java rest client to perform the search from within the `foreach` function. If you decide to do that, I would suggest doing `foreachPartition` instead of `foreach` so that you can batch up the query to send to Elasticsearch. Also, that would allow you to tear down the client after all the data in the partition is consumed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 30, 2019, 4:49pm UTC](https://discuss.elastic.co/t/how-do-i-build-results-dinamically-from-a-dataframe-apache-spark/162366/4 "2019-01-30T16:49:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
