# ES Aggregations in Spark

**URL:** <https://discuss.elastic.co/t/es-aggregations-in-spark/45588>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [March 28, 2016, 1:18pm UTC](https://discuss.elastic.co/t/es-aggregations-in-spark/45588 "2016-03-28T13:18:58Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![kucera.jan.cz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kucera.jan.cz/32/5869_2.png) [@kucera.jan.cz](https://discuss.elastic.co/u/kucera.jan.cz)\
**Post date:** [March 28, 2016, 1:18pm UTC](https://discuss.elastic.co/t/es-aggregations-in-spark/45588/1 "2016-03-28T13:18:58Z")

</div>

Hello everyone,

based on discussion about [ES use cases](https://discuss.elastic.co/t/use-cases-elasticsearch-and-spark/29746) I was wondering whether there is any way how Spark could benefit from ES aggregations and convert them to Dataframe. F.e something like:  
`val esConf = ... val esQuery = """{"agg" : {"my_agg" : {"terms" : {"field": "field_A"} } } }""" val jsonResult = client.search(esQuery, esConf) val transformer = ... val df = jsonResult.toDF(transformer) val result = df.filter(...).join(otherDf) ....`

A) Is there any plan to support something similar in ES roadmap?  
B) As I understood correctly how spark-es/hadoop-es works library based on scroll's json results "detects" dataframe's schema. Can you direct me to classes which is responsible for this detection? I was wondering whether these components could be used for building 'transformer' I had in my example.

Thanks advance for your suggestions  
-Jan

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [April 5, 2016, 2:47pm UTC](https://discuss.elastic.co/t/es-aggregations-in-spark/45588/2 "2016-04-05T14:47:53Z")

</div>

A) Aggregations are not currently supported by ES-Hadoop; it's the next major item on the roadmap.  
B) All the spark SQL classes reside under their dedicated package:

> <https://github.com/elastic/elasticsearch-hadoop/blob/master/spark/sql-13/src/main/scala/org/elasticsearch/spark/sql/DefaultSource.scala>

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:25pm UTC](https://discuss.elastic.co/t/es-aggregations-in-spark/45588/3 "2017-07-06T13:25:29Z")

</div>


