# Comparison between C\*/Spark and ES/Spark concerning data locality

**URL:** <https://discuss.elastic.co/t/comparison-between-c-spark-and-es-spark-concerning-data-locality/63595>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [October 21, 2016, 9:54am UTC](https://discuss.elastic.co/t/comparison-between-c-spark-and-es-spark-concerning-data-locality/63595 "2016-10-21T09:54:21Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![chernals](https://avatars.discourse-cdn.com/v4/letter/c/a3d4f5/32.png) [@chernals](https://discuss.elastic.co/u/chernals)\
**Post date:** [October 21, 2016, 9:54am UTC](https://discuss.elastic.co/t/comparison-between-c-spark-and-es-spark-concerning-data-locality/63595/1 "2016-10-21T09:54:21Z")

</div>

Hey guys,

Following up on my question on Github ([https://github.com/elastic/elasticsearch-hadoop/pull/819#issuecomment-255216112](https://github.com/elastic/elasticsearch-hadoop/pull/819#issuecomment-255216112)), it seems that the data locality is back for ES/Spark, which is great.

I don't know enough to get a clear view on how that compares with the Spark-Cassandra connector. Would anyone be able to provide more info on that?

Thanks

---

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [October 24, 2016, 6:50pm UTC](https://discuss.elastic.co/t/comparison-between-c-spark-and-es-spark-concerning-data-locality/63595/2 "2016-10-24T18:50:18Z")

</div>

Can you be more precise, you have already a Cassandra cluster with Spark, and you want to know how this is works with elasticsearch?

---

<div class="post-metadata">

**Author:** ![chernals](https://avatars.discourse-cdn.com/v4/letter/c/a3d4f5/32.png) [@chernals](https://discuss.elastic.co/u/chernals)\
**Post date:** [October 26, 2016, 10:37pm UTC](https://discuss.elastic.co/t/comparison-between-c-spark-and-es-spark-concerning-data-locality/63595/3 "2016-10-26T22:37:39Z")

</div>

I already used C\* with Spark, but I have very limited experience with ES and/or ES with Spark.

So the questions could be

- is there a good source information of the underlying implementation, how Spark workers will query the ES nodes, etc.

- I couldn't find performance analysis of the way the architecture would scale in term of number of Spark and ES nodes

Thanks a lot !

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:22pm UTC](https://discuss.elastic.co/t/comparison-between-c-spark-and-es-spark-concerning-data-locality/63595/4 "2017-07-06T13:22:47Z")

</div>


