# ESHadoop - Hadoop vs Spark

**URL:** <https://discuss.elastic.co/t/eshadoop-hadoop-vs-spark/47509>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [April 15, 2016, 2:30pm UTC](https://discuss.elastic.co/t/eshadoop-hadoop-vs-spark/47509 "2016-04-15T14:30:11Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pat\_Humphreys](https://avatars.discourse-cdn.com/v4/letter/p/ccd318/32.png) [@Pat\_Humphreys](https://discuss.elastic.co/u/Pat_Humphreys)\
**Post date:** [April 15, 2016, 2:30pm UTC](https://discuss.elastic.co/t/eshadoop-hadoop-vs-spark/47509/1 "2016-04-15T14:30:11Z")

</div>

I have a current Hadoop job running on AWS EMR, running ESHadoop with Cascading. It does bulk inserts of 10,000 4k records about 300M of them.  
I was wondering would there be any speed benefits of using Spark instead?

---

<div class="post-metadata">

**Author:** ![jhendric98](https://avatars.discourse-cdn.com/v4/letter/j/90ced4/32.png) [@jhendric98](https://discuss.elastic.co/u/jhendric98)\
**Post date:** [April 18, 2016, 7:59pm UTC](https://discuss.elastic.co/t/eshadoop-hadoop-vs-spark/47509/2 "2016-04-18T19:59:24Z")

</div>

Without hearing more about your job, I'll have to relate my general experience. We've found Spark to reduce runtimes on jobs over traditional MR in indexing to Elasticsearch. I am not an Elasticsearch expert but it seems data locality may play a part. We used a custom jar loader in a YARN job to load data and have replaced ours with the ES-Hadoop Spark library.

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [April 21, 2016, 6:08am UTC](https://discuss.elastic.co/t/eshadoop-hadoop-vs-spark/47509/3 "2016-04-21T06:08:35Z")

</div>

A big advantage that Spark SQL gives over other libraries, it that it allows push down - that is in Spark SQL the operations executed can be detected and thus pushed down by 3rd party plugins (like ES-Hadoop). This significantly reduces the amount of data that needs to be pulled in from ES.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:25pm UTC](https://discuss.elastic.co/t/eshadoop-hadoop-vs-spark/47509/4 "2017-07-06T13:25:01Z")

</div>


