# Use Spark to index data in HDFS

**URL:** <https://discuss.elastic.co/t/use-spark-to-index-data-in-hdfs/34842>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [November 17, 2015, 7:54pm UTC](https://discuss.elastic.co/t/use-spark-to-index-data-in-hdfs/34842 "2015-11-17T19:54:45Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![swapnilnarlawar](https://avatars.discourse-cdn.com/v4/letter/s/ebca7d/32.png) [@swapnilnarlawar](https://discuss.elastic.co/u/swapnilnarlawar)\
**Post date:** [November 17, 2015, 7:54pm UTC](https://discuss.elastic.co/t/use-spark-to-index-data-in-hdfs/34842/1 "2015-11-17T19:54:45Z")

</div>

Hi there,  
We are looking at simplest and fastest way to get data from HDFS to ES.  
One method we have been trying and having out of luck is ES- Hadoop with Spark  
Component versions we are using are below  
Versions;  
ES - 1.7.3  
Spark - 1.5.2  
Scala - 2.10.4  
JAVA - 1.7.0\_67

-#We initiate spark shell with following jar files.

./spark-shell --jars esjava/elasticsearch-spark\_2.11-2.1.2.jar esjava/elasticsearch-hadoop-mr-2.1.2.jar elasticsearch-hadoop-2.1.2.jar

-#Below we import following classes

import org.apache.spark.SparkContext  
import org.apache.spark.SparkContext.\_  
import org.elasticsearch.spark.\_  
import org.apache.spark.SparkConf  
import org.elasticsearch.spark.rdd.EsSpark  
import org.apache.spark.sql.\_  
import org.apache.spark.sql.types.\_  
import org.apache.spark.rdd.RDD  
import org.elasticsearch.spark.sql.\_  
import org.apache.spark.sql.SQLContext  
import org.apache.spark.sql.SQLContext.\_

-# Below, we define our elastic search master

val conf = new SparkConf()  
conf.set("es.nodes","hostname:9200")

-# Below we point json file and convert it to dataframe  
val sqlContext = new SQLContext(sc)  
val df = sqlContext.jsonFile("hdfs://namenode/tmp/2015-11-10.json")

-# Below We validate the schema  
println(df.printSchema)

-# and below save it to Elastic  
df.saveToEs("test/parquet")

-#And right after that where we get following error, not sure what we are doing wrong.

Error  
java.lang.NoSuchMethodError: scala.Predef$.ArrowAssoc(Ljava/lang/Object;)Ljava/lang/Object;  
at org.elasticsearch.spark.sql.EsSparkSQL$.saveToEs(EsSparkSQL.scala:42)  
at org.elasticsearch.spark.sql.package$SparkDataFrameFunctions.saveToEs(package.scala:25)

Any help is appreciated.

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [November 23, 2015, 9:45pm UTC](https://discuss.elastic.co/t/use-spark-to-index-data-in-hdfs/34842/2 "2015-11-23T21:45:30Z")

</div>

`./spark-shell --jars esjava/elasticsearch-spark_2.11-2.1.2.jar esjava/elasticsearch-hadoop-mr-2.1.2.jar elasticsearch-hadoop-2.1.2.jar`

You are setting the classpath incorrectly - you are pulling in 3 different jars, that overlap in functionality and packages for no reason at all. Use only `elasticsearch-spark` as indicated by the [docs](https://www.elastic.co/guide/en/elasticsearch/hadoop/master/install.html#install).

In addition, you are using the elasticsearch-spark compiled for Scala 2.11 while using Scala 2.10.  
These details indicate you are fairly new to Scala and Spark and are skipping simple yet critical details in the setup. Stop rushing, take a step back and start again paying attention to details - it might seem that you are moving slowly but getting stuck on bugs like these, is likely going to burn more time.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:27pm UTC](https://discuss.elastic.co/t/use-spark-to-index-data-in-hdfs/34842/3 "2017-07-06T13:27:07Z")

</div>


