# Exception when using Elasticsearch-spark and Elasticsearch-core together

**URL:** <https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [January 6, 2016, 7:29am UTC](https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471 "2016-01-06T07:29:12Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![dynamicscope](https://avatars.discourse-cdn.com/v4/letter/d/e9bcb4/32.png) [@dynamicscope](https://discuss.elastic.co/u/dynamicscope)\
**Post date:** [January 6, 2016, 7:29am UTC](https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471/1 "2016-01-06T07:29:13Z")

</div>

Hello, Elastic users.

I am trying to use Elasticsearch-spark and Elasticsearch-core together for my scala(sbt) spark app.

here is the **build.sbt**.

```
scalaVersion := "2.10.6"

val sparkVers = "1.5.1"

resolvers ++= Seq(
  "Typesafe Releases" at "http://repo.typesafe.com/typesafe/releases/",
  "Akka Repository" at "http://repo.akka.io/releases/",
  "Sonatype OSS Snapshots" at "https://oss.sonatype.org/content/repositories/snapshots")

// Base Spark-provided dependencies
libraryDependencies ++= Seq(
  "org.apache.spark" %% "spark-core" % sparkVers % "provided",
  "org.apache.spark" %% "spark-streaming" % sparkVers % "provided")

// Elasticsearch integration
libraryDependencies ++= Seq(
  ("org.elasticsearch" % "elasticsearch-spark_2.10" % "2.2.0-beta1").
    exclude("org.apache.hadoop", "hadoop-yarn-api").
    exclude("org.eclipse.jetty.orbit", "javax.mail.glassfish").
    exclude("org.eclipse.jetty.orbit", "javax.servlet").
    exclude("org.slf4j", "slf4j-api")
)

// Elasticsearch core
libraryDependencies += "org.elasticsearch" % "elasticsearch" % "2.1.1"

// Skip tests when assembling fat JAR
test in assembly := {}

// Exclude jars that conflict with Spark (see https://github.com/sbt/sbt-assembly)
libraryDependencies ~= { _ map {
  case m if Seq("org.elasticsearch").contains(m.organization) =>
    m.exclude("commons-logging", "commons-logging").
      exclude("commons-collections", "commons-collections").
      exclude("commons-beanutils", "commons-beanutils-core").
      exclude("com.esotericsoftware.minlog", "minlog").
      exclude("joda-time", "joda-time").
      exclude("org.apache.commons", "commons-lang3").
      exclude("com.google.guava", "guava")
  case m => m
}}

dependencyOverrides += "org.scala-lang" % "scala-compiler" % scalaVersion.value
dependencyOverrides += "org.scala-lang" % "scala-library" % scalaVersion.value
dependencyOverrides += "commons-net" % "commons-net" % "3.1"

assemblyMergeStrategy in assembly := {
  case PathList("org", "apache", "spark", "unused", "UnusedStubClass.class") => MergeStrategy.discard
  case x =>
    val oldStrategy = (assemblyMergeStrategy in assembly).value
    oldStrategy(x)
}

```

And the following is my very simple main class.

```
object Main {

  def main(args: Array[String]): Unit = {
    val settings = Settings.settingsBuilder.put("cluster.name", "Avengers").put("client.transport.sniff", true).build
    TransportClient.builder.settings(settings).build.addTransportAddress(new InetSocketTransportAddress(InetAddress.getByName("127.0.0.1"), 9300))
    
    val conf = new SparkConf()
    val sc = new SparkContext(conf)
    
    sc.esJsonRDD("*")
      .collect()
      .foreach(println)

  sc.stop()
  }
}

```

If I make this as a fat jar (sbt assembly) and spark-submit it to the spark cluster, I get the following exception.

```
Exception in thread "main" java.lang.NoClassDefFoundError: com/google/common/collect/ImmutableMap
	at Main$.main(Main.scala:13)
	at Main.main(Main.scala)
	at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
	at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)
	at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
	at java.lang.reflect.Method.invoke(Method.java:497)
	at org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:672)
	at org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:180)
	at org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:205)
	at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:120)
	at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)
Caused by: java.lang.ClassNotFoundException: com.google.common.collect.ImmutableMap
	at java.net.URLClassLoader.findClass(URLClassLoader.java:381)
	at java.lang.ClassLoader.loadClass(ClassLoader.java:424)
	at java.lang.ClassLoader.loadClass(ClassLoader.java:357)
	... 11 more

```

However, without the following lines, everything works fine.

```
val settings = Settings.settingsBuilder.put("cluster.name", "Avengers").put("client.transport.sniff", true).build
TransportClient.builder.settings(settings).build.addTransportAddress(new InetSocketTransportAddress(InetAddress.getByName("127.0.0.1"), 9300))

```

Does anyone confront any similar problem?  
I need to use **TransportClient** in order to manage index and cluster while MapReducing with esRDD.

Is it even possible to use Elasticsearch-spark and Elasticsearch-core together?

Any hint would be very appreciated.  
Thank you.

---

<div class="post-metadata">

**Author:** ![eliasah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eliasah/32/34741_2.png) [@eliasah](https://discuss.elastic.co/u/eliasah)\
**Post date:** [January 6, 2016, 9:48am UTC](https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471/2 "2016-01-06T09:48:00Z")

</div>

Did you try using the elasticsearch-hadoop [development snapshot](https://github.com/elastic/elasticsearch-hadoop#development-snapshot) ?

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [January 6, 2016, 2:46pm UTC](https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471/3 "2016-01-06T14:46:58Z")

</div>

Es-spark/Hadoop does not depend on Guava. Hadoop does as does Elastic. However Hadoop depends on an ancient version of Guava while ES does not. Try upgrading Guava to the same version as elastic or, if that doesn't work, try shading it (see the elastic blog site).

---

<div class="post-metadata">

**Author:** ![dynamicscope](https://avatars.discourse-cdn.com/v4/letter/d/e9bcb4/32.png) [@dynamicscope](https://discuss.elastic.co/u/dynamicscope)\
**Post date:** [January 7, 2016, 5:35am UTC](https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471/4 "2016-01-07T05:35:08Z")

</div>

Thank you for your hints.

I solved this issue by adding

`exclude("org.apache.spark", "spark-network-common_2.10")`

instead of

`exclude("com.google.guava", "guava")`

It looks like **spark-network-common\_2.10** contains guava classes.  
So the finally **build.sbt** includes the following...

```
// Exclude jars that conflict with Spark (see https://github.com/sbt/sbt-assembly)
libraryDependencies ~= { _ map {
  case m if Seq("org.elasticsearch").contains(m.organization) =>
    m.exclude("commons-logging", "commons-logging").
      exclude("commons-collections", "commons-collections").
      exclude("commons-beanutils", "commons-beanutils-core").
      exclude("com.esotericsoftware.minlog", "minlog").
      exclude("joda-time", "joda-time").
      exclude("org.apache.commons", "commons-lang3").
      exclude("org.apache.spark", "spark-network-common_2.10")
  case m => m
}}
```

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [January 9, 2016, 1:39pm UTC](https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471/5 "2016-01-09T13:39:25Z")

</div>

Thanks for the update.

If you have time, can you raise this with the Spark team - having unshaded google guava classes in a different jar is bad practice since it's hard to track it down. From what I can [tell](http://mvnrepository.com/artifact/org.apache.spark/spark-network-common_2.10/1.2.0), the jar itself has some `com.google.common` classes internally (bad) and a dependency on guava.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:26pm UTC](https://discuss.elastic.co/t/exception-when-using-elasticsearch-spark-and-elasticsearch-core-together/38471/6 "2017-07-06T13:26:38Z")

</div>


