# Writing spark Dataframe/Dataset to Elasticsearch

**URL:** <https://discuss.elastic.co/t/writing-spark-dataframe-dataset-to-elasticsearch/131023>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 8, 2018, 2:03pm UTC](https://discuss.elastic.co/t/writing-spark-dataframe-dataset-to-elasticsearch/131023 "2018-05-08T14:03:34Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![wandermonk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wandermonk/32/27338_2.png) [@wandermonk](https://discuss.elastic.co/u/wandermonk)\
**Post date:** [May 8, 2018, 2:03pm UTC](https://discuss.elastic.co/t/writing-spark-dataframe-dataset-to-elasticsearch/131023/1 "2018-05-08T14:03:34Z")

</div>

I am trying to write a JavaRDD to elasticsearch using the saveToES() method. But, we are getting the exception

`EsHadoopIllegalArgumentException: Cannot detect ES version - typically this happens if the network/Elasticsearch cluster is not accessible or when targeting a WAN/Cloud instance without the proper setting 'es.nodes.wan.only'`

Code:

```
public static void main(String[] args) {
		SparkConf conf = new SparkConf()
				.set("es.nodes", "cstg-01")
				.set("es.port", "9200")
				.set("es.scheme", "http")
				.set("spark.sql.wareh0use.dir", "/app/SmartAnalytics")
				.set("spark.serializer",
						"org.apache.spark.serializer.KryoSerializer");
		SparkSession spark = SparkSession.builder().appName("Spark Job")
				.config(conf).enableHiveSupport().getOrCreate();
		// spark.sql("SELECT * FROM sora_feed.sora_sr_details_temp").show(10);
		/*
		 * 
		 Dataset<Row> sqlDF = spark
				.sql("Select * from service_request_transformed.sr_bug_detail_text");
		sqlDF.printSchema();

		Dataset<Row> modifiedSR = sqlDF.
				select(col("*"),struct("incident_number")
						.as("srdetails"))
						 .groupBy("identifier","severity","project","product","original_version","os_version","version_text","os_type","headline","description","project_release_text","integrated_releases_text","known_fixed_release_text","status","duplicate_of" ,"verified_release_text","apply_to_text","to_be_fixed_text","found","submitted_on","last_mod_on","failed_release_text").agg(collect_list(struct("srdetails"))).coalesce(1);
		modifiedSR.write().format("json")
				.save("/app/SmartAnalytics/Apps/TestData/Bug_2_SR");*/
		
		
		Dataset<Row> sqlDF = spark
				.sql("Select * from service_request_transformed.sr_denorm_defects");
		sqlDF.printSchema();
		Dataset<Row> modifiedSR = sqlDF.
				select(col("*"),struct("incident_number")
						.as("srdetails"))
						 .groupBy("defect_number","defect_title","defect_submitted_on").agg(collect_list(struct("srdetails"))).coalesce(1);
		modifiedSR.write().format("json")
				.save("/app/SmartAnalytics/Apps/TestData/Defects_2_SR");

		modifiedSR.printSchema();

		JavaRDD<Row> srRDD = modifiedSR.toJavaRDD();
		
		
		JavaEsSpark.saveToEs(srRDD, "/sr2bug/data" , ImmutableMap.of("es.mapping.id" , "defect_number"));//indexpath 
		
		/*
		 * Map<String, String> params = Collections.emptyMap(); HttpEntity
		 * entity = new NStringEntity(inputJson, ContentType.APPLICATION_JSON);
		 * try { Response response = context.getContext().performRequest("POST",
		 * es_index + documentID.toString(), params, entity);
		 * logger.info("The json record is successfully indexed :: "
		 * +response.getEntity()); } catch (Exception e) { e.printStackTrace();
		 * }
		 */

	}
```

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [May 30, 2018, 7:28pm UTC](https://discuss.elastic.co/t/writing-spark-dataframe-dataset-to-elasticsearch/131023/2 "2018-05-30T19:28:49Z")

</div>

Do you have any more logs for that failure? This failure can have any number of reasons for occurring, ranging from a lack of route to the ES hosts, or errors in the local network layer (too many connections, incorrect SSL settings).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 27, 2018, 7:28pm UTC](https://discuss.elastic.co/t/writing-spark-dataframe-dataset-to-elasticsearch/131023/3 "2018-06-27T19:28:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
