# Error saving Spark RDD using rdd.saveToEs

**URL:** <https://discuss.elastic.co/t/error-saving-spark-rdd-using-rdd-savetoes/96888>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [August 13, 2017, 9:53pm UTC](https://discuss.elastic.co/t/error-saving-spark-rdd-using-rdd-savetoes/96888 "2017-08-13T21:53:46Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Antonio\_Ye](https://avatars.discourse-cdn.com/v4/letter/a/e95f7d/32.png) [@Antonio\_Ye](https://discuss.elastic.co/u/Antonio_Ye)\
**Post date:** [August 13, 2017, 9:53pm UTC](https://discuss.elastic.co/t/error-saving-spark-rdd-using-rdd-savetoes/96888/1 "2017-08-13T21:53:46Z")

</div>

I have a simple Spark data frame which contains a column of JSON strings. When I run the following code:

import org.elasticsearch.spark.\_  
val df = Seq("""{"name": "john", "age":44}""").toDF("json")  
df.rdd.map(x=\>x.getAs[String]("json")).saveToEs("test-index-df/post")

I get the error below:

17/08/13 21:51:58 ERROR TaskContextImpl: Error in TaskCompletionListener  
org.elasticsearch.hadoop.rest.EsHadoopInvalidRequest: Found unrecoverable error [100.127.0.5:9200] returned Bad Request(400) - failed to parse; Bailing out..  
at org.elasticsearch.hadoop.rest.RestClient.processBulkResponse(RestClient.java:251)  
at org.elasticsearch.hadoop.rest.RestClient.bulk(RestClient.java:203)  
at org.elasticsearch.hadoop.rest.RestRepository.tryFlush(RestRepository.java:220)  
at org.elasticsearch.hadoop.rest.RestRepository.flush(RestRepository.java:242)  
at org.elasticsearch.hadoop.rest.RestRepository.close(RestRepository.java:267)  
at org.elasticsearch.hadoop.rest.RestService$PartitionWriter.close(RestService.java:120)  
at org.elasticsearch.spark.rdd.EsRDDWriter$$anonfun$write$1.apply(EsRDDWriter.scala:60)  
at org.elasticsearch.spark.rdd.EsRDDWriter$$anonfun$write$1.apply(EsRDDWriter.scala:60)  
at org.apache.spark.TaskContext$$anon$1.onTaskCompletion(TaskContext.scala:123)  
at org.apache.spark.TaskContextImpl$$anonfun$markTaskCompleted$1.apply(TaskContextImpl.scala:97)  
at org.apache.spark.TaskContextImpl$$anonfun$markTaskCompleted$1.apply(TaskContextImpl.scala:95)  
at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)  
at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)

Anyone have any idea what I am doing wrong?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [August 15, 2017, 5:13pm UTC](https://discuss.elastic.co/t/error-saving-spark-rdd-using-rdd-savetoes/96888/2 "2017-08-15T17:13:03Z")

</div>

I would make sure that the JSON you are saving does not contain any non printing characters or unicode characters in standard json token locations (Like special quote characters)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 12, 2017, 5:13pm UTC](https://discuss.elastic.co/t/error-saving-spark-rdd-using-rdd-savetoes/96888/3 "2017-09-12T17:13:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
