# \[Spark\] Unable to index JSON from HDFS using SchemaRDD.saveToES()

**URL:** <https://discuss.elastic.co/t/spark-unable-to-index-json-from-hdfs-using-schemardd-savetoes/22275>\
**Category:** Elasticsearch\
**Created:** [February 19, 2015, 9:15pm UTC](https://discuss.elastic.co/t/spark-unable-to-index-json-from-hdfs-using-schemardd-savetoes/22275 "2015-02-19T21:15:43Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![m\_shirley](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/m_shirley/32/55253_2.png) [@m\_shirley](https://discuss.elastic.co/u/m_shirley)\
**Post date:** [February 19, 2015, 9:15pm UTC](https://discuss.elastic.co/t/spark-unable-to-index-json-from-hdfs-using-schemardd-savetoes/22275/1 "2015-02-19T21:15:43Z")

</div>

This is my first real attempt at spark/scala so be gentle.

I have a file called test.json on HDFS that I'm trying to read and index  
using Spark. I'm able to read the file via SQLContext.jsonFile() but when  
I try to use SchemaRDD.saveToEs() I get an invalid JSON fragment received  
error. I'm thinking that the saveToES() function isn't actually formatting  
the output in json and instead is just sending the value field of the RDD.

What am I doing wrong?

Spark 1.2.0  
Elasticsearch-hadoop 2.1.0.BUILD-20150217

test.json:  
{"key":"value"}

spark-shell:  
import org.apache.spark.SparkContext.\_  
import org.elasticsearch.spark.\_

val sqlContext = new org.apache.spark.sql.SQLContext(sc)  
import sqlContext.\_

val input =  
sqlContext.jsonFile("hdfs://nameservice1/user/mshirley/test.json")

input.saveToEs("mshirley\_spark\_test/test")

error:  
  
org.elasticsearch.hadoop.rest.EsHadoopInvalidRequest: Found unrecoverable  
error [Bad Request(400) - Invalid JSON fragment  
received[["value"]][MapperParsingException[failed to parse]; n  
ested: ElasticsearchParseException[Failed to derive xcontent from  
(offset=13, length=9): [123, 34, 105, 110, 100, 101, 120, 34, 58, 123, 125,  
125, 10, 91, 34, 118, 97, 108, 117, 101, 3  
4, 93, 10]]; ]]; Bailing out..

input:  
res2: org.apache.spark.sql.SchemaRDD =  
SchemaRDD[6] at RDD at SchemaRDD.scala:108  
== Query Plan ==  
== Physical Plan ==  
PhysicalRDD [key#0], MappedRDD[5] at map at JsonRDD.scala:47

input.printSchema():  
root  
|-- key: string (nullable = true)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/bc6caa8f-b309-488c-8b1b-4cbef1e1c9fc%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/bc6caa8f-b309-488c-8b1b-4cbef1e1c9fc%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:31am UTC](https://discuss.elastic.co/t/spark-unable-to-index-json-from-hdfs-using-schemardd-savetoes/22275/2 "2017-07-06T00:31:26Z")

</div>


