# Pushing Data to Elasticsearch from Spark

**URL:** <https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [January 6, 2017, 11:39pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827 "2017-01-06T23:39:12Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![boots](https://avatars.discourse-cdn.com/v4/letter/b/c89c15/32.png) [@boots](https://discuss.elastic.co/u/boots)\
**Post date:** [January 6, 2017, 11:39pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/1 "2017-01-06T23:39:12Z")

</div>

Hello All,

I am running into an issue when trying to use the Spark Scala library to push data to an index within our Elasticsearch cluster. I have a DStream[(String, String)] with the first String being the '\_id' I want the document to have and the value (the second String) being a JSONified version of the document.

I try to save the DStream using this line:

EsSparkStreaming.saveToEsWithMeta(transformed, "my-index/tmp", Map(ES\_INPUT\_JSON -\> true.toString))

When I do this I get the following error:

17/01/06 16:26:19 ERROR JobScheduler: Error running job streaming job 1483745150000 ms.0  
org.apache.spark.SparkException: Job aborted due to stage failure: Task 1 in stage 0.0 failed 1 times, most recent failure: Lost task 1.0 in stage 0.0 (TID 1, localhost): org.elasticsearch.hadoop.rest.EsHadoopInvalidRequest: Unexpected character ('d' (code 100)) in numeric value: Exponent indicator not followed by a digit  
at [Source: [B@33014046; line: 1, column: 23]  
at org.elasticsearch.hadoop.rest.RestClient.checkResponse(RestClient.java:488)  
at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:446)  
at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:436)  
at org.elasticsearch.hadoop.rest.RestClient.bulk(RestClient.java:185)  
at org.elasticsearch.hadoop.rest.RestRepository.tryFlush(RestRepository.java:220)  
at org.elasticsearch.hadoop.rest.RestRepository.flush(RestRepository.java:242)  
at org.elasticsearch.hadoop.rest.RestRepository.doWriteToIndex(RestRepository.java:182)  
at org.elasticsearch.hadoop.rest.RestRepository.writeToIndex(RestRepository.java:159)  
at org.elasticsearch.spark.rdd.EsRDDWriter.write(EsRDDWriter.scala:67)  
at org.elasticsearch.spark.rdd.EsSpark$$anonfun$doSaveToEs$1.apply(EsSpark.scala:102)  
at org.elasticsearch.spark.rdd.EsSpark$$anonfun$doSaveToEs$1.apply(EsSpark.scala:102)  
at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)  
at org.apache.spark.scheduler.Task.run(Task.scala:86)  
at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)  
at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)  
at java.lang.Thread.run(Thread.java:745)

This incorrect parsing results in a cascading set of failures all related to JSON parse errors.

After firing up Wireshark I was able to capture the data sent to the cluster and found this:

{"index":{"\_id":4315ede3:dff5:3127:9d15:dc692d825cc8}}  
{JSON DOCUMENT}

It looks like the library either isn't encoding the \_id correctly (not passing as a JSON string) or I have a misunderstanding and the \_id field is meant to be numeric.

Any help in troubleshooting this issue would be greatly appreciated!

---

<div class="post-metadata">

**Author:** ![eperry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eperry/32/551_2.png) [@eperry](https://discuss.elastic.co/u/eperry)\
**Post date:** [January 7, 2017, 6:10am UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/2 "2017-01-07T06:10:51Z")

</div>

well the ID is not quoated so in JSON terms it would have to be an integer. but an ID should be quoted

[https://www.elastic.co/guide/en/elasticsearch/reference/current/mapping-id-field.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/mapping-id-field.html)

beyond that is the limit of my knowledge of what your doing and from what I see

---

<div class="post-metadata">

**Author:** ![boots](https://avatars.discourse-cdn.com/v4/letter/b/c89c15/32.png) [@boots](https://discuss.elastic.co/u/boots)\
**Post date:** [January 9, 2017, 4:17pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/3 "2017-01-09T16:17:05Z")

</div>

Yeah I figured that the JSON was being improperly encoded as a string but I am not sure why this is happening.

---

<div class="post-metadata">

**Author:** ![eperry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eperry/32/551_2.png) [@eperry](https://discuss.elastic.co/u/eperry)\
**Post date:** [January 10, 2017, 12:56pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/4 "2017-01-10T12:56:28Z")

</div>

you probably want to go open an issue on Github for the library, or maybe someone else will have an idea.

---

<div class="post-metadata">

**Author:** ![boots](https://avatars.discourse-cdn.com/v4/letter/b/c89c15/32.png) [@boots](https://discuss.elastic.co/u/boots)\
**Post date:** [January 10, 2017, 4:18pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/5 "2017-01-10T16:18:27Z")

</div>

That came to my mind yesterday after I found a workaround since it seems like a bug.

For anyone that has this same issue just make quotation marks (") as part of the string itself.

E.G  
Java  
Sting s = ""THIS IS MY ID"";

Scala  
val id = s""""This is my id""""

---

<div class="post-metadata">

**Author:** ![boots](https://avatars.discourse-cdn.com/v4/letter/b/c89c15/32.png) [@boots](https://discuss.elastic.co/u/boots)\
**Post date:** [January 10, 2017, 7:24pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/6 "2017-01-10T19:24:34Z")

</div>

Bug report has been created.

> <https://github.com/elastic/elasticsearch-hadoop/issues/913>

---

<div class="post-metadata">

**Author:** ![eperry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eperry/32/551_2.png) [@eperry](https://discuss.elastic.co/u/eperry)\
**Post date:** [January 11, 2017, 2:09pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/7 "2017-01-11T14:09:24Z")

</div>

Well glad you found A work around and opened a bug report on it. It always helps everyone.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 8, 2017, 2:09pm UTC](https://discuss.elastic.co/t/pushing-data-to-elasticsearch-from-spark/70827/8 "2017-02-08T14:09:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
