# Upsert nested fields with Spark

**URL:** <https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [April 20, 2018, 7:36am UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813 "2018-04-20T07:36:39Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![cluengo](https://avatars.discourse-cdn.com/v4/letter/c/d9b06d/32.png) [@cluengo](https://discuss.elastic.co/u/cluengo)\
**Post date:** [April 20, 2018, 7:36am UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/1 "2018-04-20T07:36:39Z")

</div>

I'm trying to update a nested field using Spark and a scripted update. The code I'm using is:

update\_params = "new\_samples: samples"  
update\_script = "ctx.\_source.samples += new\_samples"

es\_conf = {  
"es.mapping.id": "id",  
"es.mapping.exclude": "id",  
"es.write.operation": "upsert",  
"es.update.script.params": update\_params,  
"es.update.script.inline": update\_script  
}

result.write.format("org.elasticsearch.spark.sql").options(\*\*es\_conf).option("es.nodes",configuration["elasticsearch"]["host"]).option("es.port",configuration["elasticsearch"]["port"] ).save(configuration["elasticsearch"]["index\_name"]+"/"+configuration["version"],mode='append')

And the schema of the field is:

|-- samples: array (nullable = true)  
| |-- element: struct (containsNull = true)  
| | |-- gq: integer (nullable = true)  
| | |-- dp: integer (nullable = true)  
| | |-- gt: string (nullable = true)  
| | |-- adBug: array (nullable = true)  
| | | |-- element: integer (containsNull = true)  
| | |-- ad: double (nullable = true)  
| | |-- sample: string (nullable = true)

And I get the following error:

py4j.protocol.Py4JJavaError: An error occurred while calling o83.save.  
: org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 1.0 failed 1 times, most recent failure: Lost task 0.0 in stage 1.0 (TID 1, localhost, executor driver): java.lang.ClassCastException: scala.collection.mutable.WrappedArray$ofRef cannot be cast to scala.Tuple2

Am I doing something wrong or is this a bug? Thanks.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [April 20, 2018, 4:42pm UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/2 "2018-04-20T16:42:31Z")

</div>

It's difficult to say without the full stack trace. Can you provide it?

---

<div class="post-metadata">

**Author:** ![cluengo](https://avatars.discourse-cdn.com/v4/letter/c/d9b06d/32.png) [@cluengo](https://discuss.elastic.co/u/cluengo)\
**Post date:** [April 21, 2018, 2:47pm UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/3 "2018-04-21T14:47:54Z")

</div>

The full stack trace is:

[https://pastebin.com/34tFHqYj](https://pastebin.com/34tFHqYj)

I couldn't paste it here because it hit the maximum amount of characters.

Thanks again!

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [April 24, 2018, 5:35pm UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/4 "2018-04-24T17:35:40Z")

</div>

Ahah, yes, this is a known issue: [https://github.com/elastic/elasticsearch-hadoop/issues/931](https://github.com/elastic/elasticsearch-hadoop/issues/931)

We have a set of interfaces that define how to convert between integration specific record types and JSON values. For SparkSQL, each value in a row needs to be placed in a specific spot in the row order. Since JSON doesn't necessarily make any guarantees of the order of fields, we need the schema available to go between Row objects and raw values. When we serialize the data in SparkSQL, we pass in a Tuple2 with the schema in one slot and the record in the other. This somewhat breaks the contract that we have built for these serialization tools in other places though, most notably here where we try to extract a value to be used in a script parameter.

---

<div class="post-metadata">

**Author:** ![cluengo](https://avatars.discourse-cdn.com/v4/letter/c/d9b06d/32.png) [@cluengo](https://discuss.elastic.co/u/cluengo)\
**Post date:** [April 25, 2018, 8:30am UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/5 "2018-04-25T08:30:30Z")

</div>

Ok, thanks for the fast reply James! Is there a workaround for the moment, or will it be fixed in future releases?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [April 26, 2018, 6:51pm UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/6 "2018-04-26T18:51:32Z")

</div>

This will be fixed in a future release. I do not believe there is any workaround for it at the moment aside from trying to keep all script parameters limited to basic data types.

---

<div class="post-metadata">

**Author:** ![cluengo](https://avatars.discourse-cdn.com/v4/letter/c/d9b06d/32.png) [@cluengo](https://discuss.elastic.co/u/cluengo)\
**Post date:** [April 26, 2018, 7:27pm UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/7 "2018-04-26T19:27:51Z")

</div>

Ok, thanks!!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 24, 2018, 7:27pm UTC](https://discuss.elastic.co/t/upsert-nested-fields-with-spark/128813/8 "2018-05-24T19:27:55Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
