# Spark RDD.saveToES

**URL:** <https://discuss.elastic.co/t/spark-rdd-savetoes/26527>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [July 29, 2015, 9:08pm UTC](https://discuss.elastic.co/t/spark-rdd-savetoes/26527 "2015-07-29T21:08:14Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![pferrel](https://avatars.discourse-cdn.com/v4/letter/p/59ef9b/32.png) [@pferrel](https://discuss.elastic.co/u/pferrel)\
**Post date:** [July 29, 2015, 9:08pm UTC](https://discuss.elastic.co/t/spark-rdd-savetoes/26527/1 "2015-07-29T21:08:14Z")

</div>

The Spark writing of an index works well if you construct the entire dataset with all fields before you write using rdd.saveToES. Is there a way to use this mechanism for upserting to change an existing field? I want to change the value of an existing field without changing the rest of the document.

If I write an rdd whose Map elements contain only one field won't the entire doc be deleted except for the Map element?

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [August 13, 2015, 7:41pm UTC](https://discuss.elastic.co/t/spark-rdd-savetoes/26527/2 "2015-08-13T19:41:04Z")

</div>

Depends on how you define the update operation; you can specify a script which can only change the value as oppose to deleting the whole document.  
Along with `mapping.include`/`exclude`, the [configuration](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/configuration.html#cfg-update) settings give you access to all the update options in Elastic.

---

<div class="post-metadata">

**Author:** ![pferrel](https://avatars.discourse-cdn.com/v4/letter/p/59ef9b/32.png) [@pferrel](https://discuss.elastic.co/u/pferrel)\
**Post date:** [August 14, 2015, 3:40pm UTC](https://discuss.elastic.co/t/spark-rdd-savetoes/26527/3 "2015-08-14T15:40:20Z")

</div>

Thanks this is good to know but not sure these mappings help. First I don't know anything about the structure of the document at the time I am trying to do the equivalent of upserting a double value into the doc properties.

This seems like a very simple use case where I'm adding a possibly new property to a doc but rdd.saveToEs overwrites the entire doc with the Map in each rdd element.

The include/excluse docs seem to be talking about pruning unneeded data from a doc so maybe I misunderstand things.

To be clear one element of the rdd is something like a Scala tuple `("doc1", Map(("popularity" -> 1.0d))`. I know doc1 has other fields but only want to write the "popularity" double field. If I use `include` mapping for "popularity won't this just erase the rest of the doc?

Should I `include` all with \* and give it the Map above? Will that leave all fields alone and overwrite the "popularity" field?

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [August 18, 2015, 9:39am UTC](https://discuss.elastic.co/t/spark-rdd-savetoes/26527/4 "2015-08-18T09:39:23Z")

</div>

I think you misunderstand how `update` works in Elasticsearch. ES-Spark doesn't change its semantics rather exposes them in a way that's convenient in Spark.  
Take a look at the Elasticsearch documentation - start for example with [this section](https://www.elastic.co/guide/en/elasticsearch/guide/current/partial-updates.html) in the reference guide on partial updates which is what you are looking for.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:27pm UTC](https://discuss.elastic.co/t/spark-rdd-savetoes/26527/5 "2017-07-06T13:27:40Z")

</div>


