# How to upsert an initial value into elasticsearch using spark?

**URL:** <https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [September 17, 2015, 1:02am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450 "2015-09-17T01:02:02Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 17, 2015, 1:02am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/1 "2015-09-17T01:02:02Z")

</div>

With HTTP POST, the following script can insert a new field `createtime` or update `lastupdatetime`:

```
curl -XPOST 'localhost:9200/test/type1/1/_update' -d '{
"doc": {
    "lastupdatetime": "2015-09-16T18:00:00"
}
"upsert" : {
    "createtime": "2015-09-16T18:00:00"
    "lastupdatetime": "2015-09-16T18:00:00",
}
}'

```

But in spark script, after setting `"es.write.operation": "upsert"`, i don't know how to insert `createtime` at all. There is only `es.update.script.*` in the [official document](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/configuration.html#cfg-update)... So, can anyone give me an example?

---

<div class="post-metadata">

**Author:** ![eliasah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eliasah/32/34741_2.png) [@eliasah](https://discuss.elastic.co/u/eliasah)\
**Post date:** [September 19, 2015, 10:11am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/2 "2015-09-19T10:11:03Z")

</div>

I'm not very sure about that. But what I'll try to do is defining as `es.mapping.id` with the key of the document you want to upsert in the SparkConf.

```
val conf = new SparkConf()
[...]
conf.set("es.write.operation", "upsert")
// you can set the the document field/property name containing the document id.
// I believe that you are able to know that you should change <id> 
// with the desired field name
conf.set("es.mapping.id",<id>) 

```

I haven't tried this but I think that it should work!

Let us know if it works for you!

---

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 20, 2015, 6:33am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/3 "2015-09-20T06:33:40Z")

</div>

Thank you for your help.

In my case, i want to save the information of Android devices from log into one elasticsearch type, and set it's first appearance time as `createtime`.

If the device appear again, only update the `lastupdatetime`, but leave the `createtime` as it was. So the document id is android ID, if the id exists, update `lastupdatetime`, else insert `createtime` and `lastupdatetime`. So the setting here is(in python):

```auto
    conf = {
        "es.resource.write": "stats-device/activation",
        "es.nodes": "NODE1:9200",
        "es.write.operation": "upsert",
        "es.mapping.id": "id"
        # ???
    }
 
    rdd.saveAsNewAPIHadoopFile(
        path='-',
        outputFormatClass="org.elasticsearch.hadoop.mr.EsOutputFormat",
        keyClass="org.apache.hadoop.io.NullWritable",
        valueClass="org.elasticsearch.hadoop.mr.LinkedMapWritable",
        conf=conf
    )

```

I don't know how to `insert` a new field if the `id` not exist...

---

<div class="post-metadata">

**Author:** ![eliasah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eliasah/32/34741_2.png) [@eliasah](https://discuss.elastic.co/u/eliasah)\
**Post date:** [September 20, 2015, 8:02am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/4 "2015-09-20T08:02:46Z")

</div>

The id doesn't exist where? In Elasticsearch or it the data you are reading?

---

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 20, 2015, 8:44am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/5 "2015-09-20T08:44:17Z")

</div>

Well, id is just in the input data.

---

<div class="post-metadata">

**Author:** ![eliasah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eliasah/32/34741_2.png) [@eliasah](https://discuss.elastic.co/u/eliasah)\
**Post date:** [September 20, 2015, 9:02am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/6 "2015-09-20T09:02:24Z")

</div>

Then theoretically speaking it should only be available in your data when you perform an upsert

---

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 20, 2015, 11:15am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/7 "2015-09-20T11:15:21Z")

</div>

So you means i can't upsert any new field not in my data when i use spark? OK, i got it.... Thank you.

---

<div class="post-metadata">

**Author:** ![eliasah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eliasah/32/34741_2.png) [@eliasah](https://discuss.elastic.co/u/eliasah)\
**Post date:** [September 20, 2015, 1:38pm UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/8 "2015-09-20T13:38:54Z")

</div>

I didn't say that. Let's agree first of the definition of an upsert action which is the following :

> If the document does not already exist, the contents of the upsert element will be inserted as a new document. If the document does exist, then the script will be executed instead.

Which in your case will be :

- update on id since you defined es.mapping.id -\> id
- insert a document with the \_id equals id

---

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 21, 2015, 3:11am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/9 "2015-09-21T03:11:41Z")

</div>

Sorry, i couldn't login this site last day.

You are right. The requirement of my case is just:

1. Update `lastupdatetime` if the `id` already exist in elasticsearch;
2. Insert `lastupdatetime` and `createtime`=`lastupdatetime` when `id` not exist in elasticsearch;

The source doc is just like this:

```auto
{
    'id': 'xxxxx',
    'lastupdatetime': '2015-09-20'
}

```

The problem is that a new field `createtime` will be added to the source doc if `id` not exist in es yet. I don't know how to solve this problem in spark.

Here is a solution which is not perfect:

1. add `createtime` to all source doc (`rdd.map(lambda d: d['createtime']=d['lastupdatetime']`)
2. save to es with `create` and ignore 409
3. remove `createtime` field
4. save to es again with `update`

After these steps, i get what i want. But if there is a better solution?

---

<div class="post-metadata">

**Author:** ![eliasah](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eliasah/32/34741_2.png) [@eliasah](https://discuss.elastic.co/u/eliasah)\
**Post date:** [September 21, 2015, 3:10pm UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/10 "2015-09-21T15:10:30Z")

</div>

What is the structure of your final rdd before writing it to es?

---

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 22, 2015, 1:51am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/11 "2015-09-22T01:51:09Z")

</div>

Just like my last post, i write the doc to es twice now.  
At first time, using `create`, rdd structure is

```auto
{
    'id': 'xxxxx',
    'lastupdatetime': '2015-09-20',
    'createtime':'2015-09-20',
}

```

At second time, using `update`, rdd structure is

```auto
{
    'id': 'xxxxx',
    'lastupdatetime': '2015-09-20'
}

```

so, if the `id` already exist, `create` will be fail, only `lastupdatetime` will be updated. However, i believe these 2 operations can be combined into 1.

---

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 24, 2015, 10:43am UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/12 "2015-09-24T10:43:46Z")

</div>

At last, i found `elasticsearch-hadoop` don't support **create if not exist, else do nothing**.

---

<div class="post-metadata">

**Author:** ![Terran\_Yiu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/terran_yiu/32/4358_2.png) [@Terran\_Yiu](https://discuss.elastic.co/u/Terran_Yiu)\
**Post date:** [September 24, 2015, 3:28pm UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/13 "2015-09-24T15:28:32Z")

</div>

from [this post](https://github.com/elastic/elasticsearch-hadoop/pull/308), i decide to build the elasticsearch-hadoop jar myself.

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [September 26, 2015, 2:35pm UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/14 "2015-09-26T14:35:55Z")

</div>

Sorry for the late reply.  
It seems the update support is not complete. I've seen you already raised an issue (great!) [here](https://github.com/elastic/elasticsearch-hadoop/issues/553) so Iet's use that to track progress.

Enhancing the doc only to do another update is far from the proper solution. And doing bulk requests which create exceptions which later on are ignored is even worse.  
This should be properly fixed in one call not 3 plus exceptions.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:27pm UTC](https://discuss.elastic.co/t/how-to-upsert-an-initial-value-into-elasticsearch-using-spark/29450/15 "2017-07-06T13:27:28Z")

</div>


