# Update

**URL:** <https://discuss.elastic.co/t/update/3559>\
**Category:** Elasticsearch\
**Created:** [November 11, 2010, 6:04pm UTC](https://discuss.elastic.co/t/update/3559 "2010-11-11T18:04:27Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![mooky](https://avatars.discourse-cdn.com/v4/letter/m/43a26b/32.png) [@mooky](https://discuss.elastic.co/u/mooky)\
**Post date:** [November 11, 2010, 6:04pm UTC](https://discuss.elastic.co/t/update/3559/1 "2010-11-11T18:04:27Z")

</div>

Any thoughts on supporting an "update" feature in ES?

We have a need to update a quantity of documents - and rather than  
rebuild the ES document and re-index, we'd rather read the data from  
ES (since it has all the data), modify & reindex. (will be much  
snappier, as re-assembling the ES document is a bit costly for us) .

What would be one step better is calling update with a query/filter &  
pass, say, a js function to do our update and have ES execute it & re-  
index all under the hood. That way we avoid having to write a bunch of  
scrolling through large results sets, batching indexing operations and  
we avoid shifting all the data back to the client and then back to the  
ES node(s).

Thoughts?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 12, 2010, 8:13am UTC](https://discuss.elastic.co/t/update/3559/2 "2010-11-12T08:13:16Z")

</div>

Hi,

Yes, that certainly make sense. The difficulty of handling this revolves  
around the distributed nature (more specifically, replication) of  
operations. There are different ways to implement it:

1. Have the update js function run on the primary shard, and then batch  
changes using a similar mechanism to the batch API. This will means  
replicating the _data_ to the replicas. This is simpler to implement, though  
the update is not atomic or blocking (other index operations on the same  
data might "get in").

2. Have the update function happen on the primary and the replicas. This is  
more efficient when it comes to not needing to transfer the data to the  
replicas, but the query will be executed on all replicas, and its _much_  
harder to maintain consistency of shard and its replicas in this case (this  
must be maintained of course).

aparo has been talking about it as well (on IRC), and even went ahead and  
implemented a proof of concept code.

-shay.banon

On Thu, Nov 11, 2010 at 8:04 PM, Mooky [nick.minutello@gmail.com](mailto:nick.minutello@gmail.com) wrote:

> Any thoughts on supporting an "update" feature in ES?
> 
> We have a need to update a quantity of documents - and rather than  
> rebuild the ES document and re-index, we'd rather read the data from  
> ES (since it has all the data), modify & reindex. (will be much  
> snappier, as re-assembling the ES document is a bit costly for us) .
> 
> What would be one step better is calling update with a query/filter &  
> pass, say, a js function to do our update and have ES execute it & re-  
> index all under the hood. That way we avoid having to write a bunch of  
> scrolling through large results sets, batching indexing operations and  
> we avoid shifting all the data back to the client and then back to the  
> ES node(s).
> 
> Thoughts?

---

<div class="post-metadata">

**Author:** ![mooky](https://avatars.discourse-cdn.com/v4/letter/m/43a26b/32.png) [@mooky](https://discuss.elastic.co/u/mooky)\
**Post date:** [November 12, 2010, 6:36pm UTC](https://discuss.elastic.co/t/update/3559/3 "2010-11-12T18:36:01Z")

</div>

Cool

On 12 November 2010 08:13, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com) wrote:

> Hi,
> 
> Yes, that certainly make sense. The difficulty of handling this revolves  
> around the distributed nature (more specifically, replication) of  
> operations. There are different ways to implement it:
> 
> 1. Have the update js function run on the primary shard, and then batch  
> changes using a similar mechanism to the batch API. This will means  
> replicating the _data_ to the replicas. This is simpler to implement, though  
> the update is not atomic or blocking (other index operations on the same  
> data might "get in").
> 
> 2. Have the update function happen on the primary and the replicas. This is  
> more efficient when it comes to not needing to transfer the data to the  
> replicas, but the query will be executed on all replicas, and its _much_  
> harder to maintain consistency of shard and its replicas in this case (this  
> must be maintained of course).
> 
> aparo has been talking about it as well (on IRC), and even went ahead and  
> implemented a proof of concept code.
> 
> -shay.banon
> 
> On Thu, Nov 11, 2010 at 8:04 PM, Mooky [nick.minutello@gmail.com](mailto:nick.minutello@gmail.com) wrote:
> 
> > Any thoughts on supporting an "update" feature in ES?
> > 
> > We have a need to update a quantity of documents - and rather than  
> > rebuild the ES document and re-index, we'd rather read the data from  
> > ES (since it has all the data), modify & reindex. (will be much  
> > snappier, as re-assembling the ES document is a bit costly for us) .
> > 
> > What would be one step better is calling update with a query/filter &  
> > pass, say, a js function to do our update and have ES execute it & re-  
> > index all under the hood. That way we avoid having to write a bunch of  
> > scrolling through large results sets, batching indexing operations and  
> > we avoid shifting all the data back to the client and then back to the  
> > ES node(s).
> > 
> > Thoughts?

---

<div class="post-metadata">

**Author:** ![Alberto\_Paro\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alberto_paro_2/32/1137_2.png) [@Alberto\_Paro\_2](https://discuss.elastic.co/u/Alberto_Paro_2)\
**Post date:** [November 16, 2010, 8:15am UTC](https://discuss.elastic.co/t/update/3559/4 "2010-11-16T08:15:59Z")

</div>

On 12/nov/2010, at 19.36, Nick Minutello wrote:

> Cool
> 
> On 12 November 2010 08:13, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com) wrote:  
> Hi,
> 
> Yes, that certainly make sense. The difficulty of handling this revolves around the distributed nature (more specifically, replication) of operations. There are different ways to implement it:
> 
> 1. Have the update js function run on the primary shard, and then batch changes using a similar mechanism to the batch API. This will means replicating the _data_ to the replicas. This is simpler to implement, though the update is not atomic or blocking (other index operations on the same data might "get in").
> 
> 2. Have the update function happen on the primary and the replicas. This is more efficient when it comes to not needing to transfer the data to the replicas, but the query will be executed on all replicas, and its _much_ harder to maintain consistency of shard and its replicas in this case (this must be maintained of course).
> 
> aparo has been talking about it as well (on IRC), and even went ahead and implemented a proof of concept code.
> 
> -shay.banon  
> I've implemented an update of ES for a my client at point 2. The problem, as Shay said, is the consistency of shard and its replicas.  
> There is a high risk to have broken/corrupted data in your index if there are problem during update (Verified in some my border cases).  
> So now I'm working to implement parent/children approuch as Shay suggested to me.

Hi,  
Alberto Paro

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 16, 2010, 8:31am UTC](https://discuss.elastic.co/t/update/3559/5 "2010-11-16T08:31:43Z")

</div>

Hi Alberto,

I am going to work on parent child post 0.13... :), no promises though...

On Tue, Nov 16, 2010 at 10:15 AM, Alberto Paro [alberto.paro@gmail.com](mailto:alberto.paro@gmail.com)wrote:

> On 12/nov/2010, at 19.36, Nick Minutello wrote:
> 
> Cool
> 
> On 12 November 2010 08:13, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> 
> > Hi,
> > 
> > Yes, that certainly make sense. The difficulty of handling this revolves  
> > around the distributed nature (more specifically, replication) of  
> > operations. There are different ways to implement it:
> > 
> > 1. Have the update js function run on the primary shard, and then batch  
> > changes using a similar mechanism to the batch API. This will means  
> > replicating the _data_ to the replicas. This is simpler to implement, though  
> > the update is not atomic or blocking (other index operations on the same  
> > data might "get in").
> > 
> > 2. Have the update function happen on the primary and the replicas. This  
> > is more efficient when it comes to not needing to transfer the data to the  
> > replicas, but the query will be executed on all replicas, and its _much_  
> > harder to maintain consistency of shard and its replicas in this case (this  
> > must be maintained of course).
> > 
> > aparo has been talking about it as well (on IRC), and even went ahead and  
> > implemented a proof of concept code.
> > 
> > -shay.banon
> 
> I've implemented an update of ES for a my client at point 2. The problem,  
> as Shay said, is the consistency of shard and its replicas.  
> There is a high risk to have broken/corrupted data in your index if there  
> are problem during update (Verified in some my border cases).  
> So now I'm working to implement parent/children approuch as Shay suggested  
> to me.
> 
> Hi,  
> Alberto Paro

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:16am UTC](https://discuss.elastic.co/t/update/3559/6 "2017-07-06T04:16:26Z")

</div>


