# Suggestion for updating documents?

**URL:** <https://discuss.elastic.co/t/suggestion-for-updating-documents/15640>\
**Category:** Elasticsearch\
**Created:** [February 6, 2014, 11:25am UTC](https://discuss.elastic.co/t/suggestion-for-updating-documents/15640 "2014-02-06T11:25:54Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ivan\_Ji](https://avatars.discourse-cdn.com/v4/letter/i/d78d45/32.png) [@Ivan\_Ji](https://discuss.elastic.co/u/Ivan_Ji)\
**Post date:** [February 6, 2014, 11:25am UTC](https://discuss.elastic.co/t/suggestion-for-updating-documents/15640/1 "2014-02-06T11:25:54Z")

</div>

Hi all,

Assume I already had lot of documents inside ES and each document represent  
one file.

But now I want to update some files' fields, so I need to find the  
document, get its id, and then apply the \_update operation.

And If I have n document to do such things and there are m document inside  
the ES, the performance to search the desired document to get its id is  
O(n\*m), right? Because each finding operation needs to scan entire  
documents inside the index, does it exist any way to find the desired  
document with unique key, return immediately when found it?

If so, it's really not a good option when to update a document's field. I  
am wondering what's the suggested workflow to update some file without  
knowing its id first.

Ideas?

Cheers,

Ivan

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/8c91ab7e-1ae5-4f97-abad-68afec17aa76%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/8c91ab7e-1ae5-4f97-abad-68afec17aa76%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [February 6, 2014, 12:14pm UTC](https://discuss.elastic.co/t/suggestion-for-updating-documents/15640/2 "2014-02-06T12:14:22Z")

</div>

How do you find the documents you need to update? I guess by executing a  
query? In that case, the search won't scan all the documents, this is the  
whole point of elasticsearch and lucene. There is an inverted index which  
makes it easy to find matches based on the terms in your queries.

Still, updating a lot of documents can be quite expensive, but the problem  
is not exactly the query part (aka finding the documents to update) but the  
update itself, as it need to get each document back, delete it and reindex  
it internally (that's how updates work in lucene). This is why the update  
by query feature has not been exposed yet.

On Thursday, February 6, 2014 12:25:54 PM UTC+1, Ivan Ji wrote:

> Hi all,
> 
> Assume I already had lot of documents inside ES and each document  
> represent one file.
> 
> But now I want to update some files' fields, so I need to find the  
> document, get its id, and then apply the \_update operation.
> 
> And If I have n document to do such things and there are m document inside  
> the ES, the performance to search the desired document to get its id is  
> O(n\*m), right? Because each finding operation needs to scan entire  
> documents inside the index, does it exist any way to find the desired  
> document with unique key, return immediately when found it?
> 
> If so, it's really not a good option when to update a document's field. I  
> am wondering what's the suggested workflow to update some file without  
> knowing its id first.
> 
> Ideas?
> 
> Cheers,
> 
> Ivan

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a777758c-1c2b-4030-9bb4-16f28ee5b0d0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a777758c-1c2b-4030-9bb4-16f28ee5b0d0%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan\_Ji](https://avatars.discourse-cdn.com/v4/letter/i/d78d45/32.png) [@Ivan\_Ji](https://discuss.elastic.co/u/Ivan_Ji)\
**Post date:** [February 6, 2014, 2:09pm UTC](https://discuss.elastic.co/t/suggestion-for-updating-documents/15640/3 "2014-02-06T14:09:09Z")

</div>

Hi Luca,

Thanks for replies. In fact, I am not familiar with the internal algorithm  
of Elasticsearch.  
Seems I need to catch up these knowledge, such as lucene. Thanks a lot for  
clearance of this question.

Ivan

Luca Cavanna於 2014年2月6日星期四UTC+8下午8時14分22秒寫道：

> How do you find the documents you need to update? I guess by executing a  
> query? In that case, the search won't scan all the documents, this is the  
> whole point of elasticsearch and lucene. There is an inverted index which  
> makes it easy to find matches based on the terms in your queries.
> 
> Still, updating a lot of documents can be quite expensive, but the problem  
> is not exactly the query part (aka finding the documents to update) but the  
> update itself, as it need to get each document back, delete it and reindex  
> it internally (that's how updates work in lucene). This is why the update  
> by query feature has not been exposed yet.
> 
> On Thursday, February 6, 2014 12:25:54 PM UTC+1, Ivan Ji wrote:
> 
> > Hi all,
> > 
> > Assume I already had lot of documents inside ES and each document  
> > represent one file.
> > 
> > But now I want to update some files' fields, so I need to find the  
> > document, get its id, and then apply the \_update operation.
> > 
> > And If I have n document to do such things and there are m document  
> > inside the ES, the performance to search the desired document to get its id  
> > is O(n\*m), right? Because each finding operation needs to scan entire  
> > documents inside the index, does it exist any way to find the desired  
> > document with unique key, return immediately when found it?
> > 
> > If so, it's really not a good option when to update a document's field. I  
> > am wondering what's the suggested workflow to update some file without  
> > knowing its id first.
> > 
> > Ideas?
> > 
> > Cheers,
> > 
> > Ivan

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/11a76615-9f8b-4914-8623-30e050dfcce1%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/11a76615-9f8b-4914-8623-30e050dfcce1%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:51am UTC](https://discuss.elastic.co/t/suggestion-for-updating-documents/15640/4 "2017-07-06T01:51:58Z")

</div>


