# ES 1.3 - Calculating+updating a field in millions of documents

**URL:** <https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114>\
**Category:** Elasticsearch\
**Created:** [December 2, 2015, 3:16am UTC](https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114 "2015-12-02T03:16:32Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![jdyck](https://avatars.discourse-cdn.com/v4/letter/j/47e85d/32.png) [@jdyck](https://discuss.elastic.co/u/jdyck)\
**Post date:** [December 2, 2015, 3:16am UTC](https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114/1 "2015-12-02T03:16:32Z")

</div>

Hey all,

I have an ES index containing several million records. I'm adding a field to the document's mapping, afterwards I'll have to go back and calculate the value for this new field for every existing document. The value can be different for every document, so I don't think the 'update by query' idea pertains here since I'll have to calculate the value elsewhere.

Is there an efficient way of doing this?

Is the best way to do this to just pull the IDs of all documents in batches of 1000 or so, then for each batch calculate the value and update the documents using the bulk API?

Thanks!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 6:10am UTC](https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114/2 "2015-12-02T06:10:57Z")

</div>

If you are updating 100% off docs, it's better to reindex.

---

<div class="post-metadata">

**Author:** ![jdyck](https://avatars.discourse-cdn.com/v4/letter/j/47e85d/32.png) [@jdyck](https://discuss.elastic.co/u/jdyck)\
**Post date:** [December 2, 2015, 6:44am UTC](https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114/3 "2015-12-02T06:44:50Z")

</div>

Would you suggest something like this to do it with minimal downtime?

1. Create new index and PUT the revised mapping to the new index
2. Use the scan and scroll method to read all of the existing documents and calculate the value for the new field
3. Send the new documents to the new index in batches using the bulk api
4. Set an alias for the new index to 'flip the switch' to use the new index

In this case, would the disk usage double if I have the data in both indexes?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 7:27am UTC](https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114/4 "2015-12-02T07:27:51Z")

</div>

Exact for all points.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 7:28am UTC](https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114/5 "2015-12-02T07:28:23Z")

</div>

Just missing the DELETE old index at the end 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:34pm UTC](https://discuss.elastic.co/t/es-1-3-calculating-updating-a-field-in-millions-of-documents/36114/6 "2017-07-05T23:34:15Z")

</div>


