# Read all documents in an index and write to a new index (with more shards)

**URL:** <https://discuss.elastic.co/t/read-all-documents-in-an-index-and-write-to-a-new-index-with-more-shards/27426>\
**Category:** Elasticsearch\
**Created:** [August 14, 2015, 7:08pm UTC](https://discuss.elastic.co/t/read-all-documents-in-an-index-and-write-to-a-new-index-with-more-shards/27426 "2015-08-14T19:08:57Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![burtonator](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/burtonator/32/44790_2.png) [@burtonator](https://discuss.elastic.co/u/burtonator)\
**Post date:** [August 14, 2015, 7:08pm UTC](https://discuss.elastic.co/t/read-all-documents-in-an-index-and-write-to-a-new-index-with-more-shards/27426/1 "2015-08-14T19:08:57Z")

</div>

Are there any tools that read all documents in an index, and write them to a new index?

This is useful if you have an index with say 50 shards, and you have 50 machines, but you want to expand this to 100 machines. This way you would just create a new index, read all documents from the old index, then write the new index, then swap the index names.

There are some problems with some naive approaches:

- how do you make it work in parallel so it works on all nodes at once?

- how do you make it resume so that if it crashes it can pick up where it left off.

I mean I think I can dive in and write this, but would be nice if something worked out of the box.

---

<div class="post-metadata">

**Author:** ![mosiddi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mosiddi/32/577_2.png) [@mosiddi](https://discuss.elastic.co/u/mosiddi)\
**Post date:** [August 15, 2015, 6:12am UTC](https://discuss.elastic.co/t/read-all-documents-in-an-index-and-write-to-a-new-index-with-more-shards/27426/2 "2015-08-15T06:12:07Z")

</div>

this is a good thread. I'm not a big fan of having 1 shard per machine. When closing on index/shard model, one should think of reasonable # of indices and shards per index so we have isolation and no single point of failure. And more importantly space for scaling in future.

Coming to this Q - this will be possible only when u have \_source ON OR you are storing all fields, otherwise you wont be able to retrieve all docs.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 15, 2015, 6:59am UTC](https://discuss.elastic.co/t/read-all-documents-in-an-index-and-write-to-a-new-index-with-more-shards/27426/3 "2015-08-15T06:59:40Z")

</div>

There are a few tools out there and you can DIY.

I use Logstash - [https://gist.github.com/markwalkom/8a7201e3f6ea4354ae06](https://gist.github.com/markwalkom/8a7201e3f6ea4354ae06)

---

<div class="post-metadata">

**Author:** ![burtonator](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/burtonator/32/44790_2.png) [@burtonator](https://discuss.elastic.co/u/burtonator)\
**Post date:** [September 9, 2015, 9:25pm UTC](https://discuss.elastic.co/t/read-all-documents-in-an-index-and-write-to-a-new-index-with-more-shards/27426/4 "2015-09-09T21:25:40Z")

</div>

Yeah. Right now you need \_source but we're doing that anyway. Not sure what % of people are doing this though.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:51pm UTC](https://discuss.elastic.co/t/read-all-documents-in-an-index-and-write-to-a-new-index-with-more-shards/27426/5 "2017-07-05T23:51:28Z")

</div>


