# Moving big indices around with two or more instances running in the same box, working separately

**URL:** <https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689>\
**Category:** Elasticsearch\
**Created:** [January 20, 2016, 6:43pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689 "2016-01-20T18:43:12Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![fforbeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fforbeck/32/6137_2.png) [@fforbeck](https://discuss.elastic.co/u/fforbeck)\
**Post date:** [January 20, 2016, 6:43pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/1 "2016-01-20T18:43:12Z")

</div>

Hi there,

I would like an opinion about this scenario:

- 4 ES instances, one node each, running in the same box but using different cluster names.  
So they won't be working together. But they all will store same type of data.
- A proxy running on top of these instances to route my indices to the right node.
- ZSF snapshot as backup

I am not complete aware of the ES capabilities and the main reason here would be easily move an entire index to other new servers, as it grows.

How is it sound for you?

Thanks

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 20, 2016, 9:13pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/2 "2016-01-20T21:13:56Z")

</div>

Why not just cluster the new node to the old one, then use [shard filtering](https://www.elastic.co/guide/en/elasticsearch/reference/2.1/allocation-filtering.html) to shift the data across.

---

<div class="post-metadata">

**Author:** ![fforbeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fforbeck/32/6137_2.png) [@fforbeck](https://discuss.elastic.co/u/fforbeck)\
**Post date:** [January 20, 2016, 9:36pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/3 "2016-01-20T21:36:17Z")

</div>

I did not know about this, but one question.  
With this approach, I would not have to declare my indices in the file?  
Because my indices are generated on the fly, so do not know their names before their creation.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 20, 2016, 9:38pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/4 "2016-01-20T21:38:35Z")

</div>

> [@fforbeck](#):
>
> With this approach, I would not have to declare my indices in the file?

In what file?

---

<div class="post-metadata">

**Author:** ![fforbeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fforbeck/32/6137_2.png) [@fforbeck](https://discuss.elastic.co/u/fforbeck)\
**Post date:** [January 20, 2016, 11:53pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/5 "2016-01-20T23:53:24Z")

</div>

Sorry, I mean, not necessarily in the conf file.  
I would have to send a PUT for each new index I have created on the fly.

for instance, new index 1000  
PUT 1000/\_settings  
{  
"index.routing.allocation.include.\_name": "my-node-A"  
}  
...

Then, sending my bulk request and repeat this process for each new index passing the correct node name.

Currently I have something around 6K indices and it will get much bigger.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 21, 2016, 1:11am UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/6 "2016-01-21T01:11:49Z")

</div>

I don't understand why you'd want to move stuff around, nor why you have this stand alone "cluster" setup.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 21, 2016, 6:59am UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/7 "2016-01-21T06:59:32Z")

</div>

I must admit that I also do not understand exactly what you are trying to achieve. The fact that you say that you have 6000 indices in a cluster that size and expect that to grow is however a concern. Each shard in Elasticsearch is a Lucene index and has some resource overhead (memory, file handles, CPU) associated with it. Having that many indices and shards will unnecessarily use up a lot of system resources and is likely to not scale well. Why are you having so many indices?

---

<div class="post-metadata">

**Author:** ![fforbeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fforbeck/32/6137_2.png) [@fforbeck](https://discuss.elastic.co/u/fforbeck)\
**Post date:** [January 21, 2016, 12:49pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/8 "2016-01-21T12:49:42Z")

</div>

Here is my case:

- One client can archive many websites.
- Each archived website has a unique id.
- For each time the same client archives the same website, I have a different snapshot id.
- The document content is the website _html/css/js/etc_
- The documents are stored following this path:  
**/website\_id/snapshot\_id/document\_id**
- Each new index is dedicated to the snapshots from a particular website.
- Each website have its own index.

That's why so many indices.

Now, about the stand alone instances. The idea was proposed in order to facilitate the data migration to another server as it grows. For instance, move one entire index data to a another dedicated server. As the data was stored in only one node we could easily move the entire index.

I am trying to understand if this would be viable. That is why I would like to hear from you guys and know if there is other ways achieve that. 🙂

Thanks a lot

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 22, 2016, 3:28am UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/9 "2016-01-22T03:28:03Z")

</div>

Why not have a cluster, then when you need, add larger nodes in and use filtering to move data off the smaller nodes.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 22, 2016, 8:18am UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/10 "2016-01-22T08:18:55Z")

</div>

Having a separate index per website wastes a lot of resources and will not scale. If the structure of the documents you are indexing for the different website are similar and/or you can control the mappings, I would recommend storing multiple, if not all, websites in a single index. You can then either add filters at the application layer or use [filtered aliases](https://www.elastic.co/guide/en/elasticsearch/guide/current/faking-it.html) when accessing the data. If you always query per website, you can also use [routing](https://www.elastic.co/blog/customizing-your-document-routing) in order to ensure all documents belonging to a single website reside in a single shard, which can improve query latency and throughput.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:22pm UTC](https://discuss.elastic.co/t/moving-big-indices-around-with-two-or-more-instances-running-in-the-same-box-working-separately/39689/11 "2017-07-05T23:22:19Z")

</div>


