# Data balancing and backup in the cluster

**URL:** <https://discuss.elastic.co/t/data-balancing-and-backup-in-the-cluster/3781>\
**Category:** Elasticsearch\
**Created:** [January 15, 2011, 8:35pm UTC](https://discuss.elastic.co/t/data-balancing-and-backup-in-the-cluster/3781 "2011-01-15T20:35:17Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Barak\_Yaish](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@Barak\_Yaish](https://discuss.elastic.co/u/Barak_Yaish)\
**Post date:** [January 15, 2011, 8:35pm UTC](https://discuss.elastic.co/t/data-balancing-and-backup-in-the-cluster/3781/1 "2011-01-15T20:35:17Z")

</div>

Hello,

Doing my first steps in ES, I've few questions:

1. In case the cluster is composed of N nodes, is data split equally  
on the nodes?
2. In case node is crashed, is its data backuped on the other nodes  
so no data loss in case of such crash? And if the client uses this  
node for queries, will it still get answers, or be notified that error  
occure?
3. In case master node is crashed, is the cluster still functioning?
4. In case a new node joins the cluster, how much time takes the  
cluster to re-balance the data (say 1B docs, 4 nodes cluster)?

Thanks

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 16, 2011, 9:40am UTC](https://discuss.elastic.co/t/data-balancing-and-backup-in-the-cluster/3781/2 "2011-01-16T09:40:59Z")

</div>

On Saturday, January 15, 2011 at 10:35 PM, barak wrote:

> Hello,
> 
> Doing my first steps in ES, I've few questions:
> 
> 1. In case the cluster is composed of N nodes, is data split equally  
> on the nodes?

The aim of the cluster is to get an even number of shards allocated on each node.

> 1. In case node is crashed, is its data backuped on the other nodes  
> so no data loss in case of such crash? And if the client uses this  
> node for queries, will it still get answers, or be notified that error  
> occure?

Each shard can have one or more replicas. If a node crashes, the replicas will consist of its backup, so no data is lost, and the shards allocated on that node will get reallocated on the rest of the nodes.

If a client uses that node to query, and that node crashes, then you need to use another node to query. If you use HTTP with the REST API, then you can simply round robin between servers.

> 1. In case master node is crashed, is the cluster still functioning?

Yes, another node will be elected as master.

> 1. In case a new node joins the cluster, how much time takes the  
> cluster to re-balance the data (say 1B docs, 4 nodes cluster)?

Depends on your network. There is no reindexing being done, just moving data around (shards).

> Thanks

---

<div class="post-metadata">

**Author:** ![Barak\_Yaish](https://avatars.discourse-cdn.com/v4/letter/b/9de053/32.png) [@Barak\_Yaish](https://discuss.elastic.co/u/Barak_Yaish)\
**Post date:** [January 16, 2011, 2:53pm UTC](https://discuss.elastic.co/t/data-balancing-and-backup-in-the-cluster/3781/3 "2011-01-16T14:53:36Z")

</div>

On Jan 16, 11:40 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> If a client uses that node to query, and that node crashes, then you need to use another node to query. If you use HTTP with the REST API, then you can simply round robin between servers.

Is this means that java clients ( TransportClient ) cannot be cached?  
In case of adding or removing nodes to the cluster, the client must  
recreated to recognize the change?

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [January 16, 2011, 3:47pm UTC](https://discuss.elastic.co/t/data-balancing-and-backup-in-the-cluster/3781/4 "2011-01-16T15:47:24Z")

</div>

Hi,

TransportClient can learn about changes in the cluster, no need to recreate  
it.

Regards,  
Lukas  
Dne 16.1.2011 15:53 "barak" [barak.yaish@gmail.com](mailto:barak.yaish@gmail.com) napsal(a):

> On Jan 16, 11:40 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > If a client uses that node to query, and that node crashes, then you need  
> > to use another node to query. If you use HTTP with the REST API, then you  
> > can simply round robin between servers.
> 
> Is this means that java clients ( TransportClient ) cannot be cached?  
> In case of adding or removing nodes to the cluster, the client must  
> recreated to recognize the change?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:13am UTC](https://discuss.elastic.co/t/data-balancing-and-backup-in-the-cluster/3781/5 "2017-07-06T04:13:48Z")

</div>


