# How index distributed into ElasticSearch cluster

**URL:** <https://discuss.elastic.co/t/how-index-distributed-into-elasticsearch-cluster/10565>\
**Category:** Elasticsearch\
**Created:** [January 31, 2013, 6:20am UTC](https://discuss.elastic.co/t/how-index-distributed-into-elasticsearch-cluster/10565 "2013-01-31T06:20:41Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![dean\_2](https://avatars.discourse-cdn.com/v4/letter/d/ce7236/32.png) [@dean\_2](https://discuss.elastic.co/u/dean_2)\
**Post date:** [January 31, 2013, 6:20am UTC](https://discuss.elastic.co/t/how-index-distributed-into-elasticsearch-cluster/10565/1 "2013-01-31T06:20:41Z")

</div>

There are 3 nodes of my ES cluster.  
The shard number is 5.  
There are 2 shards existing on 2 of the nodes and the last shard existing  
on the last node.  
I use TransportClient.addTransportAddress(...) to add all these 3 nodes.  
Then a BulkRequestBuilder is created using the TransportClient.  
Then index request is added to BulkRequestBuilder.  
At last bulkRequest.execute().actionGet() is used to send the ES cluster.

The things I want to know is:

1. If I just use one of the nodes to communicate with the ES cluster, are  
the indices be distributed to all the ES cluster? What's the Java class  
used to do this?

2. How indices are distributed into the ES cluster? Is it based on the  
nodes or based on the shard?  
For example, if there are 3000 documents to be indexed; Then 1000  
documents for each node? Or 600 documents for each shard?  
What's the java class name about this policy?

3. Is there a good tool to manage ES cluster? Like deploying, monitoring,  
upgrading, restarting, or installing new plugin.  
From the website, Chef is used as an example to manage ES cluster; but  
Chef is not a standard scm in my company.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![radu\_gheorghe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe/32/556_2.png) [@radu\_gheorghe](https://discuss.elastic.co/u/radu_gheorghe)\
**Post date:** [January 31, 2013, 1:08pm UTC](https://discuss.elastic.co/t/how-index-distributed-into-elasticsearch-cluster/10565/2 "2013-01-31T13:08:37Z")

</div>

Hello,

On Thu, Jan 31, 2013 at 8:20 AM, Dean Zhang from China \<  
[elasticbetter@gmail.com](mailto:elasticbetter@gmail.com)\> wrote:

> There are 3 nodes of my ES cluster.  
> The shard number is 5.  
> There are 2 shards existing on 2 of the nodes and the last shard existing  
> on the last node.  
> I use TransportClient.addTransportAddress(...) to add all these 3 nodes.  
> Then a BulkRequestBuilder is created using the TransportClient.  
> Then index request is added to BulkRequestBuilder.  
> At last bulkRequest.execute().actionGet() is used to send the ES cluster.
> 
> The things I want to know is:
> 
> 1. If I just use one of the nodes to communicate with the ES cluster, are  
> the indices be distributed to all the ES cluster?

No, by default it will automatically distribute your shards so that you get  
an even number of shards per node. That's regardless of indexes, of how  
many docs there are per shard, and of the performance of your nodes. In  
your case with 5 shards and 3 nodes, 1 node will have to end up with just  
one shard.

But you can change the way your shards are allocated:

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

And you can also move shards around manually:

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

In future, there will be other algorithms available for distributing  
shards. This one looks really nice:  
[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/cluster/routing/allocation/allocator/BalancedShardsAllocator.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/cluster/routing/allocation/allocator/BalancedShardsAllocator.java)

> What's the Java class used to do this?

I believe it's this one:  
[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/cluster/routing/allocation/allocator/EvenShardsCountAllocator.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/cluster/routing/allocation/allocator/EvenShardsCountAllocator.java)

> 1. How indices are distributed into the ES cluster? Is it based on the  
> nodes or based on the shard?  
> For example, if there are 3000 documents to be indexed; Then 1000  
> documents for each node? Or 600 documents for each shard?

By default, documents are distributed pretty evenly per shard - so ~600  
docs/shard in your example with 5 shards. But you can add an explicit  
routing value to control that. For more details and some other links, take  
a look here:

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

> ```
> What's the java class name about this policy?
> 
> ```

I think it's this one:  
[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/cluster/routing/operation/hash/simple/SimpleHashFunction.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/cluster/routing/operation/hash/simple/SimpleHashFunction.java)

> 1. Is there a good tool to manage ES cluster? Like deploying, monitoring,  
> upgrading, restarting, or installing new plugin.  
> From the website, Chef is used as an example to manage ES cluster; but  
> Chef is not a standard scm in my company.

For monitoring, I'd recommend our own SPM:

> **[Elasticsearch - Sematext Documentation](https://sematext.com/docs/integration/elasticsearch-integration/)**
>
> Collect and monitor key Elasticsearch metrics such as request latency, indexing rate, and segment merges with built-in anomaly detection, threshold, and heartbeat alerts. Send notifications to email and various chatops messaging services, correlate...

For managing, what is a standard scm in your company?

## Best regards, Radu

[http://sematext.com/](http://sematext.com/) -- Elasticsearch -- Solr -- Lucene

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:53am UTC](https://discuss.elastic.co/t/how-index-distributed-into-elasticsearch-cluster/10565/3 "2017-07-06T02:53:46Z")

</div>


