# Requirements per node role

**URL:** <https://discuss.elastic.co/t/requirements-per-node-role/6295>\
**Category:** Elasticsearch\
**Created:** [January 6, 2012, 2:49am UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295 "2012-01-06T02:49:48Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![plaflamme](https://avatars.discourse-cdn.com/v4/letter/p/a88e57/32.png) [@plaflamme](https://discuss.elastic.co/u/plaflamme)\
**Post date:** [January 6, 2012, 2:49am UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/1 "2012-01-06T02:49:48Z")

</div>

Hi,

A node in a cluster can be configured to serve exclusively as: a data,  
master or "client" node. By deciding a node's role a-priori, I suspect one  
should tweak its hardware (more RAM, less CPU, etc.) to fit. What would be  
the (relative) recommendations per role?

- data nodes (not master eligible, not serving requests):

- master nodes (no data, not serving requests):

- seems it's not doing much... lazy master nodes!

- client nodes (not master eligible, no data):

If there are several master-only nodes, are they all idle except for one at  
any given time?

Anyone have experience in deploying a cluster with "load-balancing" clients  
for serving requests?

Thanks,  
Philippe

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [January 6, 2012, 12:19pm UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/2 "2012-01-06T12:19:39Z")

</div>

ES does not have the concept master vs. slave.

Have a look:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Peter.

On 6 Jan., 03:49, Philippe Laflamme [philippe.lafla...@obiba.org](mailto:philippe.lafla...@obiba.org)  
wrote:

> Hi,
> 
> A node in a cluster can be configured to serve exclusively as: a data,  
> master or "client" node. By deciding a node's role a-priori, I suspect one  
> should tweak its hardware (more RAM, less CPU, etc.) to fit. What would be  
> the (relative) recommendations per role?
> 
> - data nodes (not master eligible, not serving requests):
> 
> - master nodes (no data, not serving requests):
> 
> - seems it's not doing much... lazy master nodes!
> 
> - client nodes (not master eligible, no data):
> 
> If there are several master-only nodes, are they all idle except for one at  
> any given time?
> 
> Anyone have experience in deploying a cluster with "load-balancing" clients  
> for serving requests?
> 
> Thanks,  
> Philippe

---

<div class="post-metadata">

**Author:** ![plaflamme](https://avatars.discourse-cdn.com/v4/letter/p/a88e57/32.png) [@plaflamme](https://discuss.elastic.co/u/plaflamme)\
**Post date:** [January 6, 2012, 2:29pm UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/3 "2012-01-06T14:29:38Z")

</div>

Yes, I'm aware that all nodes are equivalent by default (and they elect a  
master node themselves), but by changing the default settings, you can make  
a node not master eligible (node.master=false or node.client=true) and you  
can decide whether a node has data (node.data).

Using node.data=false and node.client=true, you're effectively creating a  
node that will not serve indices directly, but will redirect requests to  
"data nodes" and aggregate the results. I'm wondering if there are any  
advantages in creating such nodes. For example, does this change the  
requirements on hardware (requires less RAM, no disk access, etc.) If so,  
one can create a cluster topology and scale data nodes and client nodes  
independently.

Maybe this only introduces additional complexity, but for cloud-based  
solutions such as EC2, it may be interesting to have different types of  
nodes to have greater flexibility for choosing instance types.

Thanks,  
Philippe

On Fri, Jan 6, 2012 at 07:19, Karussell [tableyourtime@googlemail.com](mailto:tableyourtime@googlemail.com)wrote:

> ES does not have the concept master vs. slave.
> 
> Have a look:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/videos/2010/02/08/es-distributed-diagram.html)
> 
> Peter.
> 
> On 6 Jan., 03:49, Philippe Laflamme [philippe.lafla...@obiba.org](mailto:philippe.lafla...@obiba.org)  
> wrote:
> 
> > Hi,
> > 
> > A node in a cluster can be configured to serve exclusively as: a data,  
> > master or "client" node. By deciding a node's role a-priori, I suspect  
> > one  
> > should tweak its hardware (more RAM, less CPU, etc.) to fit. What would  
> > be  
> > the (relative) recommendations per role?
> > 
> > - data nodes (not master eligible, not serving requests):
> > 
> > - master nodes (no data, not serving requests):
> > 
> > - seems it's not doing much... lazy master nodes!
> > 
> > - client nodes (not master eligible, no data):
> > 
> > If there are several master-only nodes, are they all idle except for one  
> > at  
> > any given time?
> > 
> > Anyone have experience in deploying a cluster with "load-balancing"  
> > clients  
> > for serving requests?
> > 
> > Thanks,  
> > Philippe

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [January 7, 2012, 1:56pm UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/4 "2012-01-07T13:56:00Z")

</div>

On 6 Jan., 15:29, Philippe Laflamme [philippe.lafla...@obiba.org](mailto:philippe.lafla...@obiba.org)  
wrote:

> Yes, I'm aware that all nodes are equivalent by default (and they elect a  
> master node themselves), but by changing the default settings, you can make  
> a node not master eligible (node.master=false or node.client=true) and you  
> can decide whether a node has data (node.data).
> 
> Using node.data=false and node.client=true, you're effectively creating a  
> node that will not serve indices directly, but will redirect requests to  
> "data nodes" and aggregate the results.

why do you think that adding a separate no-data node would be  
beneficial? what should be the advantages overs directing the queries  
directly to a data node? As the data node needs to process the query  
nevertheless.

Peter.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 7, 2012, 7:48pm UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/5 "2012-01-07T19:48:46Z")

</div>

Having just "load balancing nodes" (non master, non data) will not help  
that much. I have seen cases where it was used to run it locally with the  
relevant client code for HTTP access since it was connecting over loopback  
and ES had better network handling for remote access to nodes. "Just" data  
nodes still serve requests, even if they are coming from client nodes.

Dedicated master nodes can become handy in certain situations. For very  
large clusters they can help (i.e. 200 data nodes with 3 "eligible" master  
nodes).

On Sat, Jan 7, 2012 at 3:56 PM, Karussell [tableyourtime@googlemail.com](mailto:tableyourtime@googlemail.com)wrote:

> On 6 Jan., 15:29, Philippe Laflamme [philippe.lafla...@obiba.org](mailto:philippe.lafla...@obiba.org)  
> wrote:
> 
> > Yes, I'm aware that all nodes are equivalent by default (and they elect a  
> > master node themselves), but by changing the default settings, you can  
> > make  
> > a node not master eligible (node.master=false or node.client=true) and  
> > you  
> > can decide whether a node has data (node.data).
> > 
> > Using node.data=false and node.client=true, you're effectively creating a  
> > node that will not serve indices directly, but will redirect requests to  
> > "data nodes" and aggregate the results.
> 
> why do you think that adding a separate no-data node would be  
> beneficial? what should be the advantages overs directing the queries  
> directly to a data node? As the data node needs to process the query  
> nevertheless.
> 
> Peter.

---

<div class="post-metadata">

**Author:** ![plaflamme](https://avatars.discourse-cdn.com/v4/letter/p/a88e57/32.png) [@plaflamme](https://discuss.elastic.co/u/plaflamme)\
**Post date:** [January 7, 2012, 11:30pm UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/6 "2012-01-07T23:30:51Z")

</div>

> why do you think that adding a separate no-data node would be  
> beneficial? what should be the advantages overs directing the queries  
> directly to a data node? As the data node needs to process the query  
> nevertheless.

I was wondering if putting several client-only nodes "in front" of the data  
nodes would be beneficial by offloading some work from data nodes. These  
nodes would act like load-balancing nodes and maybe would require different  
type of hardware (less RAM, more CPU, no disk access, for example).

The data nodes still have to process the query, but they wouldn't have to  
aggregate the results.

Thanks,  
Philippe

---

<div class="post-metadata">

**Author:** ![plaflamme](https://avatars.discourse-cdn.com/v4/letter/p/a88e57/32.png) [@plaflamme](https://discuss.elastic.co/u/plaflamme)\
**Post date:** [January 7, 2012, 11:33pm UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/7 "2012-01-07T23:33:33Z")

</div>

Ok, thanks for the answers.

In your example, how does having 3 "master eligible" nodes help in a 200  
node cluster? Is this for avoiding split brain situations?

Thanks,  
Philippe

On Sat, Jan 7, 2012 at 14:48, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Having just "load balancing nodes" (non master, non data) will not help  
> that much. I have seen cases where it was used to run it locally with the  
> relevant client code for HTTP access since it was connecting over loopback  
> and ES had better network handling for remote access to nodes. "Just" data  
> nodes still serve requests, even if they are coming from client nodes.
> 
> Dedicated master nodes can become handy in certain situations. For very  
> large clusters they can help (i.e. 200 data nodes with 3 "eligible" master  
> nodes).
> 
> On Sat, Jan 7, 2012 at 3:56 PM, Karussell [tableyourtime@googlemail.com](mailto:tableyourtime@googlemail.com)wrote:
> 
> > On 6 Jan., 15:29, Philippe Laflamme [philippe.lafla...@obiba.org](mailto:philippe.lafla...@obiba.org)  
> > wrote:
> > 
> > > Yes, I'm aware that all nodes are equivalent by default (and they elect  
> > > a  
> > > master node themselves), but by changing the default settings, you can  
> > > make  
> > > a node not master eligible (node.master=false or node.client=true) and  
> > > you  
> > > can decide whether a node has data (node.data).
> > > 
> > > Using node.data=false and node.client=true, you're effectively creating  
> > > a  
> > > node that will not serve indices directly, but will redirect requests to  
> > > "data nodes" and aggregate the results.
> > 
> > why do you think that adding a separate no-data node would be  
> > beneficial? what should be the advantages overs directing the queries  
> > directly to a data node? As the data node needs to process the query  
> > nevertheless.
> > 
> > Peter.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 12, 2012, 10:33am UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/8 "2012-01-12T10:33:00Z")

</div>

Yes, it mainly helps in avoiding split brain situations (as it can only  
happen between those 3 nodes).

On Sun, Jan 8, 2012 at 1:33 AM, Philippe Laflamme \<  
[philippe.laflamme@obiba.org](mailto:philippe.laflamme@obiba.org)\> wrote:

> Ok, thanks for the answers.
> 
> In your example, how does having 3 "master eligible" nodes help in a 200  
> node cluster? Is this for avoiding split brain situations?
> 
> Thanks,  
> Philippe
> 
> On Sat, Jan 7, 2012 at 14:48, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Having just "load balancing nodes" (non master, non data) will not help  
> > that much. I have seen cases where it was used to run it locally with the  
> > relevant client code for HTTP access since it was connecting over loopback  
> > and ES had better network handling for remote access to nodes. "Just" data  
> > nodes still serve requests, even if they are coming from client nodes.
> > 
> > Dedicated master nodes can become handy in certain situations. For very  
> > large clusters they can help (i.e. 200 data nodes with 3 "eligible" master  
> > nodes).
> > 
> > On Sat, Jan 7, 2012 at 3:56 PM, Karussell [tableyourtime@googlemail.com](mailto:tableyourtime@googlemail.com)wrote:
> > 
> > > On 6 Jan., 15:29, Philippe Laflamme [philippe.lafla...@obiba.org](mailto:philippe.lafla...@obiba.org)  
> > > wrote:
> > > 
> > > > Yes, I'm aware that all nodes are equivalent by default (and they  
> > > > elect a  
> > > > master node themselves), but by changing the default settings, you can  
> > > > make  
> > > > a node not master eligible (node.master=false or node.client=true) and  
> > > > you  
> > > > can decide whether a node has data (node.data).
> > > > 
> > > > Using node.data=false and node.client=true, you're effectively  
> > > > creating a  
> > > > node that will not serve indices directly, but will redirect requests  
> > > > to  
> > > > "data nodes" and aggregate the results.
> > > 
> > > why do you think that adding a separate no-data node would be  
> > > beneficial? what should be the advantages overs directing the queries  
> > > directly to a data node? As the data node needs to process the query  
> > > nevertheless.
> > > 
> > > Peter.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:42am UTC](https://discuss.elastic.co/t/requirements-per-node-role/6295/9 "2017-07-06T03:42:57Z")

</div>


