# How does heap shortage in the masters affect datanodes?

**URL:** <https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895>\
**Category:** Elasticsearch\
**Created:** [February 26, 2016, 9:26pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895 "2016-02-26T21:26:34Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 26, 2016, 9:26pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/1 "2016-02-26T21:26:34Z")

</div>

Hi,

Versions:

- Ubuntu 14.04
- ES 1.4.2 from elastic repos
- Java 1.7.0\_95

Cluster:

- 6 c3.8xlarge data nodes ( [http://www.ec2instances.info/?filter=c3.8xlarge](http://www.ec2instances.info/?filter=c3.8xlarge) ) with 30G of heap configured
- 3 dedicated masters (originally [http://www.ec2instances.info/?filter=r3.large](http://www.ec2instances.info/?filter=r3.large) ) with 7G of heap configured, all client requests are routed through these 3 servers

Index:

- Single index with 60 shards, 1 primary and 2 replicas, shards ranging from 10G to 40G
- 260Mil documents
- _not_ using doc values (looking into it)

The problem presented itself after attempting to replace the r3.large masters with m3.large ( [http://www.ec2instances.info/?filter=m3.large](http://www.ec2instances.info/?filter=m3.large) ) with 4G of heap configured. The 3 masters exhibited the following behavior:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/f/f63d697c776789484b16916e446106e1beab8276.png)

In essence: over 75% heap usage and increased amount of GC and CPU (due to GC probably)

At the same time, every datanode exhibited the following behavior:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/d/db00dfa0924b1963a5cba1542c506894388e58de.png)

i.e. the same behavior. During the time we had the smaller masters in charge of the cluster we saw it directly affecting the rest of the cluster. We replaced the masters with m3.xlarge instances ( [http://www.ec2instances.info/?filter=m3.xlarge](http://www.ec2instances.info/?filter=m3.xlarge) ) because we weren't unsure at the time if the problem was CPU or Heap. Now it looks to me as if Heap pressure was the problem, you will notice the datanodes in the last image heap usage is back to around 75% and old GC back to almost nothing.

A few questions then:

- Why the heap pressure in the masters affected the entire cluster? (in such a way that every datanode also showed heap pressure)
- In this particular use case, what would be the ideal master setup? My hunch says memory is the factor here (as far as I can tell the active master's process is single-threaded so it doesn't really matter how many cores the machine has?) and having 7G of heap size configured seems to have "fixed" this issue
- Is it particularly harmful for the active master to be serving requests? i.e. should I remove the active master from the configured URLs in the services using ES?
- Other general suggestions with regards to index/shard size appreciated

Thank you all.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 26, 2016, 10:22pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/2 "2016-02-26T22:22:47Z")

</div>

How big is your mapping for this monolithic index?

---

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 26, 2016, 10:54pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/3 "2016-02-26T22:54:23Z")

</div>

Hi @warkolm thanks for responding,

We have 9 pretty big mappings for that particular index. I would paste but:

> curl -XGET [http://localhost:9200/my\_index/\_mapping?pretty](http://localhost:9200/my_index/_mapping?pretty) \> /tmp/mappings.json  
> % Total % Received % Xferd Average Speed Time Time Time Current  
> Dload Upload Total Spent Left Speed  
> 100 85.5M 100 85.5M 0 0 15.1M 0 0:00:05 0:00:05 --:--:-- 20.3M

So not easy to share. I can tell you though that the mappings have in average 75 properties each.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 26, 2016, 11:18pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/4 "2016-02-26T23:18:24Z")

</div>

That's quite large, especially if there's 9 of those! ES needs to store that in cluster state, which is held in memory. It does compress is but that's still pretty big.

You're on a pretty old version, can you upgrade?

---

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 27, 2016, 12:41am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/5 "2016-02-27T00:41:47Z")

</div>

Looking into the upgrade path to possibly 1.7.5, which will help I am sure.  
With regards to my original question, are you suggesting the heap needed by the master is directly tied to the number and size of the mappings?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 27, 2016, 12:42am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/6 "2016-02-27T00:42:46Z")

</div>

It's likely, what was heap use like before the change? Or do you not have metrics from that time.

---

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 27, 2016, 12:51am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/7 "2016-02-27T00:51:39Z")

</div>

I do, actually, here's how the old r3.large masters looked like before we attempted to replace them with the smaller instances:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/2/297a2ae672442f7299d246f75df583094bde040f.png)

The graphing is kind of messed up because the sampling size is 1mo. You can tell the heap usage on the active master (green) is steady at 75% and the other 2 much lower.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 27, 2016, 4:05am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/8 "2016-02-27T04:05:40Z")

</div>

> [@thekad](#):
>
> 3 dedicated masters (originally [Amazon EC2 Instance Comparison](http://www.ec2instances.info/?filter=r3.large) ) with 7G of heap configured, all client requests are routed through these 3 servers

If you are sending all requests through these nodes, they are not truly dedicated master nodes, but rather master/client nodes. It is recommended that dedicated master nodes do not serve traffic, but are allowed to just manage the cluster. This will prevent them from suffering from memory pressure intense garbage collection and make the cluster more stable.

If you instead set up dedicated client nodes, you may be able to reduce the size of the dedicated master nodes as they will be doing a lot less work.

---

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 27, 2016, 5:23am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/9 "2016-02-27T05:23:30Z")

</div>

Interesting. My initial thought was that resource usage in the masters was so low they could be used for routing, specially the passive masters since they are doing close to nothing. If your suggestion is that I can have routing (client) nodes then what would be the minimum requirements to run these?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 28, 2016, 3:05am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/10 "2016-02-28T03:05:04Z")

</div>

Truly dedicated master nodes can often be considerable smaller than the ones you used and have heap as a larger portion of the available memory (~75%) than data nodes as they do not need as much file system cache as nodes that hold data. Instances suitable here may be t2.medium or m3.medium unless the cluster state is very large.

The same recommendation around heap size also applies to client nodes, as they also do not store data. In addition to acting as routers, these nodes also coordinate requests and perform the final aggregation steps once data from the underlying shards have been gathered, which is why they can suffer from garbage collection depending on the work load. These are more difficult to size, but a good starting point may be the instances types you previously used together with the higher heap setting.

---

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 29, 2016, 4:39am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/11 "2016-02-29T04:39:55Z")

</div>

Ah, here we go we're converging back to size of cluster state, think that's my key problem. I believe I have 2 final questions:

1. Is there a formula to size the JVM according to the cluster state? i.e. can I grab the mappings and some how get a number of Gb I absolutely need for my heap size?
2. Is the active master process single-threaded? So far I _think so_ but I haven't dug through the code to make sure.

Many thanks.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 29, 2016, 4:45am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/12 "2016-02-29T04:45:17Z")

</div>

1. Not really, cluster state is kept in binary form in memory, so the json output is simply a human readable version.
2. What do you mean by master process exactly?

---

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 29, 2016, 4:48am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/13 "2016-02-29T04:48:37Z")

</div>

In my case I have 3 masters, but only one is active at any given time. This active master shows cpu usage on a single core all the time, that to me means that the ES master process is always bound to a single CPU, but I am not 100% positive.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 29, 2016, 5:06am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/14 "2016-02-29T05:06:21Z")

</div>

Check `hot_threads` against that node to see what is happening.

---

<div class="post-metadata">

**Author:** ![thekad](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thekad/32/8068_2.png) [@thekad](https://discuss.elastic.co/u/thekad)\
**Post date:** [February 29, 2016, 9:34pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/15 "2016-02-29T21:34:56Z")

</div>

sorry, I didn't mean to say 1 core is always pegged, I meant to say _when_ there is activity in the active master it's always a) brief and b) only on a single core. Kinda hard to grab a hot\_threads dump when it spikes so briefly.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 29, 2016, 11:40pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/16 "2016-02-29T23:40:20Z")

</div>

Most of the work a dedicated master node does is probably linked to maintaining the cluster state, which could very well be single threaded in nature. It is generally not very CPU intensive, which is why I recommended using instances with relatively little CPU.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 1, 2016, 1:16am UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/17 "2016-03-01T01:16:08Z")

</div>

Cluster state is single threaded, to ensure consistency.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:12pm UTC](https://discuss.elastic.co/t/how-does-heap-shortage-in-the-masters-affect-datanodes/42895/18 "2017-07-05T23:12:43Z")

</div>


