# 4 Data nodes, struggling with simultaneous multiple heavy aggregations.. How to scale?

**URL:** <https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551>\
**Category:** Elasticsearch\
**Created:** [March 4, 2016, 8:39pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551 "2016-03-04T20:39:37Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![geebee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/geebee/32/8297_2.png) [@geebee](https://discuss.elastic.co/u/geebee)\
**Post date:** [March 4, 2016, 8:39pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/1 "2016-03-04T20:39:37Z")

</div>

Hello, hoping someone can give me some more insight on my issue:

I have an ES cluster with 4 data-only nodes (4 core/32GB RAM), under heavy aggregation scenarios (mostly large Kibana dashboards with multiple complex visualizations over longer time frames) the heap crosses 95% used, and node(s) crash.

I have the chance to scale by either adding a 5th identical data node, or by doubling the specs on the 4 existing nodes (to 4 core/64GB RAM each)

We use doc\_values extensively as well as pretty strict limits (by config) on fielddata cache size, and I don't believe it is a field data issue.

I'm not sure where else to look next, or which scaling strategy will be most effective for this use case and would greatly appreciate any advice on either that is available.

Thanks!

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 4, 2016, 9:09pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/2 "2016-03-04T21:09:35Z")

</div>

This will be a case of scaling horizontally more than anything.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 4, 2016, 9:13pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/3 "2016-03-04T21:13:00Z")

</div>

In this case it may make sense to scale up to 64GB of RAM before scaling out as it will give you more available heap space compared to adding an additional node.

---

<div class="post-metadata">

**Author:** ![geebee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/geebee/32/8297_2.png) [@geebee](https://discuss.elastic.co/u/geebee)\
**Post date:** [March 4, 2016, 9:15pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/4 "2016-03-04T21:15:27Z")

</div>

So you think that increasing the ability to distribute the queries across nodes by 20% won't alleviate heap pressure with the same level of effectiveness?

@warkolm just above you said basically the exact opposite, so I'm just trying to to get all my ducks in a row so to speak

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 4, 2016, 9:16pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/5 "2016-03-04T21:16:49Z")

</div>

> [@geebee](#):
>
> 4 core/32GB RAM

Is that heap or total?  
If it's the former then you can't scale vertically any more, if it's the latter then what Christian said can apply.

---

<div class="post-metadata">

**Author:** ![geebee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/geebee/32/8297_2.png) [@geebee](https://discuss.elastic.co/u/geebee)\
**Post date:** [March 4, 2016, 9:17pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/6 "2016-03-04T21:17:01Z")

</div>

Thanks! Can you explain a bit why you think adding the 5th node (20% greater distribution of queries) would be better than doubling the available resources on each existing node?

---

<div class="post-metadata">

**Author:** ![geebee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/geebee/32/8297_2.png) [@geebee](https://discuss.elastic.co/u/geebee)\
**Post date:** [March 4, 2016, 9:17pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/7 "2016-03-04T21:17:58Z")

</div>

Total. As per the docs, I'm running the nodes at ~16GB (50% of total system RAM)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 4, 2016, 9:21pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/8 "2016-03-04T21:21:21Z")

</div>

I misunderstood your original post.

---

<div class="post-metadata">

**Author:** ![geebee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/geebee/32/8297_2.png) [@geebee](https://discuss.elastic.co/u/geebee)\
**Post date:** [March 4, 2016, 9:30pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/9 "2016-03-04T21:30:36Z")

</div>

Thank you for taking the time to respond and clarify. Does that mean you agree with @Christian_Dahlqvist's assessment that scaling vertically in this case is likely to make most sense?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 4, 2016, 9:38pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/10 "2016-03-04T21:38:45Z")

</div>

Do your existing 4 CPU cores and disks have room? If so, and it makes business sense to just increate the memory, then yes.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:11pm UTC](https://discuss.elastic.co/t/4-data-nodes-struggling-with-simultaneous-multiple-heavy-aggregations-how-to-scale/43551/11 "2017-07-05T23:11:03Z")

</div>


