# Regarding Coordinator Node Heap Usage

**URL:** <https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080>\
**Category:** Elasticsearch\
**Created:** [October 11, 2018, 3:02pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080 "2018-10-11T15:02:08Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 11, 2018, 3:02pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/1 "2018-10-11T15:02:09Z")

</div>

Hi,  
I have a problem with coordinator node memory usage on one of our elasticsearch 5.6.3 clusters. We have 3 x coordinator nodes sitting in front of 6 data nodes. Normally, I size our coordinators with 8GB RAM and 5-6GB HEAP...  
However, in this case I've been seeing oom errors and have raised heap size to 13GB and then 26GB and heap usage is sitting at over 90% most of the time. Can someone help me understand what is causing this and to fix it?

Workload is as follows:

- client search requests from kibana (kibana can't get a response from the coordinators)
- bulk indexing from multiple fluentd indexers running in 4 x kubernetes clusters
- direct searches to the api (probably low)

We're ingesting about 600 million to 1 billion log lines per day with indices in the 250GB to 350GB range. The data nodes are under load and I'm planning to add more. But the coordinator node behaviour has me confused - they're normally pretty quiet. Any help would be greatly appreciated...

Regards,  
D

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 11, 2018, 8:36pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/2 "2018-10-11T20:36:11Z")

</div>

You should install X-Pack and enable Monitoring to see what is happening.  
Upgrading to 6.4.2 is also useful.

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 11, 2018, 8:50pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/3 "2018-10-11T20:50:43Z")

</div>

How does upgrading to 6.x from 5.x compare, relatively speaking, to upgrading to 5.x from 2.x?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 11, 2018, 8:51pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/4 "2018-10-11T20:51:17Z")

</div>

It's much simpler, just make sure you are on the latest 5.X release and you can do a rolling upgrade 🙂

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 11, 2018, 8:53pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/5 "2018-10-11T20:53:39Z")

</div>

Is it paid or free x-pack that will give the insight you're referring to?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 11, 2018, 8:54pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/6 "2018-10-11T20:54:02Z")

</div>

It's part of the Basic license, which is free.

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 12, 2018, 7:49am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/7 "2018-10-12T07:49:57Z")

</div>

Why does this page state that a full cluster restart is required?

[https://www.elastic.co/guide/en/elasticsearch/reference/6.0/breaking-changes.html](https://www.elastic.co/guide/en/elasticsearch/reference/6.0/breaking-changes.html)

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 17, 2018, 7:21am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/8 "2018-10-17T07:21:30Z")

</div>

@warkolm, did you miss my comment above? Just looking for clarification if you could...

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 17, 2018, 7:32am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/9 "2018-10-17T07:32:48Z")

</div>

[https://www.elastic.co/guide/en/elasticsearch/reference/6.0/rolling-upgrades.html](https://www.elastic.co/guide/en/elasticsearch/reference/6.0/rolling-upgrades.html) is the one you want, it mentions you can do rolling from 5.6.X . to 6.X.

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 18, 2018, 12:49pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/10 "2018-10-18T12:49:16Z")

</div>

Ok thx. How should the x-pack plugin be treated when upgrading to 6.4.2? Uninstall before, or after?

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [October 18, 2018, 2:08pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/11 "2018-10-18T14:08:38Z")

</div>

Uninstall before.

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 26, 2018, 7:35am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/12 "2018-10-26T07:35:33Z")

</div>

Hi @warkolm ,  
I'd like to return to this topic if we could. I have installed x-pack and am still seeing coordinator nodes dying periodically by running out of heap. Am attaching screenshots to provide more context:

 ![25](https://us1.discourse-cdn.com/elastic/original/3X/a/b/ab46f78641ea859bf8f4492f5122dc69ce591882.png)  
 ![09](https://us1.discourse-cdn.com/elastic/original/3X/1/f/1fff49a2c9e0a8e1d76fd862156ac2ee025de562.png)  
 ![32](https://us1.discourse-cdn.com/elastic/original/3X/4/d/4d78b7571473e523d89539eafa01833b46f84177.png)

Our workload is logging to daily indices and the flow rate you see in the screenshots is about normal. The ingest nodes, atm, work as coordinator nodes only. They've also been allocated larger heaps but that hasn't resolved the problem. It's not clear exactly what is exhausting the heap or what steps we should take to remedy it...

Regards,  
D

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 27, 2018, 12:35am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/13 "2018-10-27T00:35:41Z")

</div>

Can you just confirm what times the OOM occurs in the graphs?

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 29, 2018, 11:39am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/14 "2018-10-29T11:39:45Z")

</div>

Here's the last line in the log from the evening prior to the day I posted:

```auto
java.lang.OutOfMemoryError: Java heap space
[2018-10-25T21:23:57,321][INFO][o.e.m.j.JvmGcMonitorService] [ingest-001] [gc][old][138330][415] duration [18.1s], collections [3]/[9.8s], total [18.1s]/[14.3m], memory [13.6gb]->[13.6gb]/[13.6gb], all_pools {[young] [133.1mb]->[133.1mb]/[133.1mb]}{[survivor] [16.2mb]->[16.6mb]/[16.6mb]}{[old] [13.5gb]->[13.5gb]/[13.5gb]}

```

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 29, 2018, 11:43am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/15 "2018-10-29T11:43:51Z")

</div>

@warkolm Note that I see bulk rejections in the log prior to heap exhaustion. Bulk queue size is 400 on the data nodes. Does x-pack graph bulk queue size and rejections? I don't see that...

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 29, 2018, 8:02pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/16 "2018-10-29T20:02:41Z")

</div>

Oh, you're using a lot of Graph?

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 30, 2018, 8:31am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/17 "2018-10-30T08:31:35Z")

</div>

@warkolm Am I right in thinking that coordinator nodes hold inbound bulks on heap before relaying them to the data nodes? So, as data nodes fall behind the heap will fill?

If i enable an ingestion pipeline on these nodes will also mean the ingest nodes maintain their own bulk queue? It appears to me that while the data nodes are rejecting, the coordinators are not and that may be the cause of the issue here?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 30, 2018, 8:48am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/18 "2018-10-30T08:48:05Z")

</div>

> [@dawiro](#):
>
> Am I right in thinking that coordinator nodes hold inbound bulks on heap before relaying them to the data nodes? So, as data nodes fall behind the heap will fill?

Yes.

> [@dawiro](#):
>
> If i enable an ingestion pipeline on these nodes will also mean the ingest nodes maintain their own bulk queue?

I don't know enough of how that works to comment with any authority sorry ☹

---

<div class="post-metadata">

**Author:** ![dawiro](https://avatars.discourse-cdn.com/v4/letter/d/71e660/32.png) [@dawiro](https://discuss.elastic.co/u/dawiro)\
**Post date:** [October 30, 2018, 10:55am UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/19 "2018-10-30T10:55:10Z")

</div>

@warkolm Could you see if you can get an answer on that? What is the recommendation here? Should we be sending bulks direct to the data nodes?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 30, 2018, 8:56pm UTC](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080/20 "2018-10-30T20:56:59Z")

</div>

> [@dawiro](#):
>
> Should we be sending bulks direct to the data nodes?

We'd suggest that, yes. The other main recommendation would be to not send them to master-only nodes.

[Next page](https://discuss.elastic.co/t/regarding-coordinator-node-heap-usage/152080.md?page=2)
