# Troubleshooting High heap usage

**URL:** https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554
**Category:** Elasticsearch
**Created:** [February 6, 2018, 4:03am UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554 "2018-02-06T04:03:00Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![animageofmine](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@animageofmine](https://discuss.elastic.co/u/animageofmine)
#### Post date: [February 6, 2018, 4:03am UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554/1 "2018-02-06T04:03:00Z")

</div>

I have looked at many posts in the discussion forum and googled quite a bit as well. Either these posts are for older versions or there wasn't anything that matched the issue I am seeing. Hence, posting it one more time. Apologies if there is something really obvious that I am overlooking.

There are a 5 data nodes in our cluster and data seem to be distributed evenly, but only one or two data nodes consistently seem to have high heap usage (between 85-90%). Can someone help me out with troubleshooting steps? Few things I have looked at:

1. Cluster health / unassigned shards : 0
2. Number of pending tasks: 0
3. Script stats: look normal
4. Field data: between 10-15% of memory

Data Node configuration: 16 cores, 30 GB RAM, 15 GB for lucene and ES each  
ES version: 5.3.2

I couldn't interpret anything meaningful from other stats. Please see node stats [in my onedrive](https://1drv.ms/u/s!Av3dfIYtwY0T4RGFk9xH3A7lyv-A)

We don't have X-Pack and installing one in production is out of scope as of now. Let me know if you need more info.

---

<div class="post-metadata">

### Author: ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)
#### Post date: [February 6, 2018, 6:55am UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554/2 "2018-02-06T06:55:13Z")

</div>

What is the full output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/5.3/cluster-stats.html)?

---

<div class="post-metadata">

### Author: ![animageofmine](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@animageofmine](https://discuss.elastic.co/u/animageofmine)
#### Post date: [February 6, 2018, 8:23am UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554/3 "2018-02-06T08:23:02Z")

</div>

I had to reboot the node. Please find cluster stats output [here](https://1drv.ms/u/s!Av3dfIYtwY0T4RNuhirN8oHl2CGt)

We are working on reducing number of shards, but that does not seem to be an issue since other nodes are holding up fine (we have a cluster with over 25k shards and that is also doing alright).

---

<div class="post-metadata">

### Author: ![loren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loren/32/44942_2.png) [@loren](https://discuss.elastic.co/u/loren)
#### Post date: [February 6, 2018, 6:37pm UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554/4 "2018-02-06T18:37:40Z")

</div>

Are you doing a lot of updates to existing documents? I see a lot of deleted docs reported in your cluster stats.

I ran into a [similar issue](https://discuss.elastic.co/t/es6-1-heap-memory-used-by-the-index-writer/116535) recently where one or two data nodes would have high heap usage and the high CPU that goes along with constant garbage collections.

The culprit was lots of updates to existing documents.

---

<div class="post-metadata">

### Author: ![animageofmine](https://avatars.discourse-cdn.com/v4/letter/a/7feea3/32.png) [@animageofmine](https://discuss.elastic.co/u/animageofmine)
#### Post date: [February 8, 2018, 7:27am UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554/5 "2018-02-08T07:27:10Z")

</div>

@loren We update in batches (bulk). Our updates are essentially overwriting whole document (no partial updates). Since ES handles updates via delete followed by index, I guess it does end up deleting a lot of documents.

In your thread you seemed to be using X-Pack. Do you have some performance counters / metrics that you particular looked at? We don't use X-Pack, however have installed [telegraf](https://github.com/influxdata/telegraf/tree/master/plugins/inputs/elasticsearch) plugin.

---

<div class="post-metadata">

### Author: ![loren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loren/32/44942_2.png) [@loren](https://discuss.elastic.co/u/loren)
#### Post date: [February 8, 2018, 5:56pm UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554/6 "2018-02-08T17:56:32Z")

</div>

In my case X-Pack was not of much help diagnosing the problem as it doesn't graph per-node merge rates. I used `iostat` on the busy node to see that the write rate was 4X more than other nodes and took a guess that the problem was due to frequent merges. I reduced the number of updates by a wide margin, and this solved my problem.

Good luck!

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [March 8, 2018, 5:57pm UTC](https://discuss.elastic.co/t/troubleshooting-high-heap-usage/118554/7 "2018-03-08T17:57:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
