# Suspicious memory leak due to netty PoolThreadCache

**URL:** <https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730>\
**Category:** Elasticsearch\
**Created:** [January 10, 2019, 12:11pm UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730 "2019-01-10T12:11:01Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![howardhuang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/howardhuang/32/69365_2.png) [@howardhuang](https://discuss.elastic.co/u/howardhuang)\
**Post date:** [January 10, 2019, 12:11pm UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/1 "2019-01-10T12:11:01Z")

</div>

Hi,

We have 10 nodes based on ES 5.6.4, and each node 8Core, 8GB memory. The cluster only have one index with 10 shards 1 replica. Each shard has around 180GB data (single big shard is a history issue).

One day, cluster got a few bulk reject errors, each node's heap memory used up to around 80%, after manually triggered old gc, the memory still cannot be reduced. We dumped one of the node's memory, and rolling restarted all the nodes, cluster got recovered, and each node's memory usage stable at around 20%.

From the dumped file, we found most of the memory used by netty pool cache.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/7/57e38e5e57d67ea53b4a2f5eefbb48e8f0182ee6.png)

257 netty pool chunks, each chunk is 16MB, total around 4GB:  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/d/f/dfec2a7fc9f898d3c8a71465e3b00ad6b2559a80.png)

Here is the GC root path of byte, only kept strong reference:  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/e/b/eb578407f5e4fb7f41b1dec9b3bc92f0acb82c9b.png)

Why the bulk thread local buffer cache cannot be released? Any idea about the huge memory consumption?

Thanks,  
Howard

---

<div class="post-metadata">

**Author:** ![howardhuang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/howardhuang/32/69365_2.png) [@howardhuang](https://discuss.elastic.co/u/howardhuang)\
**Post date:** [January 11, 2019, 1:50am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/2 "2019-01-11T01:50:29Z")

</div>

Any feedback is appreciate!

---

<div class="post-metadata">

**Author:** ![howardhuang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/howardhuang/32/69365_2.png) [@howardhuang](https://discuss.elastic.co/u/howardhuang)\
**Post date:** [January 14, 2019, 2:13am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/3 "2019-01-14T02:13:46Z")

</div>

Any comment is appreciate!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 14, 2019, 6:43am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/4 "2019-01-14T06:43:16Z")

</div>

I am not the right person to look at this, but I suspect it would help if you could provide your Elasticsearch config as well as any non-default JVM settings you are using.

What type of load is the cluster under? What is the use-case?

---

<div class="post-metadata">

**Author:** ![howardhuang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/howardhuang/32/69365_2.png) [@howardhuang](https://discuss.elastic.co/u/howardhuang)\
**Post date:** [January 15, 2019, 4:04am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/5 "2019-01-15T04:04:18Z")

</div>

Thank you @Christian_Dahlqvist ,

Here is our es config:

```auto
cluster.routing.allocation.disk.include_relocations: false
node.ingest: true
indices.memory.index_buffer_size: 15%
path.data: /data1/containers/1539769914008324911/es/data
processors: 8
node.name: 1539769914008324911
search.remote.connect: false
indices.queries.cache.count: 500
bootstrap.seccomp: false
action.destructive_requires_name: true
indices.queries.cache.size: 5%
cluster.routing.allocation.awareness.attributes: ip
cluster.routing.allocation.disk.watermark.low: 90%
ingest.new_date_format: true
discovery.zen.minimum_master_nodes: 6
http.port: 9200
cluster.routing.allocation.disk.watermark.high: 95%
node.master: true
thread_pool.bulk.queue_size: 230
node.data: true
network.publish_host: 100.125.51.15
node.attr.ip: 100.125.51.15
network.host: 0.0.0.0
cluster.name: Lambda
discovery.zen.ping.unicast.hosts: [ten nodes ip:port list]

```

And except xms/xmx, we don't have non-default JVM settings:

```auto
-Xms7896m
-Xmx7896m

```

Here is the average JVM usage for each node during that time, and eacho node cpu utilization around 20%. After rolling restart nodes, jvm usage comes down.  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/8/3830f1456c9ac82df55c7a449be4c8506bf622f5.png)

Cluster has single index with 10 shards 1 replica:

```auto
green open lambda 68_xZphGTG6aGhXhQsjcYw 10 1 1754643122 2859 1.8tb 1.8tb

```

And we could see some bulk reject, and node memory cannot be reduced by old gc, it's used up by netty pool buffer as above description.

Thanks,  
Howard

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 15, 2019, 7:30am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/6 "2019-01-15T07:30:42Z")

</div>

What is the average size of your documents? What is the average size of your bulk requests? How many shards are you actively indexing into?

---

<div class="post-metadata">

**Author:** ![howardhuang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/howardhuang/32/69365_2.png) [@howardhuang](https://discuss.elastic.co/u/howardhuang)\
**Post date:** [January 15, 2019, 7:53am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/7 "2019-01-15T07:53:24Z")

</div>

Thank you @Christian_Dahlqvist , average size of doc is 5kb, and average bulk size is 5000, the target index has 10 primary shards and each shard has one replica, total 20 shards.

---

<div class="post-metadata">

**Author:** ![howardhuang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/howardhuang/32/69365_2.png) [@howardhuang](https://discuss.elastic.co/u/howardhuang)\
**Post date:** [January 15, 2019, 7:57am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/8 "2019-01-15T07:57:32Z")

</div>

Currently cluster runs correct without exception, I also dumped current memory for analysising. I found less then 10 bytes array which contains netty pool chunk.  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/4/f4ab3d5126bbab94e13830f93c65dcd2f76aaf39.png)

But in previous exception case, dump one node has around 500 netty pool chunk bytes array which used up more than 4GB heap. (total heap size 8GB)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 12, 2019, 8:07am UTC](https://discuss.elastic.co/t/suspicious-memory-leak-due-to-netty-poolthreadcache/163730/9 "2019-02-12T08:07:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
