# Node goes down showing fatal error in network layer , thread and java heapspace error

**URL:** <https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852>\
**Category:** Elasticsearch\
**Created:** [July 4, 2019, 7:04am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852 "2019-07-04T07:04:21Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 4, 2019, 7:04am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/1 "2019-07-04T07:04:21Z")

</div>

I am having a cluster with 5 master nodes,12 coordinator nodes and 60 data nodes .Currently i am doing heavy indexing in this es cluster around 15 billion documents spread through the day.We have 3 index which are undergoing heavy indexing there are four rollover in a day for each indexes.Each indexes having 100 shards and the replica is set to 1.The nodes are up on a physical server having 200gb of RAM ,each nodes have around 32gb of heap and the translog durability is set to async.

I am getting the following error and the node goes down.

```
[2019-07-03T21:28:59,331][ERROR][o.e.t.n.Netty4Utils] fatal error on the network layer
	at org.elasticsearch.transport.netty4.Netty4Utils.maybeDie(Netty4Utils.java:140)

[2019-07-03T21:28:59,336][ERROR][o.e.b.ElasticsearchUncaughtExceptionHandler] [data_8] fatal error in thread [Thread-34032], exiting
java.lang.OutOfMemoryError: Java heap space

[2019-07-03T21:28:59,397][WARN][o.e.t.n.Netty4Transport] [data_8] exception caught on transport layer [[id: 0x07b04e0b, L:/56.241.23.137:9303 - R:/56.241.23.147:35014]], closing connection
org.elasticsearch.ElasticsearchException: java.lang.OutOfMemoryError: Java heap space

```

Please suggest .This is a very trivial issue we are facing.

---

<div class="post-metadata">

**Author:** ![kumarpiyush](https://avatars.discourse-cdn.com/v4/letter/k/85e7bf/32.png) [@kumarpiyush](https://discuss.elastic.co/u/kumarpiyush)\
**Post date:** [July 4, 2019, 7:12am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/2 "2019-07-04T07:12:35Z")

</div>

I am having same issue.  
Whenever a heavy query is fired on my cluster node timeout exception occur and sometime node goes down.

Hoping to hear you soon.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 4, 2019, 7:21am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/3 "2019-07-04T07:21:37Z")

</div>

There have been several posts about a very similar cluster over the last few days and suggestions have been given. Have any of these made any difference? If the problem is trivial, why has it not been resolved?

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 4, 2019, 7:31am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/4 "2019-07-04T07:31:00Z")

</div>

Hi @Christian_Dahlqvist

I have tried those suggestions but this haven't made any difference for me.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 4, 2019, 8:45am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/5 "2019-07-04T08:45:46Z")

</div>

Are you working with @Sourabh?

---

<div class="post-metadata">

**Author:** ![kumarpiyush](https://avatars.discourse-cdn.com/v4/letter/k/85e7bf/32.png) [@kumarpiyush](https://discuss.elastic.co/u/kumarpiyush)\
**Post date:** [July 4, 2019, 8:54am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/6 "2019-07-04T08:54:27Z")

</div>

No @dadoonet .  
But going through his errors and similar posts I found out that we are doing same thing but at slightly less scale and slightly smaller cluster as compared to @Sourabh . But I am finding the same issue.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 4, 2019, 2:18pm UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/7 "2019-07-04T14:18:12Z")

</div>

So it's better to open your own question and give the details of your cluster such as:

```auto
GET /
GET /_cat/nodes?v
GET /_cat/health?v
GET /_cat/indices?v

```

If some outputs are too big, please share them on [gist.github.com](http://gist.github.com) and link them here.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 4, 2019, 6:29pm UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/8 "2019-07-04T18:29:01Z")

</div>

I did suggest decreasing the number of primary shards, but you still report having 100. Did you add any nodes to increase the available amount of heap? I thought you had 60 data nodes before I suggested adding more.

Exactly which suggestions did you try and what was the effect?

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 11, 2019, 5:45am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/9 "2019-07-11T05:45:13Z")

</div>

I have decreased the number of primary shards to 60 but still i am getting the following error.

> [2019-07-11T07:54:17,148][ERROR][o.e.b.ElasticsearchUncaughtExceptionHandler] [data\_15] fatal error in thread [Thread-142209], exiting  
> java.lang.OutOfMemoryError: Java heap space  
> at io.netty.buffer.PoolArena$HeapArena.newChunk(PoolArena.java:656) ~[?:?]  
> at io.netty.buffer.PoolArena.allocateNormal(PoolArena.java:237) ~[?:?]  
> at io.netty.buffer.PoolArena.allocate(PoolArena.java:221) ~[?:?]  
> at io.netty.buffer.PoolArena.allocate(PoolArena.java:141) ~[?:?]  
> at io.netty.buffer.PooledByteBufAllocator.newHeapBuffer(PooledByteBufAllocator.java:272) ~[?:?]  
> at io.netty.buffer.AbstractByteBufAllocator.heapBuffer(AbstractByteBufAllocator.java:160) ~[?:?]  
> at io.netty.buffer.AbstractByteBufAllocator.heapBuffer(AbstractByteBufAllocator.java:151) ~[?:?]  
> at io.netty.buffer.AbstractByteBufAllocator.ioBuffer(AbstractByteBufAllocator.java:133) ~[?:?]  
> at io.netty.channel.DefaultMaxMessagesRecvByteBufAllocator$MaxMessageHandle.allocate(DefaultMaxMessagesRecvByteBufAllocator.java:73) ~[?:?]  
> at io.netty.channel.nio.AbstractNioByteChannel$NioByteUnsafe.read(AbstractNioByteChannel.java:117) ~[?:?]  
> at io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:642) ~[?:?]  
> at io.netty.channel.nio.NioEventLoop.processSelectedKeysPlain(NioEventLoop.java:527) ~[?:?]  
> at io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:481) ~[?:?]  
> at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:441) ~[?:?]  
> at io.netty.util.concurrent.SingleThreadEventExecutor$5.run(SingleThreadEventExecutor.java:858) ~[?:?]  
> at java.lang.Thread.run(Thread.java:748) [?:1.8.0\_161]  
> Caused by: java.lang.OutOfMemoryError: Java heap space  
> at io.netty.buffer.PoolArena$HeapArena.newChunk(PoolArena.java:656) ~[?:?]  
> at io.netty.buffer.PoolArena.allocateNormal(PoolArena.java:237) ~[?:?]  
> at io.netty.buffer.PoolArena.allocate(PoolArena.java:221) ~[?:?]  
> at io.netty.buffer.PoolArena.allocate(PoolArena.java:141) ~[?:?]  
> at io.netty.buffer.PooledByteBufAllocator.newHeapBuffer(PooledByteBufAllocator.java:272) ~[?:?]  
> at io.netty.buffer.AbstractByteBufAllocator.heapBuffer(AbstractByteBufAllocator.java:160) ~[?:?]  
> at io.netty.buffer.AbstractByteBufAllocator.heapBuffer(AbstractByteBufAllocator.java:151) ~[?:?]  
> at io.netty.buffer.AbstractByteBufAllocator.ioBuffer(AbstractByteBufAllocator.java:133) ~[?:?]  
> at io.netty.channel.DefaultMaxMessagesRecvByteBufAllocator$MaxMessageHandle.allocate(DefaultMaxMessagesRecvByteBufAllocator.java:73) ~[?:?]  
> at io.netty.channel.nio.AbstractNioByteChannel$NioByteUnsafe.read(AbstractNioByteChannel.java:117) ~[?:?]  
> ... 6 more  
> SymbolTable statistics:  
> Number of buckets : 20011 = 160088 bytes, avg 8.000  
> Number of entries : 148658 = 3567792 bytes, avg 24.000  
> Number of literals : 148658 = 9189616 bytes, avg 61.817  
> Total footprint : = 12917496 bytes  
> Average bucket size : 7.429  
> Variance of bucket size : 7.497  
> Std. dev. of bucket size: 2.738  
> Maximum bucket size : 20  
> StringTable statistics:  
> Number of buckets : 500009 = 4000072 bytes, avg 8.000  
> Number of entries : 23025 = 552600 bytes, avg 24.000  
> Number of literals : 23025 = 3138896 bytes, avg 136.326  
> Total footprint : = 7691568 bytes  
> Average bucket size : 0.046  
> Variance of bucket size : 0.046  
> Std. dev. of bucket size: 0.215  
> Maximum bucket size : 3

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2019, 6:57am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/10 "2019-07-11T06:57:42Z")

</div>

Did you increase the number of data nodes and thus the total amount of heap available to the cluster?

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 11, 2019, 7:15am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/11 "2019-07-11T07:15:13Z")

</div>

Due to resource restriction i can't increase the number of data nodes but i have reduced the number of shards to 60.The amount of heap available to each datanode is around 32gb.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2019, 7:26am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/12 "2019-07-11T07:26:52Z")

</div>

A typical data node has 64GB RAM out of which ~30GB is allocated for heap. If your hosts have 200GB RAM you should be able to run 3 data nodes per host. This will naturally reduce the size of the OS page cache.

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 12, 2019, 7:17am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/13 "2019-07-12T07:17:46Z")

</div>

The server that has has 200gb of ram has low storage capacity if increase the number of datanodes then there will be lot of reallocation of shards across datanodes.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 12, 2019, 7:33am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/14 "2019-07-12T07:33:39Z")

</div>

How much RAM do the nodes with large storage capacity have? If you could highlight the hardware profile of the different node types it would be easier to provide guidance.

---

<div class="post-metadata">

**Author:** ![jonathan\_rowe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jonathan_rowe/32/49275_2.png) [@jonathan\_rowe](https://discuss.elastic.co/u/jonathan_rowe)\
**Post date:** [July 12, 2019, 2:32pm UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/15 "2019-07-12T14:32:38Z")

</div>

don't set a 32gb heap, 64-bit pointers will kick in, stick to \<30gb

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 12, 2019, 3:16pm UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/16 "2019-07-12T15:16:37Z")

</div>

What does 64-bit pointer mean?

And everyday one or more node goes down with heap space error.And in the logs i see old gc.

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 12, 2019, 3:21pm UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/17 "2019-07-12T15:21:41Z")

</div>

And in logs today i saw new IndexWriter is closed exception.

> org.apache.lucene.store.AlreadyClosedException: this IndexWriter is closed

When i run free -g command in the servers where i am running the nodes sometimes the there is no free space available all the free space is in the buffer.I have to free the cache using the command.

---

<div class="post-metadata">

**Author:** ![jonathan\_rowe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jonathan_rowe/32/49275_2.png) [@jonathan\_rowe](https://discuss.elastic.co/u/jonathan_rowe)\
**Post date:** [July 13, 2019, 9:08am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/18 "2019-07-13T09:08:58Z")

</div>

It means that object references on the heap are 8 bytes instead of 4 and consume more memory/cache/memory bandwidth etc

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [July 15, 2019, 6:55am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/19 "2019-07-15T06:55:16Z")

</div>

Hi @Christian_Dahlqvist

I have 14 server of 500gb ram in which one data node is running and remaining servers have around 200gb of ram in which 2 data nodes are running in each servers.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 12, 2019, 6:55am UTC](https://discuss.elastic.co/t/node-goes-down-showing-fatal-error-in-network-layer-thread-and-java-heapspace-error/188852/20 "2019-08-12T06:55:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
