# \[ Frequent OOME on coordinator node \]

**URL:** <https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539>\
**Category:** Elasticsearch\
**Created:** [April 18, 2018, 1:10pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539 "2018-04-18T13:10:24Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![jerome831361](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome831361/32/30125_2.png) [@jerome831361](https://discuss.elastic.co/u/jerome831361)\
**Post date:** [April 18, 2018, 1:10pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/1 "2018-04-18T13:10:24Z")

</div>

Hi,  
Since some 5 days, my coordinating nodes are both crashing frequenty with a OOME.

I have tried to gather some usefull information to help investigating the problem; but I can't figure out what is wrong.

I'm not really sure what to check first; so I'm trying to find some help here 🙂  
That's strange because I didn't find any corelable change in the configuration or in the data volume (maybe I missed it thought)

## **Here are some information to help investigating:**

Problem:  
My 2 coordinator nodes are crashing with a OOME

Logs:  
[2018-04-17T15:14:51,535][WARN][i.n.c.AbstractChannelHandlerContext] An exception 'java.lang.OutOfMemoryError: Java heap space' [enable DEBUG level for full stacktrace] was thrown by a user handler's exceptionCaught() method while handling the following exception:  
java.lang.OutOfMemoryError: Java heap space  
[2018-04-17T15:14:53,038][ERROR][o.e.x.m.c.n.NodeStatsCollector] [coordinator1] collector [node\_stats] timed out when collecting data  
[2018-04-17T15:14:46,205][WARN][i.n.c.AbstractChannelHandlerContext] An exception 'java.lang.OutOfMemoryError: Java heap space' [enable DEBUG level for full stacktrace] was thrown by a user handler's exceptionCaught() method while handling the following exception:  
java.lang.OutOfMemoryError: Java heap space  
[2018-04-17T15:14:58,029][ERROR][o.e.t.n.Netty4Utils] fatal error on the network layer  
at org.elasticsearch.transport.netty4.Netty4Utils.maybeDie(Netty4Utils.java:185)  
at org.elasticsearch.transport.netty4.Netty4MessageChannelHandler.exceptionCaught(Netty4MessageChannelHandler.java:73)  
at io.netty.channel.AbstractChannelHandlerContext.invokeExceptionCaught(AbstractChannelHandlerContext.java:285)  
at io.netty.channel.AbstractChannelHandlerContext.invokeExceptionCaught(AbstractChannelHandlerContext.java:264)  
at io.netty.channel.AbstractChannelHandlerContext.fireExceptionCaught(AbstractChannelHandlerContext.java:256)  
at io.netty.channel.ChannelInboundHandlerAdapter.exceptionCaught(ChannelInboundHandlerAdapter.java:131)  
at io.netty.channel.AbstractChannelHandlerContext.invokeExceptionCaught(AbstractChannelHandlerContext.java:285)  
at io.netty.channel.AbstractChannelHandlerContext.notifyHandlerException(AbstractChannelHandlerContext.java:850)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:364)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:348)  
at io.netty.channel.AbstractChannelHandlerContext.fireChannelRead(AbstractChannelHandlerContext.java:340)  
at io.netty.channel.ChannelInboundHandlerAdapter.channelRead(ChannelInboundHandlerAdapter.java:86)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:362)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:348)  
at io.netty.channel.AbstractChannelHandlerContext.fireChannelRead(AbstractChannelHandlerContext.java:340)  
at io.netty.handler.logging.LoggingHandler.channelRead(LoggingHandler.java:241)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:362)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:348)  
at io.netty.channel.AbstractChannelHandlerContext.fireChannelRead(AbstractChannelHandlerContext.java:340)  
at io.netty.channel.DefaultChannelPipeline$HeadContext.channelRead(DefaultChannelPipeline.java:1334)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:362)  
at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:348)  
at io.netty.channel.DefaultChannelPipeline.fireChannelRead(DefaultChannelPipeline.java:926)  
at io.netty.channel.nio.AbstractNioByteChannel$NioByteUnsafe.read(AbstractNioByteChannel.java:134)  
at io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:644)  
at io.netty.channel.nio.NioEventLoop.processSelectedKeysPlain(NioEventLoop.java:544)  
at io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:498)  
at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:458)  
at io.netty.util.concurrent.SingleThreadEventExecutor$5.run(SingleThreadEventExecutor.java:858)  
at java.base/java.lang.Thread.run(Thread.java:844)  
[2018-04-17T15:14:48,801][ERROR][o.e.b.ElasticsearchUncaughtExceptionHandler] [coordinator1] fatal error in thread [Thread-7], exiting  
java.lang.OutOfMemoryError: Java heap space

Metrics:

- Indices : 2093 (daily based)
- Total shards : 4184
- ES version : 6.1.1
- 3 master nodes : 4 cpu / 8 GB ram
- 4 data nodes : 16 cpu / 32 GB ram
- 2 coordinator nodes : 2 cpu / 4 GB ram

Heap dump analysis:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/e/6e5bc1a77666837309c066dfc83bcf0b864ba6a4.png)

The heap dump analysis is quite the same on both coordinator nodes

Thank you for your help  
Best regards  
Jérôme

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 18, 2018, 1:56pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/2 "2018-04-18T13:56:08Z")

</div>

Which version is it?

4184 shards for 4 data nodes seem a bit too much.

---

<div class="post-metadata">

**Author:** ![jerome831361](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome831361/32/30125_2.png) [@jerome831361](https://discuss.elastic.co/u/jerome831361)\
**Post date:** [April 18, 2018, 2:16pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/3 "2018-04-18T14:16:21Z")

</div>

Hi,

The ES version is: 6.1.1

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 18, 2018, 6:49pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/4 "2018-04-18T18:49:47Z")

</div>

Could you upgrade?

---

<div class="post-metadata">

**Author:** ![jerome831361](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome831361/32/30125_2.png) [@jerome831361](https://discuss.elastic.co/u/jerome831361)\
**Post date:** [April 19, 2018, 7:38am UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/5 "2018-04-19T07:38:50Z")

</div>

OK, I will upgrade to the lastest version and keep this topic up to date.

Thanks for your help

Regards

---

<div class="post-metadata">

**Author:** ![jerome831361](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome831361/32/30125_2.png) [@jerome831361](https://discuss.elastic.co/u/jerome831361)\
**Post date:** [April 19, 2018, 12:21pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/6 "2018-04-19T12:21:22Z")

</div>

Just to keep you up to date:

I have rolling-ugraded my ES cluster to 6.2.4

Operations have ended at ~ 01:30 AM

For now, it seems the Heap of my coordinators nodes is doing well:

![image](https://us1.discourse-cdn.com/elastic/original/3X/1/3/13ceb1b5e034db877db373207c7aae9caaa16150.png)

I will give some news tomorrow

Best regards

---

<div class="post-metadata">

**Author:** ![jerome831361](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerome831361/32/30125_2.png) [@jerome831361](https://discuss.elastic.co/u/jerome831361)\
**Post date:** [April 20, 2018, 7:15am UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/7 "2018-04-20T07:15:43Z")

</div>

Hi,

Since the upgrade; the HEAP of my coordinator nodes is much more stable.

![image](https://us1.discourse-cdn.com/elastic/original/3X/b/8/b84fa93fd752a84eff83978956b6077bfe6d01cf.png)

So I think we can consider this issue as closed.

Many thanks for your help

Best regards  
Jérôme

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 20, 2018, 12:51pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/8 "2018-04-20T12:51:15Z")

</div>

Great.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 18, 2018, 12:51pm UTC](https://discuss.elastic.co/t/frequent-oome-on-coordinator-node/128539/9 "2018-05-18T12:51:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
