# Data-only node keeps crashing with oom error

**URL:** <https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286>\
**Category:** Elasticsearch\
**Created:** [August 12, 2015, 7:30pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286 "2015-08-12T19:30:34Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![slee](https://avatars.discourse-cdn.com/v4/letter/s/f4b2a3/32.png) [@slee](https://discuss.elastic.co/u/slee)\
**Post date:** [August 12, 2015, 7:30pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/1 "2015-08-12T19:30:34Z")

</div>

Hello, we have our cluster currently set up across 2 machines, with 1 machine having 1 data+master node, and the other machine having a data only node as well as 2 master nodes. The data only node, as well as the data+master node, are configured to have a reserved JAVA heap size of 32 GB. The 2 master nodes each have 16GB reserved heap size. Pretty frequently, I notice that the data only node will fail every so often, maybe every hour or so? When it fails, I see the error:

> [2015-08-12 14:33:50,813][WARN][netty.channel.DefaultChannelFuture] An exception was thrown by ChannelFutureListener.  
> java.lang.OutOfMemoryError: unable to create new native thread

I'm not sure why this is, as I can see that there are about 6GB used for fielddata, and that the used heap size % is at about 50%...I see a lot of other error messages, but they're all basically giving me the same outofmemory error message. I can't figure out why this node is crashing, especially since it's essentially the same as the data+master node in terms of configuration, and if anything that node should be using more memory than the data only node! Can anybody help?

---

<div class="post-metadata">

**Author:** ![msimos](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@msimos](https://discuss.elastic.co/u/msimos)\
**Post date:** [August 12, 2015, 9:38pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/2 "2015-08-12T21:38:48Z")

</div>

Try setting indices.fielddata.cache.size to 20% and see if that stabilizes things.

[https://www.elastic.co/guide/en/elasticsearch/reference/current/index-modules-fielddata.html#fielddata-monitoring](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-modules-fielddata.html#fielddata-monitoring)

You may want to lower the amount of heap below 32GB. As around this point is where java decompresses its pointers.

[https://www.elastic.co/guide/en/elasticsearch/guide/current/heap-sizing.html#compressed\_oops](https://www.elastic.co/guide/en/elasticsearch/guide/current/heap-sizing.html#compressed_oops)

You can read this as to why:

> **[Why 35GB Heap is Less Than 32GB - Java JVM Memory Oddities - codecentric AG Blog](https://blog.codecentric.de/en/2014/02/35gb-heap-less-32gb-java-jvm-memory-oddities/)**
>
> When I run the following Java program, which consumes all free memory with artificial data structures, something interesting happens. Something very similar happens to all real applications: https://gist.github.com/CodingFabian/8708393 \[node04\] ~ ➜...

Try using 30.5GB and see if you see any effect as well.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [August 13, 2015, 6:14am UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/3 "2015-08-13T06:14:22Z")

</div>

> Try using 30.5GB and see if you see any effect as well.

Is there a reference for the 30.5 GB number? I've always believed that as long as the heap is \<32 GB compressed pointers would be used and I've been using 31 GB to be on the safe side.

---

<div class="post-metadata">

**Author:** ![slee](https://avatars.discourse-cdn.com/v4/letter/s/f4b2a3/32.png) [@slee](https://discuss.elastic.co/u/slee)\
**Post date:** [August 13, 2015, 2:13pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/4 "2015-08-13T14:13:37Z")

</div>

Thanks, I'll try setting the fielddata cache size. Although, I did enable doc\_values, and it seems that stabilized the data+master node, but the data node is still unstable. I misspoke earlier, in the sysconfig/elasticsearch file I set the ES\_HEAP\_SIZE to be 31G, but when running TOP it shows the reserved size as 32G. As long as it's below 32G, shouldn't it still use the smaller pointers? And since the config for both the data+master and the data-only node is the same, why would one by unstable but not the other? Both machines are identical, and in fact there is more load on the data+master machine since we have all the logstash instances as well as redis running on it.

---

<div class="post-metadata">

**Author:** ![nilsga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nilsga/32/482_2.png) [@nilsga](https://discuss.elastic.co/u/nilsga)\
**Post date:** [August 13, 2015, 5:34pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/5 "2015-08-13T17:34:18Z")

</div>

> [@slee](#):
>
> unable to create new native thread

I don't think this is due to lack of memory. The error message indicates that your application is starting too many threads, or you have reached your process limit for the user running the process. Check with `ulimit -a` and check the settings for max user processes and file descriptors.

---

<div class="post-metadata">

**Author:** ![slee](https://avatars.discourse-cdn.com/v4/letter/s/f4b2a3/32.png) [@slee](https://discuss.elastic.co/u/slee)\
**Post date:** [August 13, 2015, 5:44pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/6 "2015-08-13T17:44:00Z")

</div>

Thanks for the tip. Taking a look, when I run ulimit -a as root I see that the max user processes is 1031433. Seems pretty high, I don't imagine that's the issue? Elasticsearch is running under the context of the elasticsearch user, does that max user processes value change for each user? How do I check the value for Elasticsearch, since it's not a valid login?

edit: the very first error message I see is

> [logstash-firewall-2015.08.13][0] failed engine [out of memory]  
> java.lang.OutOfMemoryError: unable to create new native thread

would this point to the process limit?

---

<div class="post-metadata">

**Author:** ![msimos](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@msimos](https://discuss.elastic.co/u/msimos)\
**Post date:** [August 13, 2015, 6:31pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/7 "2015-08-13T18:31:57Z")

</div>

Also check ulimit -n (open files). This may prevent creating another thread, if its set too low.

[http://localhost:9200/\_nodes/process?pretty&human](http://localhost:9200/_nodes/process?pretty&human)

Look at max\_file\_descriptors. I don't think Elasticsearch reports the max processes, you'll need to change the shell for the elasticsearch user (chsh) and then do a su -l elasticsearch -c "ulimit -a"

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:55pm UTC](https://discuss.elastic.co/t/data-only-node-keeps-crashing-with-oom-error/27286/8 "2017-07-05T23:55:55Z")

</div>


