# OOM killer triggered and machine crashes after using up all memory

**URL:** <https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725>\
**Category:** Elasticsearch\
**Created:** [April 13, 2019, 9:21am UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725 "2019-04-13T09:21:11Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jalaluddeen\_Mohammed](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jalaluddeen_mohammed/32/44079_2.png) [@Jalaluddeen\_Mohammed](https://discuss.elastic.co/u/Jalaluddeen_Mohammed)\
**Post date:** [April 13, 2019, 9:21am UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/1 "2019-04-13T09:21:11Z")

</div>

I have a 12 node cluster with 24G RAM running only ELasticsearch.  
Elasticsearch is configured to use 11G of RAM as below.

-Xms11g  
-Xmx11g

There is no indexing or querying happening on this cluster, this is a fresh deployment, basically sitting idle.  
But after for a few days, machine crashes and I have to reboot the machine.  
There is no Heap dump or any GC events in the elasticsearch logs.  
THis has happened in 3 machines in the cluster now.

Upon checking dmesg it says  
//Out of memory: Kill process 19590 (java) score 507 or sacrifice child

Below is the relevant dmesg output.

\<\>  
Node 0 Normal free:45816kB min:59172kB low:73964kB high:88756kB active\_anon:8kB inactive\_anon:0kB active\_file:544kB inactive\_file:412kB unevictable:12650952kB isolated(anon):0kB isolated(file):128kB present:21719040kB mlocked:207080kB dir  
ty:0kB writeback:0kB mapped:116920kB shmem:0kB slab\_reclaimable:9872kB slab\_unreclaimable:42008kB kernel\_stack:5680kB pagetables:29844kB unstable:0kB bounce:0kB writeback\_tmp:0kB pages\_scanned:0 all\_unreclaimable? no  
lowmem\_reserve: 0 0 0 0  
Node 0 DMA: 1_4kB 1_8kB 2_16kB 1_32kB 2_64kB 0_128kB 0_256kB 0_512kB 1_1024kB 1_2048kB 3_4096kB = 15564kB  
Node 0 DMA32: 12_4kB 9_8kB 7_16kB 29_32kB 16_64kB 4_128kB 6_256kB 9_512kB 8_1024kB 7_2048kB 15_4096kB = 92808kB  
Node 0 Normal: 10433_4kB 23_8kB 0_16kB 0_32kB 0_64kB 0_128kB 0_256kB 0_512kB 0_1024kB 0_2048kB 1\*4096kB = 46012kB  
29638 total pagecache pages  
109 pages in swap cache  
Swap cache stats: add 116875, delete 116766, find 4926/7159  
Free swap = 631592kB  
Total swap = 1042140kB  
6291440 pages RAM  
139727 pages reserved  
35020 pages shared  
6079389 pages non-shared  
[pid] uid tgid total\_vm rss cpu oom\_adj oom\_score\_adj name  
[638] 0 638 2731 69 3 -17 -1000 udevd  
[1149] 0 1149 2730 69 1 -17 -1000 udevd  
[1264] 0 1264 2730 66 7 -17 -1000 udevd  
[1714] 0 1714 7441 132 5 -17 -1000 auditd  
[1738] 0 1738 1540 121 0 0 0 portreserve  
[1748] 0 1748 63919 217 1 0 0 rsyslogd  
[1763] 0 1763 4586 104 2 0 0 irqbalance  
[1785] 32 1785 4745 124 5 0 0 rpcbind  
[1809] 29 1809 5838 180 6 0 0 rpc.statd  
[1839] 0 1839 1671 92 2 0 0 vnstatd  
[1850] 81 1850 5359 87 6 0 0 dbus-daemon  
[1888] 0 1888 1019 131 6 0 0 acpid  
[1900] 68 1900 9581 200 2 0 0 hald  
[1901] 0 1901 5099 132 1 0 0 hald-runner  
[1934] 0 1934 5629 119 2 0 0 hald-addon-inpu  
[1948] 68 1948 4501 162 6 0 0 hald-addon-acpi  
[2126] 0 2126 43169 249 1 0 0 vmtoolsd  
[2179] 0 2179 96539 180 2 0 0 automount  
[2311] 0 2311 7979 124 2 -17 -1000 sshd  
[2324] 0 2324 5428 161 6 0 0 xinetd  
[2335] 38 2335 7685 239 7 0 0 ntpd  
[2481] 0 2481 20253 238 1 0 0 master  
[2491] 89 2491 20316 238 7 0 0 qmgr  
[2495] 0 2495 45773 202 3 0 0 abrtd  
[2523] 0 2523 29221 148 5 0 0 crond  
[2538] 0 2538 5276 68 4 0 0 atd  
[2958] 0 2958 148548 232 0 0 0 salt-minion  
[2959] 0 2959 123917 59 6 0 0 salt-minion  
[2987] 0 2987 16119 84 0 0 0 certmonger  
[3020] 0 3020 1015 115 3 0 0 mingetty  
[3024] 0 3024 1015 115 3 0 0 mingetty  
[3027] 0 3027 1015 115 7 0 0 mingetty  
[3030] 0 3030 1015 115 1 0 0 mingetty  
[3034] 0 3034 1015 115 4 0 0 mingetty  
[3038] 0 3038 1015 115 3 0 0 mingetty  
[19590] 61947 19590 4583234 3162869 1 0 0 java  
[12304] 89 12304 20273 225 1 0 0 pickup  
Out of memory: Kill process 19590 (java) score 507 or sacrifice child  
Killed process 19590, UID 61947, (java) total-vm:18332936kB, anon-rss:12534556kB, file-rss:116920kB  
[0]: VMCI: Updating context from (ID=0xea4a4c50) to (ID=0xea4a4c50) on event (type=0).

\</\>

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 13, 2019, 9:45am UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/2 "2019-04-13T09:45:41Z")

</div>

Elasticsearch used memory in addition to the heap you have configured, and it is generally recommended to give it the same amount of off-heap memory as you have heap. If you have configured the heap to be 11GB, Elasticsearch should ideally according to this rule of thumb have access to 22GB or memory in total. If you have other components or services running on this host it is possible that this is what is causing problems. I also see that you are using swap space, which is something that is not recommended with Elasticsearch.

It might also help if you were able to provide a bit more information about your environment, e.g. size of data and the version used.

---

<div class="post-metadata">

**Author:** ![Jalaluddeen\_Mohammed](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jalaluddeen_mohammed/32/44079_2.png) [@Jalaluddeen\_Mohammed](https://discuss.elastic.co/u/Jalaluddeen_Mohammed)\
**Post date:** [April 13, 2019, 3:16pm UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/3 "2019-04-13T15:16:17Z")

</div>

Thanks!

I using Elasticsearch version 5.6.16 and have around 480G of data in total from 3 indices, 12 node cluster with 3.4TB total available on each node.

There are no other apps except for salt-minion on the nodes.  
OS is Centos 6.10 .. Kernel - 2.6.32-754.10.1.el6.x86\_64

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 13, 2019, 3:58pm UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/4 "2019-04-13T15:58:32Z")

</div>

Is there anything in the Elasticsearch logs?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 13, 2019, 5:19pm UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/5 "2019-04-13T17:19:37Z")

</div>

> [@Jalaluddeen\_Mohammed](#):
>
> `Node 0 Normal: 10433 *4kB 23* 8kB 0 *16kB 0* 32kB 0 *64kB 0* 128kB 0 *256kB 0* 512kB 0 *1024kB 0* 2048kB 1*4096kB = 46012kB`

This indicates that your off-heap memory is very fragmented. Do you have any nonstandard kernel settings? In particular, what is `vm.overcommit_memory` set to?

---

<div class="post-metadata">

**Author:** ![Jalaluddeen\_Mohammed](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jalaluddeen_mohammed/32/44079_2.png) [@Jalaluddeen\_Mohammed](https://discuss.elastic.co/u/Jalaluddeen_Mohammed)\
**Post date:** [April 13, 2019, 5:27pm UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/6 "2019-04-13T17:27:29Z")

</div>

Hi,

/proc/sys/vm/overcommit\_memory = 0

/proc/sys/vm/overcommit\_ratio = 50

Kernel version 2.6.32-754.10.1.el6.x86\_64

OS Centos 6.10

Below are the log entries from Elasticsearch around the time of crash.

[2019-04-12T03:16:50,736][WARN][o.e.m.j.JvmGcMonitorService] [AAzI9z\_] [gc][young][175648][31] duration [12.3s], collections [1]/[12.8s], total [12.3s]/[14.2s], memory [1gb]-\>[563.1mb]/[10.9gb], all\_pools {[young] [524.1mb]-\>[5.5mb]/[532.5mb]}{[survivor] [237.1kb]-\>[51.9mb]/[66.5mb]}{[old] [505.6mb]-\>[505.6mb]/[10.3gb]}

[2019-04-12T03:16:50,738][WARN][o.e.m.j.JvmGcMonitorService] [AAzI9z\_] [gc][175648] overhead, spent [12.3s] collecting in the last [12.8s]

[2019-04-12T03:17:18,525][WARN][o.e.m.j.JvmGcMonitorService] [AAzI9z\_] [gc][young][175664][32] duration [12.1s], collections [1]/[12.6s], total [12.1s]/[26.3s], memory [1gb]-\>[657.5mb]/[10.9gb], all\_pools {[young] [508mb]-\>[4.3mb]/[532.5mb]}{[survivor] [51.9mb]-\>[66.5mb]/[66.5mb]}{[old] [505.6mb]-\>[587mb]/[10.3gb]}

[2019-04-12T03:17:18,526][WARN][o.e.m.j.JvmGcMonitorService] [AAzI9z\_] [gc][175664] overhead, spent [12.1s] collecting in the last [12.6s]

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 14, 2019, 7:30am UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/7 "2019-04-14T07:30:06Z")

</div>

> [@Jalaluddeen\_Mohammed](#):
>
> ```plaintext
> /proc/sys/vm/overcommit_memory = 0
> 
> ```

Thanks, ok, this is the default so I think that's not it.

> [@Jalaluddeen\_Mohammed](#):
>
> ```plaintext
> [gc][175664] overhead, spent [12.1s] collecting in the last [12.6s]
> 
> ```

This suggests that this node is under a lot of heap pressure, which is weird if the cluster is idle. Can you tell us a bit more about this cluster: for instance how many shards does it have?

I agree with the advice above that it'd be a good idea to disable swap for [the reasons described in the manual](https://www.elastic.co/guide/en/elasticsearch/reference/5.6/setup-configuration-memory.html). Also note that the 5.6 series is now [past the end of its supported life](https://www.elastic.co/support/eol) so it could be tricky to dig into this in detail. Upgrading is recommended, although I don't know of any changes in 6.x that would address these specific symptoms.

Ideally I think I'd like to see a heap dump taken when a node goes OOM, but unfortunately we don't get a heap dump if a process is killed by the OOM killer. Are there GC warnings on other nodes like this, suggesting that other nodes are heading towards failure? If so, can you take a heap dump from such a node?

---

<div class="post-metadata">

**Author:** ![Jalaluddeen\_Mohammed](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jalaluddeen_mohammed/32/44079_2.png) [@Jalaluddeen\_Mohammed](https://discuss.elastic.co/u/Jalaluddeen_Mohammed)\
**Post date:** [April 14, 2019, 8:09am UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/8 "2019-04-14T08:09:30Z")

</div>

Thank you.

The cluster has 42 shards in total.

Unfortunately there are no heap dumps in any of the 12 nodes in the cluster.  
And around the time of crash there are no GC warnings in the ES logs.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 14, 2019, 8:13am UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/9 "2019-04-14T08:13:36Z")

</div>

> [@Jalaluddeen\_Mohammed](#):
>
> Unfortunately there are no heap dumps in any of the 12 nodes in the cluster.

Sure, you'll have to [capture one by hand](https://www.baeldung.com/java-heap-dump-capture).

> [@Jalaluddeen\_Mohammed](#):
>
> And around the time of crash there are no GC warnings in the ES logs.

I'm confused, because previously you said:

> [@Jalaluddeen\_Mohammed](#):
>
> Below are the log entries from Elasticsearch around the time of crash.

---

<div class="post-metadata">

**Author:** ![Jalaluddeen\_Mohammed](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jalaluddeen_mohammed/32/44079_2.png) [@Jalaluddeen\_Mohammed](https://discuss.elastic.co/u/Jalaluddeen_Mohammed)\
**Post date:** [April 14, 2019, 3:59pm UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/10 "2019-04-14T15:59:38Z")

</div>

What I meant is there are no GC warnings from other nodes, the above listed GC log is from the machine that crashed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 12, 2019, 3:59pm UTC](https://discuss.elastic.co/t/oom-killer-triggered-and-machine-crashes-after-using-up-all-memory/176725/11 "2019-05-12T15:59:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
