# Nodes slowly climbing to high memory usage

**URL:** <https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266>\
**Category:** Elasticsearch\
**Created:** [December 30, 2016, 10:43am UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266 "2016-12-30T10:43:21Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Aldo\_Stracquadanio](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aldo_stracquadanio/32/11978_2.png) [@Aldo\_Stracquadanio](https://discuss.elastic.co/u/Aldo_Stracquadanio)\
**Post date:** [December 30, 2016, 10:43am UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/1 "2016-12-30T10:43:21Z")

</div>

Hi, we are experiencing some issues with our Elasticsearch cluster: after around 24 hours of usage the nodes report more than 95% of heap usage and they start becoming unresponsive bringing down the whole cluster; we need to constantly restart them if we want to keep it alive. This is the command line we use to start the Elasticsearch process:

```auto
/usr/bin/java -Xms25g -Xmx25g -Djava.awt.headless=true -XX:+UseParNewGC -XX:+UseConcMarkSweepGC -XX:CMSInitiatingOccupancyFraction=75 -XX:+UseCMSInitiatingOccupancyOnly -XX:+HeapDumpOnOutOfMemoryError -XX:+DisableExplicitGC -Dfile.encoding=UTF-8 -Djna.nosys=true -Des.path.home=/usr/share/elasticsearch -cp /usr/share/elasticsearch/lib/elasticsearch-2.4.1.jar:/usr/share/elasticsearch/lib/* org.elasticsearch.bootstrap.Elasticsearch start -p /var/run/elasticsearch/elasticsearch.pid -d -Des.default.path.home=/usr/share/elasticsearch -Des.default.path.logs=/var/log/elasticsearch -Des.default.path.data=/var/lib/elasticsearch -Des.default.path.conf=/etc/elasticsearch

```

Another interesting detail is that in our CQRS architecture if we attach a cluster only to the write path the heap issue doesn't seem to occur but nodes start climbing up as soon as we attach them to the read path of the system.

This seems to point to some caching issue or to some query that is triggering the behaviour; on the latter, we do have a query to which we provide up to 5000 ids to exclude from the result as in a `WHERE NOT IN (...)` SQL query. Is there anything we should do to take care of this specific query?

Any advice on what could be causing this issue or how we can keep the heap usage from blowing up the cluster would be really appreciated.

Many thanks!

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [December 30, 2016, 2:52pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/2 "2016-12-30T14:52:34Z")

</div>

Could you post [node stats](https://www.elastic.co/guide/en/elasticsearch/reference/5.1/cluster-nodes-stats.html) for a node that reached 95% of heap?

---

<div class="post-metadata">

**Author:** ![Aldo\_Stracquadanio](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aldo_stracquadanio/32/11978_2.png) [@Aldo\_Stracquadanio](https://discuss.elastic.co/u/Aldo_Stracquadanio)\
**Post date:** [January 4, 2017, 9:47am UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/3 "2017-01-04T09:47:02Z")

</div>

Sorry for the late answer, I'm attaching all the stats I could find about a faulty node:

- Node stats: [https://gist.github.com/Astrac/779ab0057699d0caca8292acf8e8fe44](https://gist.github.com/Astrac/779ab0057699d0caca8292acf8e8fe44)
- Nodes: [https://gist.github.com/Astrac/a4dc9aabc638565ce07f26aa6d4fe15c](https://gist.github.com/Astrac/a4dc9aabc638565ce07f26aa6d4fe15c)
- Hot threads: [https://gist.github.com/Astrac/9a31630c4cbb7ee1d897987e7ebff154](https://gist.github.com/Astrac/9a31630c4cbb7ee1d897987e7ebff154)

Thanks again!

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 4, 2017, 4:50pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/4 "2017-01-04T16:50:35Z")

</div>

I couldn't see a smoking guns by looking at stats, but this portion might be interesting to investigate:

```auto
         "script":{
            "compilations":387991,
            "cache_evictions":387891
         }

```

Which scripting language are you using and how do you generate your scripts? Is it possible to move the dynamic portion of your script into parameters instead of generating scripts on the fly?

---

<div class="post-metadata">

**Author:** ![Aris\_Koliopoulos](https://avatars.discourse-cdn.com/v4/letter/a/858c86/32.png) [@Aris\_Koliopoulos](https://discuss.elastic.co/u/Aris_Koliopoulos)\
**Post date:** [January 4, 2017, 5:38pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/5 "2017-01-04T17:38:27Z")

</div>

Hi Igor,  
Thanks for replying.  
We use groovy. We generate the scripts in the application code, alongside the queries, and we pass them to ES (as opposed to stored groovy scripts). The scripts access document fields to compute a couple of values and change the score accordingly.  
This is used in two simple cases, it would be non-trivial to refactor and it is used quite a lot.

Do you reckon those get cached and never evicted?

Thanks,  
Aris

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 4, 2017, 6:01pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/6 "2017-01-04T18:01:03Z")

</div>

> Do you reckon those get cached and never evicted?

I don't know for sure. To figure out exactly what's leaking memory we would need to take a look at heap dump. Without it I can only guess and poke at things that don't look quite right (i.e. different from typical installations that don't experience memory leaks). In your case it looks like all these script are getting evicted, but you recompile them on every request which was causing [issues](https://github.com/elastic/elasticsearch/issues/18572) before. So I just assume that it might be possible that this issue wasn't quite fixed in your particular use case.

Any chance we can take a look at the heap dump?

---

<div class="post-metadata">

**Author:** ![Aris\_Koliopoulos](https://avatars.discourse-cdn.com/v4/letter/a/858c86/32.png) [@Aris\_Koliopoulos](https://discuss.elastic.co/u/Aris_Koliopoulos)\
**Post date:** [January 5, 2017, 9:45am UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/7 "2017-01-05T09:45:38Z")

</div>

There is a good chance you are right:  
I ran the eclipse analyzer in the dump and it seems that:  
`The class "org.codehaus.groovy.reflection.ClassInfo", loaded by "java.net.FactoryURLClassLoader @ 0x1c42d78b8", occupies 1,899,607,464 (43.61%) bytes. The memory is accumulated in class "org.codehaus.groovy.reflection.ClassInfo", loaded by "java.net.FactoryURLClassLoader @ 0x1c42d78b8".`  
which indicates that it is a prime leak suspect.

Looking at the histogram in `jhat`:

```auto

Class	Instance Count	Total Size
class [C	1064936	1229061570
class org.codehaus.groovy.runtime.metaclass.MetaMethodIndex$Entry	12038167	1011206028
class [B	422619	712977967
class [Ljava.lang.Object;	9556045	387915336
class [Lorg.codehaus.groovy.util.ComplexKeyHashMap$Entry;	968641	263472144
class [Lorg.codehaus.groovy.runtime.metaclass.MetaMethodIndex$Entry;	69207	142814832
class java.util.HashMap$Node	5040441	141132348
class [Ljava.util.HashMap$Node;	548700	130578824
class org.codehaus.groovy.util.FastArray	8716940	104603280
class java.lang.reflect.Method	585909	76168170
class org.codehaus.groovy.util.SingleKeyHashMap$Entry	2283352	63933856
class java.util.HashMap	1268142	60870816
class java.lang.invoke.MethodHandleImpl$CountingWrapper	1115156	60218424

```

I will take another dump today and I will share it. The HEAP increased by 3gb per node (on average) overnight,  
so I guess the diff of the two will reveal additional info.

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [January 5, 2017, 12:03pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/8 "2017-01-05T12:03:11Z")

</div>

> [@Aris\_Koliopoulos](#):
>
> I will take another dump today and I will share it.

That would be great. Just don't share the entire heap dump publicly since it might contain sensitive information about your cluster.

---

<div class="post-metadata">

**Author:** ![tanguy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tanguy/32/6030_2.png) [@tanguy](https://discuss.elastic.co/u/tanguy)\
**Post date:** [January 10, 2017, 1:27pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/9 "2017-01-10T13:27:21Z")

</div>

Hi,

It looks like another Groovy memory leak issue. Can you please share a sample of scripts that are executed? That might help to reproduce the issue.

Thanks

---

<div class="post-metadata">

**Author:** ![Aris\_Koliopoulos](https://avatars.discourse-cdn.com/v4/letter/a/858c86/32.png) [@Aris\_Koliopoulos](https://discuss.elastic.co/u/Aris_Koliopoulos)\
**Post date:** [January 16, 2017, 4:47pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/10 "2017-01-16T16:47:59Z")

</div>

Indeed Groovy was the issue (sorry for the late reply). Removing groovy fixed the issue.  
The scripts looked like his:

```auto
scriptScore(
      s"doc['SomeDoubleField'].value / (1 + Math.log(1 + doc['SomeIntField'].value)) "
    )

```

There were 2 variations of this thing, single line script scores performing basic math calculations on relatively small result sets.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 13, 2017, 4:48pm UTC](https://discuss.elastic.co/t/nodes-slowly-climbing-to-high-memory-usage/70266/11 "2017-02-13T16:48:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
