# Garbage Collection in ES

**URL:** <https://discuss.elastic.co/t/garbage-collection-in-es/6736>\
**Category:** Elasticsearch\
**Created:** [February 17, 2012, 10:20am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736 "2012-02-17T10:20:54Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Brad\_Lhotsky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/brad_lhotsky/32/2293_2.png) [@Brad\_Lhotsky](https://discuss.elastic.co/u/Brad_Lhotsky)\
**Post date:** [February 17, 2012, 10:20am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/1 "2012-02-17T10:20:54Z")

</div>

Hello,

We're seeing some issues with timeouts lasting 6 seconds in  
ElasticSearch during what appear to be massive Garbage Collection  
sessions. We have 4 nodes, and 7 shards, the index is only 4.3G  
loaded into memory. We've allocated 10gigs of memory on each server  
to ES and have followed the recommendations to:

elasticsearch.yml:  
bootstrap.mlockall: true

ENV:  
ES\_MIN\_MEM=10g  
ES\_MAX\_MEM=10g

This is a CentOS 5 box, and I've set the /proc/sys/vm/swappiness to 0.  
On one box I have completely disabled swap, but as a System  
Administrator, this makes me really nervous. Disabling virtual memory  
is not a production solution and doesn't appear to have fixed the  
issue anyways.

I hacked together a Perl script to pull data from the ElasticSearch  
nodes and output it to graphite or cacti. It's attached as  
perf\_elastic\_search.pl.

I'm seeing some interesting behavior from the garbage collection at  
the point this timeout occurs. The Graphlot-2h.png shows the garbage  
collection time\_ms spiking as the heap.used\_bytes drops significantly.  
This seems to be a pattern, see Graphlot-24h.png.

It seems once the node has more than 8gb in heap.used\_bytes, it  
garbage collects itself down to 2gb. However, the GC between these  
points is relatively unconcerned by the expanding heap.

Is there a way to force the garbage collection to favor smaller, more  
frequent GC than to simply do it all at once every 2-3 hours? This  
Garbage collection results in nodes timing out their connections for 6  
seconds every 2-3 hours.

I'm not a Java guy, so any nudges in the right direction would be appreciated.

Thanks,

--  
Brad Lhotsky

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 17, 2012, 10:27am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/2 "2012-02-17T10:27:41Z")

</div>

Hi Brad

> We're seeing some issues with timeouts lasting 6 seconds in  
> Elasticsearch during what appear to be massive Garbage Collection  
> sessions. We have 4 nodes, and 7 shards, the index is only 4.3G  
> loaded into memory. We've allocated 10gigs of memory on each server  
> to ES and have followed the recommendations to:
> 
> elasticsearch.yml:  
> bootstrap.mlockall: true
> 
> ENV:  
> ES\_MIN\_MEM=10g  
> ES\_MAX\_MEM=10g

One question - are you sure that you have "ulimit -l unlimited" set for  
the user running elasticsearch? If not, then the mlockall will be  
ignored.

eg I have this in my /etc/security/limits.conf:

elasticsearch - nofile 32000  
elasticsearch - memlock unlimited

but I know that, at least on ubuntu, this file is ignored by default.  
Don't know what the situation is on Centos.

clint

---

<div class="post-metadata">

**Author:** ![Brad\_Lhotsky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/brad_lhotsky/32/2293_2.png) [@Brad\_Lhotsky](https://discuss.elastic.co/u/Brad_Lhotsky)\
**Post date:** [February 17, 2012, 10:41am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/3 "2012-02-17T10:41:59Z")

</div>

Sorry, forgot to include that, yes, I have this in /etc/security/limits.conf:  
elasticsearch - memlock unlimited

Have you graphed the jvm stats? Do you see the same massive GC  
Sessions every few hours?

On Fri, Feb 17, 2012 at 11:27 AM, Clinton Gormley [clint@traveljury.com](mailto:clint@traveljury.com) wrote:

> Hi Brad
> 
> > We're seeing some issues with timeouts lasting 6 seconds in  
> > Elasticsearch during what appear to be massive Garbage Collection  
> > sessions. We have 4 nodes, and 7 shards, the index is only 4.3G  
> > loaded into memory. We've allocated 10gigs of memory on each server  
> > to ES and have followed the recommendations to:
> > 
> > elasticsearch.yml:  
> > bootstrap.mlockall: true
> > 
> > ENV:  
> > ES\_MIN\_MEM=10g  
> > ES\_MAX\_MEM=10g
> 
> One question - are you sure that you have "ulimit -l unlimited" set for  
> the user running elasticsearch? If not, then the mlockall will be  
> ignored.
> 
> eg I have this in my /etc/security/limits.conf:
> 
> elasticsearch - nofile 32000  
> elasticsearch - memlock unlimited
> 
> but I know that, at least on ubuntu, this file is ignored by default.  
> Don't know what the situation is on Centos.
> 
> clint

--  
Brad Lhotsky

---

<div class="post-metadata">

**Author:** ![Danil\_TORIN](https://avatars.discourse-cdn.com/v4/letter/d/dec6dc/32.png) [@Danil\_TORIN](https://discuss.elastic.co/u/Danil_TORIN)\
**Post date:** [February 17, 2012, 10:58am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/4 "2012-02-17T10:58:02Z")

</div>

Try to reduce the memory.

If after GC your memory usage is ~2GB, set your limits to 3Gb.

On Fri, Feb 17, 2012 at 12:41, Brad Lhotsky [brad.lhotsky@gmail.com](mailto:brad.lhotsky@gmail.com) wrote:

> Sorry, forgot to include that, yes, I have this in /etc/security/limits.conf:  
> elasticsearch - memlock unlimited
> 
> Have you graphed the jvm stats? Do you see the same massive GC  
> Sessions every few hours?
> 
> On Fri, Feb 17, 2012 at 11:27 AM, Clinton Gormley [clint@traveljury.com](mailto:clint@traveljury.com) wrote:
> 
> > Hi Brad
> > 
> > > We're seeing some issues with timeouts lasting 6 seconds in  
> > > Elasticsearch during what appear to be massive Garbage Collection  
> > > sessions. We have 4 nodes, and 7 shards, the index is only 4.3G  
> > > loaded into memory. We've allocated 10gigs of memory on each server  
> > > to ES and have followed the recommendations to:
> > > 
> > > elasticsearch.yml:  
> > > bootstrap.mlockall: true
> > > 
> > > ENV:  
> > > ES\_MIN\_MEM=10g  
> > > ES\_MAX\_MEM=10g
> > 
> > One question - are you sure that you have "ulimit -l unlimited" set for  
> > the user running elasticsearch? If not, then the mlockall will be  
> > ignored.
> > 
> > eg I have this in my /etc/security/limits.conf:
> > 
> > elasticsearch - nofile 32000  
> > elasticsearch - memlock unlimited
> > 
> > but I know that, at least on ubuntu, this file is ignored by default.  
> > Don't know what the situation is on Centos.
> > 
> > clint
> 
> --  
> Brad Lhotsky

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 17, 2012, 11:19am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/5 "2012-02-17T11:19:04Z")

</div>

Hi Brad

On Fri, 2012-02-17 at 11:41 +0100, Brad Lhotsky wrote:

> Sorry, forgot to include that, yes, I have this in /etc/security/limits.conf:  
> elasticsearch - memlock unlimited

Sorry, just to repeat: are you sure that it is being applied? When you  
start ES, does it immediately take up the full amount specified in  
ES\_MAX\_MEM?

> Have you graphed the jvm stats? Do you see the same massive GC  
> Sessions every few hours?

I haven't done so for a long time, but yes, I used to see those big GCs,  
but they are fast enough so that they don't cause a problem. We're  
currently running 0.17.9 in production, so if you're on a more recent  
version, behaviour might have changed.

I'll use bigdesk to capture some data today, and come back to you

clint

---

<div class="post-metadata">

**Author:** ![Aurelien\_2](https://avatars.discourse-cdn.com/v4/letter/a/a587f6/32.png) [@Aurelien\_2](https://discuss.elastic.co/u/Aurelien_2)\
**Post date:** [February 17, 2012, 11:25am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/6 "2012-02-17T11:25:02Z")

</div>

Hello.

this is a quite huge heap size for a JVM.

Gil Tene as explained the problem quite well: [Understanding Java Garbage Collection and What You Can Do about It](http://www.infoq.com/presentations/Understanding-Java-Garbage-Collection) Ok, he's Azul co-founder so it's a kind of advertising for Azul, but the presentation is still very interesting.

So, I think Danil is right: reduce the heap size. Or give a try to Azul 🙂

regards.  
----- Mail original -----

> De: "Brad Lhotsky" [brad.lhotsky@gmail.com](mailto:brad.lhotsky@gmail.com)  
> À: [elasticsearch@googlegroups.com](mailto:elasticsearch@googlegroups.com)  
> Envoyé: Vendredi 17 Février 2012 11:20:54  
> Objet: Garbage Collection in ES

> Hello,

> We're seeing some issues with timeouts lasting 6 seconds in  
> Elasticsearch during what appear to be massive Garbage Collection  
> sessions. We have 4 nodes, and 7 shards, the index is only 4.3G  
> loaded into memory. We've allocated 10gigs of memory on each server  
> to ES and have followed the recommendations to:

> elasticsearch.yml:  
> bootstrap.mlockall: true

> ENV:  
> ES\_MIN\_MEM=10g  
> ES\_MAX\_MEM=10g

> This is a CentOS 5 box, and I've set the /proc/sys/vm/swappiness to  
> 0.  
> On one box I have completely disabled swap, but as a System  
> Administrator, this makes me really nervous. Disabling virtual memory  
> is not a production solution and doesn't appear to have fixed the  
> issue anyways.

> I hacked together a Perl script to pull data from the Elasticsearch  
> nodes and output it to graphite or cacti. It's attached as  
> perf\_elastic\_search.pl.

> I'm seeing some interesting behavior from the garbage collection at  
> the point this timeout occurs. The Graphlot-2h.png shows the garbage  
> collection time\_ms spiking as the heap.used\_bytes drops  
> significantly.  
> This seems to be a pattern, see Graphlot-24h.png.

> It seems once the node has more than 8gb in heap.used\_bytes, it  
> garbage collects itself down to 2gb. However, the GC between these  
> points is relatively unconcerned by the expanding heap.

> Is there a way to force the garbage collection to favor smaller, more  
> frequent GC than to simply do it all at once every 2-3 hours? This  
> Garbage collection results in nodes timing out their connections for  
> 6  
> seconds every 2-3 hours.

> I'm not a Java guy, so any nudges in the right direction would be  
> appreciated.

> Thanks,

> --  
> Brad Lhotsky

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 17, 2012, 12:25pm UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/7 "2012-02-17T12:25:23Z")

</div>

Hi Brad

> I'll use bigdesk to capture some data today, and come back to you

Attached is a screenshot of bigdesk from the past hour or so

(again, version 0.17.9)

clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 17, 2012, 6:58pm UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/8 "2012-02-17T18:58:05Z")

</div>

First, there is no problem with this size of heap (as mentioned by someone on this thread), people are running ES with 30gb happily. You mentioned you store the index in memory? Whats the config that you use?

On Friday, February 17, 2012 at 12:20 PM, Brad Lhotsky wrote:

> Hello,
> 
> We're seeing some issues with timeouts lasting 6 seconds in  
> Elasticsearch during what appear to be massive Garbage Collection  
> sessions. We have 4 nodes, and 7 shards, the index is only 4.3G  
> loaded into memory. We've allocated 10gigs of memory on each server  
> to ES and have followed the recommendations to:
> 
> elasticsearch.yml:  
> bootstrap.mlockall: true
> 
> ENV:  
> ES\_MIN\_MEM=10g  
> ES\_MAX\_MEM=10g
> 
> This is a CentOS 5 box, and I've set the /proc/sys/vm/swappiness to 0.  
> On one box I have completely disabled swap, but as a System  
> Administrator, this makes me really nervous. Disabling virtual memory  
> is not a production solution and doesn't appear to have fixed the  
> issue anyways.
> 
> I hacked together a Perl script to pull data from the Elasticsearch  
> nodes and output it to graphite or cacti. It's attached as  
> perf\_elastic\_search.pl (http://perf\_elastic\_search.pl).
> 
> I'm seeing some interesting behavior from the garbage collection at  
> the point this timeout occurs. The Graphlot-2h.png shows the garbage  
> collection time\_ms spiking as the heap.used\_bytes drops significantly.  
> This seems to be a pattern, see Graphlot-24h.png.
> 
> It seems once the node has more than 8gb in heap.used\_bytes, it  
> garbage collects itself down to 2gb. However, the GC between these  
> points is relatively unconcerned by the expanding heap.
> 
> Is there a way to force the garbage collection to favor smaller, more  
> frequent GC than to simply do it all at once every 2-3 hours? This  
> Garbage collection results in nodes timing out their connections for 6  
> seconds every 2-3 hours.
> 
> I'm not a Java guy, so any nudges in the right direction would be appreciated.
> 
> Thanks,
> 
> --  
> Brad Lhotsky
> 
> Attachments:
> 
> - perf\_elastic\_search.pl
> 
> - Graphlot-2h.png
> 
> - Graphlot-24h.png

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:38am UTC](https://discuss.elastic.co/t/garbage-collection-in-es/6736/9 "2017-07-06T03:38:54Z")

</div>


