# Promotion failures (GC issues)

**URL:** <https://discuss.elastic.co/t/promotion-failures-gc-issues/15851>\
**Category:** Elasticsearch\
**Created:** [February 17, 2014, 3:34pm UTC](https://discuss.elastic.co/t/promotion-failures-gc-issues/15851 "2014-02-17T15:34:29Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![nicolas\_long](https://avatars.discourse-cdn.com/v4/letter/n/43a26b/32.png) [@nicolas\_long](https://discuss.elastic.co/u/nicolas_long)\
**Post date:** [February 17, 2014, 3:34pm UTC](https://discuss.elastic.co/t/promotion-failures-gc-issues/15851/1 "2014-02-17T15:34:29Z")

</div>

Hey all,

we regularly (several times a week) get longish GCs (20s or more) due to  
promotion failures.

From what I understand this type of major GC is caused by fragmentation of  
the heap.

So I'm wondering:

1. What is all the stuff ES puts into the heap that ends up in the Old Gen?
2. Are there any recommended strategies for dealing with this specific kind  
of problem.

For example, would allowing more filter caching help or cause even more  
problems? And so on.

To give a little more info on our usage, we're read heavy, nearly entirely  
filter operations. Our heap is at ~10g. nearly all of which is used by the  
Old Gen (until a major GC runs).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4b3ef926-94b7-4de0-b076-d5fdbc44021c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4b3ef926-94b7-4de0-b076-d5fdbc44021c%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 17, 2014, 8:50pm UTC](https://discuss.elastic.co/t/promotion-failures-gc-issues/15851/2 "2014-02-17T20:50:30Z")

</div>

Maybe it's the field cache that moves to old gen, when using facets.

I am tackling this challenge by a combination of several strategies

- tuning index.indices.fielddata.cache.size

- working around the issue by increasing node transport and ping timeout  
from 5s to something high like 30s (so GCs are allowed to run 20s without  
node disconnects)

- reducing number of shards per node (this just means to reduce the number  
of docs / index size / filter cache per node somehow), simplest method is  
adding nodes

- using heap sizes as small as possible - in my use case 6G are sufficient

- not sure if you want to go the path on the bleeding edge, but using Java  
8 and G1GC with XX:MaxGCPauseMillis of ~100-1000ms helps me. CPU load is a  
bit higher with G1GC, but since I have 32 cores on a node, it does not  
matter that much.

- otherwise, there are lots of CMS GC tuning options (needs deep GC  
analysis)

Jörg

On Mon, Feb 17, 2014 at 4:34 PM, Nic Long [nicolas.long@guardian.co.uk](mailto:nicolas.long@guardian.co.uk)wrote:

> Hey all,
> 
> we regularly (several times a week) get longish GCs (20s or more) due to  
> promotion failures.
> 
> From what I understand this type of major GC is caused by fragmentation of  
> the heap.
> 
> So I'm wondering:
> 
> 1. What is all the stuff ES puts into the heap that ends up in the Old Gen?
> 2. Are there any recommended strategies for dealing with this specific  
> kind of problem.
> 
> For example, would allowing more filter caching help or cause even more  
> problems? And so on.
> 
> To give a little more info on our usage, we're read heavy, nearly entirely  
> filter operations. Our heap is at ~10g. nearly all of which is used by the  
> Old Gen (until a major GC runs).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/4b3ef926-94b7-4de0-b076-d5fdbc44021c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4b3ef926-94b7-4de0-b076-d5fdbc44021c%40googlegroups.com)  
> .  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoG7nX7YRy7dEnfDToWaPXvVTjfwP%3DXYdPzjRk91YJ0d%2BA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoG7nX7YRy7dEnfDToWaPXvVTjfwP%3DXYdPzjRk91YJ0d%2BA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![nicolas\_long](https://avatars.discourse-cdn.com/v4/letter/n/43a26b/32.png) [@nicolas\_long](https://discuss.elastic.co/u/nicolas_long)\
**Post date:** [February 18, 2014, 3:02pm UTC](https://discuss.elastic.co/t/promotion-failures-gc-issues/15851/3 "2014-02-18T15:02:15Z")

</div>

Hey Jörg,

thanks for the detailed reply.

We don't really run facets and our field data cache size is very small.  
Increasing the node transport and ping timeouts is definitely something  
we'll consider. Reducing the number of shards per node is also something to  
consider, but am reluctant to add more nodes at the moment (already  
spending lots of cash).

I think a deep dive into GC tuning is possibly called for, and we've done  
some of that already.

Java 8 with G1GC is an interesting suggestion too!

Thanks again,

Nic

On Monday, 17 February 2014 20:50:30 UTC, Jörg Prante wrote:

> Maybe it's the field cache that moves to old gen, when using facets.
> 
> I am tackling this challenge by a combination of several strategies
> 
> - tuning index.indices.fielddata.cache.size
> 
> - working around the issue by increasing node transport and ping timeout  
> from 5s to something high like 30s (so GCs are allowed to run 20s without  
> node disconnects)
> 
> - reducing number of shards per node (this just means to reduce the number  
> of docs / index size / filter cache per node somehow), simplest method is  
> adding nodes
> 
> - using heap sizes as small as possible - in my use case 6G are sufficient
> 
> - not sure if you want to go the path on the bleeding edge, but using Java  
> 8 and G1GC with XX:MaxGCPauseMillis of ~100-1000ms helps me. CPU load is a  
> bit higher with G1GC, but since I have 32 cores on a node, it does not  
> matter that much.
> 
> - otherwise, there are lots of CMS GC tuning options (needs deep GC  
> analysis)
> 
> Jörg
> 
> On Mon, Feb 17, 2014 at 4:34 PM, Nic Long \<[nicola...@guardian.co.uk](mailto:nicola...@guardian.co.uk)\<javascript:\>
> 
> > wrote:
> 
> > Hey all,
> > 
> > we regularly (several times a week) get longish GCs (20s or more) due to  
> > promotion failures.
> > 
> > From what I understand this type of major GC is caused by fragmentation  
> > of the heap.
> > 
> > So I'm wondering:
> > 
> > 1. What is all the stuff ES puts into the heap that ends up in the Old  
> > Gen?
> > 2. Are there any recommended strategies for dealing with this specific  
> > kind of problem.
> > 
> > For example, would allowing more filter caching help or cause even more  
> > problems? And so on.
> > 
> > To give a little more info on our usage, we're read heavy, nearly  
> > entirely filter operations. Our heap is at ~10g. nearly all of which is  
> > used by the Old Gen (until a major GC runs).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/4b3ef926-94b7-4de0-b076-d5fdbc44021c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4b3ef926-94b7-4de0-b076-d5fdbc44021c%40googlegroups.com)  
> > .  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/d8803723-0cec-40ba-a095-4fe73f123e75%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/d8803723-0cec-40ba-a095-4fe73f123e75%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:49am UTC](https://discuss.elastic.co/t/promotion-failures-gc-issues/15851/4 "2017-07-06T01:49:26Z")

</div>


