# Frequency old gc of some nodes in cluster

**URL:** <https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286>\
**Category:** Elasticsearch\
**Created:** [May 9, 2019, 6:07am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286 "2019-05-09T06:07:32Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 9, 2019, 6:07am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/1 "2019-05-09T06:07:32Z")

</div>

hi  
During busy hours of Day, heap usage of some of nodes in Es cluster began to raise and old gc was more frequency. Full gc was also triggered and last about 20-40 seconds.

old gc times:

![image](https://us1.discourse-cdn.com/elastic/original/3X/e/d/ed99476748f3eb299b262bb363be63c318dd335e.png)

Only If I reduce index rate, the heap usage was back to normal. I made a dump of heap and used MAT to do Leak Suspects.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/6/f6b377a7574f6139268ddb190d02ab684257d2cd.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/7/6728cc64e6c233e9c7c9aea53d3cb6ff711a5f04.png)

Es Cluster Info:  
Es version 5.4.3  
3 \* master node, 26 hot node, 52 code node.  
Only one hot node and 3 code node meet gc frequency problem at the same time.

Is there any anomaly according to the Leak Suspects reports? I will provide more info if needed.

Thanks.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2019, 6:30am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/2 "2019-05-09T06:30:54Z")

</div>

Hey.

Not sure it will solve your current problem but what about upgrading to the latest 5.x version which contains a lot of bug fixes?  
Even better, upgrade to 6.x or better than better, upgrade to 7.0?

What is the output of:

```auto
GET /_cat/health?v
GET /_cat/indices?v
GET /_cat/shards?v

```

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 9, 2019, 6:48am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/3 "2019-05-09T06:48:25Z")

</div>

Thanks @dadoonet  
During problem time, this cluster is in Green state and also there is no shard relocation.

Doing elasticsearch version update is a huge task for us at this moment. We have 6 Es clusters and use tribe node as proxy. So we have to update all clusters.

Only one cluster meet gc problem recently and the index load of this cluster is not very high in my opinion.

Do you think that it is some bug of 5.4.3 ES caused this gc problem?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2019, 7:23am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/4 "2019-05-09T07:23:02Z")

</div>

But could you answer the questions I asked?

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 9, 2019, 7:28am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/5 "2019-05-09T07:28:25Z")

</div>

I will list the output of these requests next time when gc problem happen.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 9, 2019, 7:37am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/6 "2019-05-09T07:37:11Z")

</div>

Is this issue related to the same cluster discussed in [this thread](https://discuss.elastic.co/t/lots-of-disconnected-logs-and-then-oom/177758)? If so, how much data do you have in the cluster? Have you followed the guidelines laid out in [this webinar](https://www.elastic.co/webinars/optimizing-storage-efficiency-in-elasticsearch)? It also seems like you have a quite high index and shard count, which could be contributing to heap pressure. Please see [this blog post](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) for some practical guidelines.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2019, 7:41am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/7 "2019-05-09T07:41:37Z")

</div>

Please do it now. No need to wait.

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 9, 2019, 7:48am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/8 "2019-05-09T07:48:03Z")

</div>

This issus is not related to [this thread](https://discuss.elastic.co/t/lots-of-disconnected-logs-and-then-oom/177758)  
The ES cluster INFO:  
Es version 5.4.3  
3 \* master node, 26 \* hot node, 52 \* code node.  
each node has 31GB heap.  
2,910 indices 6,518 shards

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2019, 7:49am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/9 "2019-05-09T07:49:46Z")

</div>

Is it different than your initial question?

> Es Cluster Info:  
> Es version 5.4.3  
> 3 \* master node, 20 hot node, 40 code node.

Also in which node the GC is happening?

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 9, 2019, 7:52am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/10 "2019-05-09T07:52:51Z")

</div>

GET /\_cat/health?v

> <https://gist.github.com/wangxiangyu/32a3ac9de2496923782a31bf982d984f>

GET /\_cat/indices?v

> <https://gist.github.com/wangxiangyu/58d309b3ae0e7e106ab35ba2e240b341>

GET /\_cat/shards?v

> <https://gist.github.com/wangxiangyu/c05df65e762b86095c5f11e176bf8a41>

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 9, 2019, 7:58am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/11 "2019-05-09T07:58:30Z")

</div>

I make a mistake in initial question

NOTE: Date is only index to hot node. The data is moved from hot node to cold node daily.

The nodes with gc problems are:  
jssz-billions-es-40-datanode\_hot  
jssz-billions-es-22-datanode\_stale  
jssz-billions-es-39-datanode\_stale01  
jssz-billions-es-48-datanode\_stale  
jssz-billions-es-26-datanode\_stale01  
jssz-billions-es-24-datanode\_stale

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2019, 8:32am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/12 "2019-05-09T08:32:08Z")

</div>

Thanks for sharing.  
It looks like you have plenty of small shards. May be something you should consider.  
Some shards are overloaded IMO. Like

```auto
billions-video.vod.playurl-@2019.05.09-jssz03-1	1	p	STARTED	89054241	131.2gb	10.69.175.31	jssz-billions-es-55-datanode_hot
billions-video.vod.playurl-@2019.05.09-jssz03-1	0	p	STARTED	89083881	131.2gb	10.69.175.32	jssz-billions-es-56-datanode_hot
billions-video.vod.playurl-@2019.05.09-jssz03-1	5	p	STARTED	88789888	131.3gb	10.69.67.14	jssz-billions-es-19-datanode_hot
billions-video.vod.playurl-@2019.05.09-jssz03-1	2	p	STARTED	88941625	133.4gb	10.69.34.17	jssz-billions-es-40-datanode_hot
billions-video.vod.playurl-@2019.05.09-jssz03-1	9	p	STARTED	88851385	133.5gb	10.69.67.20	jssz-billions-es-28-datanode_hot
billions-video.vod.playurl-@2019.05.09-jssz03-1	8	p	STARTED	89018956	133.8gb	10.69.67.18	jssz-billions-es-26-datanode_hot

```

We recommend no more than 50gb per shard.

Not sure if you are using [rollover API](https://www.elastic.co/guide/en/elasticsearch/reference/7.0/indices-rollover-index.html) but I'd use it in your case to reduce the number of shards and try to keep them around 50gb per shard.

The total number of shards/indices you have in your cluster has also the consequence I think of a very big cluster state. Those big "objects" needs to be Gc'ed sometime. Because you have a very big HEAP (31gb), the old GC can take several minutes sadly.

My opinion is that you should consider at some point to upgrade elasticsearch and your JVM. In 7.x you will have a more recent JVM which different GC algorithms.

But I'll be happy to hear other thoughts. 🙂

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 9, 2019, 9:19am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/13 "2019-05-09T09:19:41Z")

</div>

Thanks for your reply.  
I think "big cluster state" maybe is not the case of gc problem. I have another cluster ( B for short)which is the same size(hardware size) as this cluster( A for short) with gc problem. Cluster B has 13,508 indices and 21,615 shards as much as twice of cluster A. Index load of Cluster B is also larger then Cluster A. But cluster B never met the gc problem.

According to the leak suspect report, this suspect is very suspicious.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/3/3374946dcdc8ae95e84e2a62c502213fb3cb4ce4.png)

what's your opinion?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 9, 2019, 11:28am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/14 "2019-05-09T11:28:07Z")

</div>

What is the output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-stats.html)?

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 10, 2019, 8:24am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/15 "2019-05-10T08:24:00Z")

</div>

> <https://gist.github.com/wangxiangyu/ceae49818d29d6ae67fcd11f3aa550f7>

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 10, 2019, 8:53am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/16 "2019-05-10T08:53:11Z")

</div>

It looks like you have a 3rd party SQL plugin installed. Do all environments have this? Is usage of this plugin consistent across the environments?

---

<div class="post-metadata">

**Author:** ![shjdwxy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shjdwxy/32/43102_2.png) [@shjdwxy](https://discuss.elastic.co/u/shjdwxy)\
**Post date:** [May 10, 2019, 10:02am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/17 "2019-05-10T10:02:08Z")

</div>

I will try to remove sql plugin, it is useless now.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 7, 2019, 10:02am UTC](https://discuss.elastic.co/t/frequency-old-gc-of-some-nodes-in-cluster/180286/18 "2019-06-07T10:02:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
