# Garbage collector question

**URL:** <https://discuss.elastic.co/t/garbage-collector-question/124776>\
**Category:** Elasticsearch\
**Created:** [March 20, 2018, 1:42pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776 "2018-03-20T13:42:19Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 20, 2018, 1:42pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/1 "2018-03-20T13:42:19Z")

</div>

Nodes: 3  
Indices: 142  
Memory: 17GB / 45GB  
Total Shards: 1284  
Documents: 442,160,568  
Data: 578GB  
Version: 5.3.0

Since a few days I am seeing a lot of these messages in the logfile:

[2018-03-20T14:35:04,269][INFO][o.e.m.j.JvmGcMonitorService] [7xHqegG] [gc][3661] overhead, spent [373ms] collecting in the last [1.1s]  
[2018-03-20T14:36:03,578][INFO][o.e.m.j.JvmGcMonitorService] [7xHqegG] [gc][3720] overhead, spent [286ms] collecting in the last [1s]  
[2018-03-20T14:36:05,608][INFO][o.e.m.j.JvmGcMonitorService] [7xHqegG] [gc][3722] overhead, spent [297ms] collecting in the last [1s]

These eventually turn into these:  
2018-03-20T13:06:20,875][ERROR][o.e.x.m.c.c.ClusterStatsCollector] [7xHqegG] collector [cluster-stats-collector] timed out when collecting data  
[2018-03-20T13:08:21,702][ERROR][o.e.x.m.c.c.ClusterStatsCollector] [7xHqegG] collector [cluster-stats-collector] timed out when collecting data  
[2018-03-20T13:09:18,651][ERROR][o.e.x.m.c.c.ClusterStatsCollector] [7xHqegG] collector [cluster-stats-collector] timed out when collecting data  
[2018-03-20T13:10:30,348][ERROR][o.e.x.m.c.i.IndexStatsCollector] [7xHqegG] collector [index-stats-collector] timed out when collecting data  
[2018-03-20T13:13:32,115][ERROR][o.e.x.m.c.c.ClusterStatsCollector] [7xHqegG] collector [cluster-stats-collector] timed out when collecting data  
[2018-03-20T13:31:10,813][ERROR][o.e.x.m.c.c.ClusterStatsCollector] [7xHqegG] collector [cluster-stats-collector] timed out when collecting data  
[2018-03-20T13:51:23,536][ERROR][o.e.x.m.c.c.ClusterStatsCollector] [7xHqegG] collector [cluster-stats-collector] timed out when collecting data

I read a suggestion on the forum, that increasing heap size would solve this .. (which I did from 30GB to 45GB) (15GB per server)

But I am still seeing the errors occur.

Any ideas on how to debug/resolve this?

thanks,

---

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 20, 2018, 5:05pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/2 "2018-03-20T17:05:28Z")

</div>

I have also tried to reduce heap size, but I am still seeing the gc overhead messages.

---

<div class="post-metadata">

**Author:** ![loren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loren/32/44942_2.png) [@loren](https://discuss.elastic.co/u/loren)\
**Post date:** [March 20, 2018, 6:12pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/3 "2018-03-20T18:12:05Z")

</div>

A 30GB heap is already incredibly large. The ~1300 shards jumps out at me first. That is a big number especially for so little data. I would try to reduce the number of indices and the number of shards per index by 10x and see if your problem goes away.

---

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 20, 2018, 7:17pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/4 "2018-03-20T19:17:09Z")

</div>

Thanks for the reply.

The indices are just daily logstash and metricbeat - so I don't I think I can reduce indices.

But I guess I can reduce shards.

---

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 20, 2018, 7:52pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/5 "2018-03-20T19:52:45Z")

</div>

I've bumped shard allocation to 2 for logstash for now.

(Average logstash index is about 12GB)

---

<div class="post-metadata">

**Author:** ![loren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/loren/32/44942_2.png) [@loren](https://discuss.elastic.co/u/loren)\
**Post date:** [March 20, 2018, 9:16pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/6 "2018-03-20T21:16:15Z")

</div>

Have you looked at using size-based [Rollover Indices](https://www.elastic.co/guide/en/elasticsearch/reference/master/indices-rollover-index.html) instead of time-based indices? Dividing so little data across so many indices and shards is like cutting birthday cake into slices: the more cuts you make, the more cake ends up on the knife.

---

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 21, 2018, 10:51am UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/7 "2018-03-21T10:51:18Z")

</div>

I hadn’t. But I am now 🙂

However, I thought ~12GB indices were decently sized? The rollover API doc also references 5GB turnover. That might just be a bad example though

Thanks

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 21, 2018, 11:01am UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/8 "2018-03-21T11:01:29Z")

</div>

In most discussions about appropriate index and shard size, it is the shard size that is most interesting and used. An average shard size of 12Gb sounds quite reasonable, but based on the data you provided, you seem to have an average shard size of less than 500MB, which is small.

If you are using the rollover index API and target a certain shard size, you naturally need to consider how many shards it will have when you make your calculation.

---

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 21, 2018, 12:01pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/9 "2018-03-21T12:01:49Z")

</div>

Indeed. I've cleaned up the indices.

Now you are talking about shard sizes, and I was talking about index size.

So average logstash index is about 12GB. And with the modifications I made yesterday - this index is allocated 2 shards and 1 replicate.

Does that sound reasonable?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 21, 2018, 12:17pm UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/10 "2018-03-21T12:17:29Z")

</div>

It you had an average shard size of 6GB, you would have around 100 shards instead of 1300, which sounds much more reasonable for that data volume.

---

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 22, 2018, 7:59am UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/11 "2018-03-22T07:59:15Z")

</div>

Do you think 10GB heap per node is too much?

Cause I am still seeing the “[INFO][o.e.m.j.JvmGcMonitorService] [7xHqegG] [gc][3722] overhead, spent [297ms] collecting in the last [1s]”

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 22, 2018, 8:04am UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/12 "2018-03-22T08:04:56Z")

</div>

That depends on the workload. Once you reduce the number of shards you should be able to install monitoring and see if it is too much from looking at heap usage over time.

---

<div class="post-metadata">

**Author:** ![michbsd](https://avatars.discourse-cdn.com/v4/letter/m/b77776/32.png) [@michbsd](https://discuss.elastic.co/u/michbsd)\
**Post date:** [March 22, 2018, 10:18am UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/13 "2018-03-22T10:18:43Z")

</div>

I already have monitoring installed.

The only noticable thing are peak waves on the heap usage :

 ![skitch%20image](https://us1.discourse-cdn.com/elastic/original/3X/d/f/df1cab9263abee95a28ba524d361851235357f05.jpg)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 19, 2018, 10:19am UTC](https://discuss.elastic.co/t/garbage-collector-question/124776/14 "2018-04-19T10:19:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
