# Getting gc overhead on elastic search cluster on each instance

**URL:** <https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898>\
**Category:** Elasticsearch\
**Created:** [September 10, 2018, 7:11am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898 "2018-09-10T07:11:08Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![manish\_jaiswal](https://avatars.discourse-cdn.com/v4/letter/m/9de0a6/32.png) [@manish\_jaiswal](https://discuss.elastic.co/u/manish_jaiswal)\
**Post date:** [September 10, 2018, 7:11am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898/1 "2018-09-10T07:11:08Z")

</div>

I am getting gc overhead on elastic search cluster on each instance. (16data node 3master)

instance memory 60  
heap size ES :32  
queue size:7000  
thread pool:write

Indices: 1,541  
Primary Shards 7,557  
Replica Shards 7,557

please help .because of this too many data loss we are not able to see logs on kibana.

we directly sending data directly from file beat

in filebeat error:

2018-09-04T17:39:22.326+0530 INFO elasticsearch/client.go:690 Connected to Elasticsearch version 6.3.1  
2018-09-04T17:39:22.331+0530 INFO template/load.go:73 Template already exists and will not be overwritten.  
2018-09-04T17:39:22.331+0530 INFO [publish] pipeline/retry.go:172 retryer: send unwait-signal to consumer  
2018-09-04T17:39:22.331+0530 INFO [publish] pipeline/retry.go:174 done  
2018-09-04T17:39:22.341+0530 INFO [publish] pipeline/retry.go:149 retryer: send wait signal to consumer  
2018-09-04T17:39:22.341+0530 INFO [publish] pipeline/retry.go:151 done  
2018-09-04T17:39:22.973+0530 INFO [monitoring] log/log.go:124 Non-zero metrics in the last 30s {"monitoring": {"metrics": {"beat":{"cpu":{"system":{"ticks":150580,"time":{"ms":3}},"total":{"ticks":1106600,"time":{"ms":75},"value":1106600},"user":{"ticks":956020,"time":{"ms":72}}},"info":{"ephemeral\_id":"c0c77725-d7ec-4d04-9778-6c3e87caf483","uptime":{"ms":271440046}},"memstats":{"gc\_next":20433952,"memory\_alloc":18759592,"memory\_total":81088606696}},"filebeat":{"harvester":{"open\_files":10,"running":10}},"libbeat":{"config":{"module":{"running":0}},"output":{"events":{"batches":29,"failed":89,"total":89},"read":{"bytes":30628},"write":{"bytes":53426}},"pipeline":{"clients":96,"events":{"active":4126,"retry":178}}},"registrar":{"states":{"current":11}},"system":{"load":{"1":0.73,"15":0.21,"5":0.27,"norm":{"1":0.0913,"15":0.0263,"5":0.0338}}},"xpack":{"monitoring":{"pipeline":{"events":{"published":3,"total":3},"queue":{"acked":3}}}}}}}  
2018-09-04T17:39:23.341+0530 ERROR pipeline/output.go:92 Failed to publish events: temporary bulk send failure  
2018-09-04T17:39:23.341+0530 INFO [publish] pipeline/retry.go:172 retryer: send unwait-signal to consumer  
2018-09-04T17:39:23.341+0530 INFO [publish] pipeline/retry.go:174 done  
2018-09-04T17:39:23.341+0530 INFO [publish] pipeline/retry.go:149 retryer: send wait signal to consumer  
2018-09-04T17:39:23.341+0530 INFO [publish] pipeline/retry.go:151 done

please help

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 10, 2018, 7:13am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898/2 "2018-09-10T07:13:54Z")

</div>

What version are you on?  
What OS? What JVM?

Are you sure the heap is 32GB?

How much data, in GB, do those shards represent?

---

<div class="post-metadata">

**Author:** ![manish\_jaiswal](https://avatars.discourse-cdn.com/v4/letter/m/9de0a6/32.png) [@manish\_jaiswal](https://discuss.elastic.co/u/manish_jaiswal)\
**Post date:** [September 10, 2018, 7:42am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898/3 "2018-09-10T07:42:18Z")

</div>

ES version 6.3.1  
kibana version 6.3.1  
filebeat 6.3.1

es os:CentOS Linux 7

java version "1.8.0\_151"

data is not uniform some indices are of 1mb and some of 100gb.(least is 2.6kb to 500 gb max)

all data node have 32 gb heap.(master have less heap)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 10, 2018, 7:50am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898/4 "2018-09-10T07:50:19Z")

</div>

Having a lot of small shards can, [as described in this blog post](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster), be very inefficient. I would therefore recommend you try to reduce the shard count by changing your sharding strategy. It also seems like you may have been suffering from bulk rejections as you have dramatically increased the index queue size. This will most certainly not help with heap usage, and [reducing the number of shards you are actively indexing into can help with this too](https://www.elastic.co/blog/why-am-i-seeing-bulk-rejections-in-my-elasticsearch-cluster).

---

<div class="post-metadata">

**Author:** ![manish\_jaiswal](https://avatars.discourse-cdn.com/v4/letter/m/9de0a6/32.png) [@manish\_jaiswal](https://discuss.elastic.co/u/manish_jaiswal)\
**Post date:** [September 10, 2018, 10:06am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898/5 "2018-09-10T10:06:49Z")

</div>

thanks for quick help,

how much queue size should i keep for write thread pool.

what should be the shard size after reducing.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 10, 2018, 10:11am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898/6 "2018-09-10T10:11:02Z")

</div>

This typically depends on your use-case. A good shard size to aim for is somewhere between 10GB and 30GB, but can sometimes be slightly lower or even higher. The size of the bulk queue depends on how you index, but increasing it as much as you have done is often just applying a band-aid instead of addressing the underlying issue.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 8, 2018, 10:25am UTC](https://discuss.elastic.co/t/getting-gc-overhead-on-elastic-search-cluster-on-each-instance/147898/7 "2018-10-08T10:25:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
