# 100% CPU and GC on all nodes

**URL:** <https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485>\
**Category:** Elasticsearch\
**Created:** [July 20, 2019, 2:03pm UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485 "2019-07-20T14:03:33Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [July 20, 2019, 2:03pm UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485/1 "2019-07-20T14:03:33Z")

</div>

Hi guys,  
we recently did a rolling upgrade on our ES Cluster (40+ Nodes, 120+ indices) from 5.6.14 to 6.8.1.  
We had many issues which we could eventually fix.

One of the issues we still have however is, that the whole Cluster is at 100% CPU with the logs looking like this:

```
[2019-07-20T15:43:34,571][INFO][o.e.m.j.JvmGcMonitorService] [server1] [gc][289] overhead, spent [309ms] collecting in the last [1s]
[2019-07-20T15:43:38,308][INFO][o.e.m.j.JvmGcMonitorService] [server1] [gc][292] overhead, spent [580ms] collecting in the last [1.4s]
[2019-07-20T15:43:39,561][INFO][o.e.m.j.JvmGcMonitorService] [server1] [gc][293] overhead, spent [358ms] collecting in the last [1.2s]

```

...

Hot threads Output:  
[Pastebin](https://pastebin.com/5KgS1Zd8)

Also: We're running up to 4 nodes on one Server. Each node has 30GB JVM heap.

Any idea how to debug this properly?

Thanks! 😉

---

<div class="post-metadata">

**Author:** ![Bernt\_Rostad](https://avatars.discourse-cdn.com/v4/letter/b/3ab097/32.png) [@Bernt\_Rostad](https://discuss.elastic.co/u/Bernt_Rostad)\
**Post date:** [July 21, 2019, 8:55am UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485/2 "2019-07-21T08:55:24Z")

</div>

> [@xenoid](#):
>
> One of the issues we still have however is, that the whole Cluster is at 100% CPU with the logs looking like this:

I'm not sure what is the cause for your GC issue, but it seems clear that Elasticsearch is struggling to free memory it needs to operate. This could be because of query caches filling up quickly or that too many tasks are queued up (you could look for [rejected tasks](https://www.elastic.co/guide/en/elasticsearch/reference/current/cat-thread-pool.html), a sign of cluster overloading).

> [@xenoid](#):
>
> We're running up to 4 nodes on one Server. Each node has 30GB JVM heap.

As for running 4 nodes on one physical server you should take into account that Lucene uses file system caching to speed up queries, ideally Lucene should be able to read the most queried data from file system caches and only access the disk for less frequent data.

By assigning too much RAM to the JVMs you may starve the file system cache; if each of your 4 nodes are running JVMs with 30GB RAM you've effectively locked 120GB of the physical memory on that server which may not leave much for the file system cache, causing more of your queries to access the disk to fetch data.

You could try to experiment with the JVM sizes, reducing them to say 16GB or 8GB (which is what I'm using in my clusters) to see if that reduces the GC frequency and duration. Since GC only kicks in above a certain memory percentage, a smaller JVM will be garbage collected more often but also quicker than a big one - which may help in your case.

Good luck!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 21, 2019, 9:12am UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485/3 "2019-07-21T09:12:50Z")

</div>

What is the full output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-stats.html)? What is the hardware specification of the nodes this cluster is running on?

---

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [July 21, 2019, 10:07am UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485/4 "2019-07-21T10:07:20Z")

</div>

Over the last few hours the situation seems to have cleared up _slightly_.  
Only one node is still struggling and indexing only goes on bit for bit.  
Just as a test I closed the indices with active shard movement. This might have helped a bit?

> [@Bernt\_Rostad](#):
>
> This could be because of query caches filling up quickly or that too many tasks are queued up (you could look for [rejected tasks](https://www.elastic.co/guide/en/elasticsearch/reference/current/cat-thread-pool.html), a sign of cluster overloading).

My nodes seem to have rejected writes and searches all the way from 1000 to over 800000.  
This is the output of the node which still has high cpu load:

```
server2171-2 search 97 974 737559
server2171-2 write 7 0 8394
server2171-4 search 97 919 709452
server2171-4 write 9 2 5543
server2171-3 search 97 954 616195
server2171-3 write 23 0 7021
server2171 search 97 997 876564
server2171 write 7 0 9643

```

> [@Christian\_Dahlqvist](#):
>
> What is the hardware specification of the nodes this cluster is running on?

> [@Bernt\_Rostad](#):
>
> By assigning too much RAM to the JVMs you may starve the file system cache; if each of your 4 nodes are running JVMs with 30GB RAM you've effectively locked 120GB of the physical memory on that server which may not leave much for the file system cache, causing more of your queries to access the disk to fetch data.

As for the hardware specifications:  
CPU: 64T+  
RAM: (2 \* [Number of ES instances] \* 30GB) + 15GB+ extra headroom (In this case 280GB)  
Disk: ZFS volume on SSD or SAS RAID

> [@Bernt\_Rostad](#):
>
> You could try to experiment with the JVM sizes

I will try this tomorrow.

> [@Christian\_Dahlqvist](#):
>
> What is the full output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-stats.html)?

Here is an output:  
[Pastebin](https://pastebin.com/aGdDWqJ9)

And here is a screenshot of htop:

 ![htop_171](https://us1.discourse-cdn.com/elastic/original/3X/7/6/761abb0795962134185e6d95456e048b66948ebe.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 21, 2019, 10:52am UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485/5 "2019-07-21T10:52:28Z")

</div>

I have a couple of comments:

- It looks like you only have 2 master eligible nodes. This is bad given that Elasticsearch is based on consensus argorithms requiring a majority of master eligible nodes to be present to elect a master. You should therefore always have at least 3 master eligible nodes and [make sure minimum\_master\_nodes is set correctly](https://www.elastic.co/guide/en/elasticsearch/reference/6.8/modules-node.html#split-brain). As it is now it is likely you cluster either is not highly available or misconfigured, which can lead to data loss.
- It seems like a significant portion of your shards are not replicated and your cluster is currently in a red state.
- The nodes are using a variety of OS and JVM versions. Not sure what effect this may or may not have.
- It seems all nodes are configured as ingest nodes. Are you using ingest pipelines extensively or is this just the default setting?
- You are using SearchGuard to secure your cluster. I have never used this, so do not know to what effect this affects heap usage, GC and GC patterns.
- Given that you have different types of storage across the hosts, are you ru nning a hot-warm architecture? If you are - how is work distributed across them? Are the problems spread across all nodes?

---

<div class="post-metadata">

**Author:** ![xenoid](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xenoid/32/20672_2.png) [@xenoid](https://discuss.elastic.co/u/xenoid)\
**Post date:** [July 21, 2019, 12:11pm UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485/6 "2019-07-21T12:11:10Z")

</div>

This cluster was created in ES 1 and upgraded all the way to ES 6. That's why some things might not be optimal.

> [@Christian\_Dahlqvist](#):
>
> It seems like a significant portion of your shards are not replicated and your cluster is currently in a red state.

We don't use replicas for our data. (yet)  
The red state is caused by some nodes being offline (We can't run 6 nodes at once since upgrading to ES6 but this is a different issue.)  
The recent indices are green though.

> [@Christian\_Dahlqvist](#):
>
> It seems all nodes are configured as ingest nodes. Are you using ingest pipelines extensively or is this just the default setting?

This is the default setting and a relic of ES2

> [@Christian\_Dahlqvist](#):
>
> Given that you have different types of storage across the hosts, are you ru nning a hot-warm architecture? If you are - how is work distributed across them? Are the problems spread across all nodes?

We don't use the built in hot-warm architecture (yet).  
We have two tiers (high-performance/low-performance) and allocate indices by age.

EDIT:  
Here are some graphs from ES Monitoring:

 ![server2257](https://us1.discourse-cdn.com/elastic/original/3X/8/6/86d2f7694323d39862152a6a9bdec54376da06c6.jpeg)

EDIT2:  
If I **close every index except the current one** indexing works like a charm.  
Once I open **yesterdays index** everything goes to 100% again.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 18, 2019, 12:11pm UTC](https://discuss.elastic.co/t/100-cpu-and-gc-on-all-nodes/191485/7 "2019-08-18T12:11:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
