# Poor performance and a lot of GC overhead

**URL:** <https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145>\
**Category:** Elasticsearch\
**Created:** [July 16, 2018, 12:01pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145 "2018-07-16T12:01:28Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [July 16, 2018, 12:01pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/1 "2018-07-16T12:01:28Z")

</div>

Hi,

we are running elasticsearch with the following spec:

1 node on a hardware server  
16GB of RAM  
2GB assigned for heap  
8 CPU cores (Intel Xeon CPU E5620 @ 2.40GHz)  
HDD disks  
Elasticsearch 5.6.3  
JVM 1.8.0\_171

we have around 20 indices, some daily, some weekly, some monthly.  
each index has the default 5 shards and 1 replica.  
each index gets might get around 1M documents per day at most.  
there is indexing operation that is happening constantly, and every minute or so, there are also search operations that are happening.

the spec described above is what we have to work with 🙂 and I know it sucks... but nothing much to do there.  
the reason that there are only 2GB allocated for heap is that there are other components running on this machine and they also need resources...

the immediate issue we can pinpoint is that we are seeing a lot of GC overhead alerts in elasticsearch logs such as these:

> [gc][999] overhead, spent [842ms] collecting in the last [1.4s]

so my questions are as follows:

1. what might be the cause for those GC overhead? and are they an issue to worry about (seems to me that it is an issue...)
2. what can we do to properly measure the actual capacity of this server? I mean to understand what is the performance I can get expect from elasticsearch given those constraints?
3. what would be a better sharding strategy given the above spec (that we have only 1 server in this setup).

thanks in advance!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 16, 2018, 12:12pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/2 "2018-07-16T12:12:05Z")

</div>

It sounds like you may be having a too many shards, especially given the small heap size. Read [this blog post](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster) for a discussion around shards and sharding practices.

Until you can revise your sharding strategy, I would recommend increasing the heap size.

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [July 16, 2018, 12:14pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/3 "2018-07-16T12:14:56Z")

</div>

thanks, so I was actually thinking to reduce the shard amount to 1 primary 0 replicas. since this is only a single node, so I don't see a point in having more than a single shard on a single node. but is that true or am I missing something?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 16, 2018, 12:16pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/4 "2018-07-16T12:16:25Z")

</div>

With single node there is no point in having replicas configured as Elasticsearch will not be able to allocate them anyway. How many shards do you have in the cluster? What is the average shard size?

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [July 16, 2018, 12:17pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/5 "2018-07-16T12:17:35Z")

</div>

as I mentioned above, I have the default 5 shards per index. but the indices are rather small. none of them go beyond 1 GB in size...

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 16, 2018, 12:19pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/6 "2018-07-16T12:19:02Z")

</div>

That is wasteful, but you may be having too many indices as well. What is the output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/6.3/cluster-stats.html)?

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [July 16, 2018, 12:33pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/7 "2018-07-16T12:33:31Z")

</div>

sorry for the weird way to share output... but I just don't have internet access (or copy paste for that matter to the machine at the moment... ), but here it is:

 ![1](https://us1.discourse-cdn.com/elastic/original/3X/9/5/95fd7c2ecd3eb92285e25a0af57622218d4505b7.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 16, 2018, 12:38pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/8 "2018-07-16T12:38:25Z")

</div>

62 shards for 407 MB of data is far too much. That data would easily fit within a single shard.

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [July 16, 2018, 12:41pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/9 "2018-07-16T12:41:29Z")

</div>

I know.  
this is a test machine. on production there will be more data.  
but still it would be less than 1-2GB per index.  
so that is why I was thinking to have just 1 shard per index.  
(the separation to indices is mandatory, since those are different types of data from different applications).

are there any things to consider against using a single shard per index (other than HA, etc... since we already don't have that when running on a single node)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 16, 2018, 12:47pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/10 "2018-07-16T12:47:52Z")

</div>

> [@dboss101](#):
>
> so that is why I was thinking to have just 1 shard per index.

That would be recommended. You may also want to consider having each index cover a longer retention period.

> [@dboss101](#):
>
> are there any things to consider against using a single shard per index (other than HA, etc... since we already don't have that when running on a single node)

No.

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [July 16, 2018, 12:51pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/11 "2018-07-16T12:51:14Z")

</div>

thank you very much.

what about my other question at the beginning of the thread?  
how can I effectively benchmark this system with real life data, so that I can understand our limits with the given system?  
any pointers on that one?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 16, 2018, 1:00pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/12 "2018-07-16T13:00:23Z")

</div>

Have a look at the following resources:

[https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing](https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing)

[https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right](https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right)

[https://www.elastic.co/elasticon/conf/2018/sf/the-seven-deadly-sins-of-elasticsearch-benchmarking](https://www.elastic.co/elasticon/conf/2018/sf/the-seven-deadly-sins-of-elasticsearch-benchmarking)

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [July 16, 2018, 1:04pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/13 "2018-07-16T13:04:47Z")

</div>

thank you very much.

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [August 9, 2018, 6:20pm UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/14 "2018-08-09T18:20:03Z")

</div>

Hi again,  
so I've been doing some benchmarking, and I do think that it's possible that our HD is actually the or a bottleneck in this case.  
I'm getting results of around 60 MB/s for read speed when simply using hdparm, and something similar for write speed when using dd with blocks of 8k to test. results actually vary when testing with different counts of blocks. for example when trying 10k blocks (80MB for the test file in total), the speed is around 250 MB/s, but when trying 100k blocks, it goes down to around 100 MB/s, and when going up to 1000k (8GB file) it goes back to around 250 MB/s. so i'm not sure I really understand what's the deal here entirely. maybe the testing is faulty at some paint, but still the write speed never goes below 100 MB/s, that's for sure.

anyhow, the question I'm struggling with is how do I find out what is the real bottleneck for elasticsearch?

I've ran esrally with various tracks.  
for example I've ran this track:

`esrally --distribution-version=5.6.3 --track noaa --car 4gheap --challenge=append-no-conflicts --include-tasks="index" --track-params="bulk_size:10000,ingest_percentage:5"`

and the results look like this:

 ![esrally_noaa_4gheap_10kBulk](https://us1.discourse-cdn.com/elastic/original/3X/3/a/3ad79b52e596e13634a49f4c9e8e1595b1336eb5.png)

can anyone give me some lead on what I can deduct from those and other metrics?  
what is the capacity that I can count on, with the current setup?  
what can I look to improve? what's most important to improve?

many thanks!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 10, 2018, 6:24am UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/15 "2018-08-10T06:24:35Z")

</div>

Elasticsearch performs a good amount of random access reads and writes, so comparing the throughput to a test that reads and/or writes quite large blocks may not be very accurate. I would recommend monitoring I/O performance, e.g. using `iostat`, while your cluster is running to see what I/O latencies and how much iowait you are experiencing. This will give a good indication whether the disk is the bottleneck or not (I suspect it likely is).

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [August 10, 2018, 8:31am UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/16 "2018-08-10T08:31:13Z")

</div>

thank you. I have done some monitoring like `iostat` and other tools during normal operation of the cluster, and they all indicate some levels of iowait that the cpu is experiencing. sometimes around 20%, sometimes less, in some extreme cases I've seen even 70%.

but assuming this is the hardware I have to deal with, how can I best benchmark the cluster with `esrally`, but with my own data. because I'm not certain that the existing tracks are representing accurately my situation.

any pointers there?

thanks!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 10, 2018, 8:36am UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/17 "2018-08-10T08:36:53Z")

</div>

The best way is to create a custom Rally track that simulates your use case. Exactly how to best do that depends on the use-case you want to simulate. Can you provide some additional details about your use-case and the types of data you have?

---

<div class="post-metadata">

**Author:** ![dboss101](https://avatars.discourse-cdn.com/v4/letter/d/7993a0/32.png) [@dboss101](https://discuss.elastic.co/u/dboss101)\
**Post date:** [August 10, 2018, 8:43am UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/18 "2018-08-10T08:43:54Z")

</div>

well the main use case is ingestion of metric data from multiple network devices. some security data, some network traffic data, etc. most of the fields are either strings or different numeric fields.  
what other additional details can I provide?

[Edit]:  
and regarding the other side of things, this data is being queried by custom dashboards (not kibana), that are used as near live monitors and generate some alerts from this data.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 10, 2018, 9:26am UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/19 "2018-08-10T09:26:15Z")

</div>

You could have a look at the [rally-eventdata-track](https://github.com/elastic/rally-eventdata-track), which simulates indexing and querying of nginx access logs. You might be able to adapt this or use it for inspiration when creating your own track. You may even get valuable information from running it as is even though it does not accurately reflect your data and access patterns.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 7, 2018, 9:26am UTC](https://discuss.elastic.co/t/poor-performance-and-a-lot-of-gc-overhead/140145/20 "2018-09-07T09:26:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
