# APM server tuning for heavy workload

**URL:** <https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455>\
**Category:** APM\
**Tags:** java, server\
**Created:** [February 1, 2023, 4:01pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455 "2023-02-01T16:01:54Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![senyam08](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/senyam08/32/106230_2.png) [@senyam08](https://discuss.elastic.co/u/senyam08)\
**Post date:** [February 1, 2023, 4:01pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/1 "2023-02-01T16:01:54Z")

</div>

After deploying elastic cluster for tracing, we see following errors due to heavy workload in APM server/Agents  
co.elastic.apm.agent.report.ApmServerReporter - dropped events because of full queue: 305

[elastic-apm-server-reporter] ERROR co.elastic.apm.agent.report.IntakeV2ReportingEventHandler - Error sending data to APM server: Read timed out, response code is -1

Could you provide info on how to monitor APM server/java agent queue and ways to tune to resolve above errors? we are in 8.4 version and planning to upgrade to 8.61

---

<div class="post-metadata">

**Author:** ![Jonas\_Kunz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jonas_kunz/32/110740_2.png) [@Jonas\_Kunz](https://discuss.elastic.co/u/Jonas_Kunz)\
**Post date:** [February 2, 2023, 12:16pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/2 "2023-02-02T12:16:21Z")

</div>

The error messages you are seeing indicates that the APM server is overloaded as you said.  
You should start by [monitoring the APM server] ([Monitor APM Server | APM User Guide [8.6] | Elastic](https://www.elastic.co/guide/en/apm/guide/current/monitor-apm.html)) and either scale your deployment vertically or horizontally or reduce the amount of data generated on the agent side via [sampling](https://www.elastic.co/guide/en/apm/guide/8.6/sampling.html).

---

<div class="post-metadata">

**Author:** ![Eyal\_Koren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eyal_koren/32/36830_2.png) [@Eyal\_Koren](https://discuss.elastic.co/u/Eyal_Koren)\
**Post date:** [February 5, 2023, 6:33am UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/3 "2023-02-05T06:33:19Z")

</div>

Adding to what @Jonas_Kunz wrote, reducing the amount of data produced by the agents may be crucial, so in addition to sampling, check your configuration and make sure you did not set any that may cause this. For example, the `trace_methods` config may be the cause of creation of large amounts of spans if set to capture too many methods or methods that are executed very frequently.  
In addition, [`span_min_duration`](https://www.elastic.co/guide/en/apm/agent/java/current/config-core.html#config-span-min-duration) and [`exit_span_min_duration`](https://www.elastic.co/guide/en/apm/agent/java/current/config-huge-traces.html#config-exit-span-min-duration) can reduce the number of captured and sent spans by discarding the very fast ones.  
Lastly, take a look at our short [tuning guide](https://www.elastic.co/guide/en/apm/agent/java/current/tuning-and-overhead.html#tuning-agent), where there is a bit more info.

---

<div class="post-metadata">

**Author:** ![senyam08](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/senyam08/32/106230_2.png) [@senyam08](https://discuss.elastic.co/u/senyam08)\
**Post date:** [February 13, 2023, 5:18pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/4 "2023-02-13T17:18:40Z")

</div>

Thanks. Is there document with parameters to tune APM server with 8.X version?  
I see the below parameters in legacy APM server document  
apm-server.max\_event\_size  
queue.mem.events  
apm-server.read\_timeout | apm-server.write\_timeout  
output.elasticsearch.worker  
output.elasticsearch.bulk\_max\_size  
output.elasticsearch.timeout

---

<div class="post-metadata">

**Author:** ![senyam08](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/senyam08/32/106230_2.png) [@senyam08](https://discuss.elastic.co/u/senyam08)\
**Post date:** [February 13, 2023, 5:51pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/5 "2023-02-13T17:51:31Z")

</div>

could someone send information on how do i update with different value for some of the properties defined in apm-server.yml file

> <https://github.com/elastic/apm-server/blob/main/apm-server.yml>

---

<div class="post-metadata">

**Author:** ![marclop](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marclop/32/119339_2.png) [@marclop](https://discuss.elastic.co/u/marclop)\
**Post date:** [February 21, 2023, 8:38am UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/6 "2023-02-21T08:38:46Z")

</div>

> [@senyam08](#):
>
> Thanks. Is there document with parameters to tune APM server with 8.X version?

APM Server doesn't need any output tunning, and versions \>= 8.6.0 ship with major performance improvements. These are most noticeable given a powerful enough machine (\>=6 cores / threads).

APM Server needs to be scaled in conjunction with Elasticsearch, making sure that your Elasticsearch is scaled out and up enough to be able to handle the load created by APM Server. In the majority of cases, Elasticsearch isn't scaled up or out to be able to process the APM Server's throughput, be sure to keep an eye on Elasticsearch's CPU usage and APM Server's usage. Very long response times may be an indicator that Elasticsearch is overloaded (needs more resources, more machines or the indices need to be tuned with a higher number of shards (generally 1 shard per Elasticsearch node)). By default the APM indices use a single shard.

My colleagues have pointed you to the agent specific configuration that can be tuned to reduce the amount of data sent to the APM Server.

---

<div class="post-metadata">

**Author:** ![senyam08](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/senyam08/32/106230_2.png) [@senyam08](https://discuss.elastic.co/u/senyam08)\
**Post date:** [February 21, 2023, 3:09pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/7 "2023-02-21T15:09:52Z")

</div>

Thank you for detailed explanation. is there any prebuilt dashboard to proactively monitor health of APM server/Elastic search and find out any issues with bulk processing/scaling?

---

<div class="post-metadata">

**Author:** ![marclop](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marclop/32/119339_2.png) [@marclop](https://discuss.elastic.co/u/marclop)\
**Post date:** [February 22, 2023, 12:09am UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/8 "2023-02-22T00:09:29Z")

</div>

There isn't one as far as I know. However, you can set up Stack Monitoring and use it to monitor both Elasticsearch and APM Server metrics: [Stack Monitoring | Kibana Guide [8.6] | Elastic](https://www.elastic.co/guide/en/kibana/current/xpack-monitoring.html).

---

<div class="post-metadata">

**Author:** ![senyam08](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/senyam08/32/106230_2.png) [@senyam08](https://discuss.elastic.co/u/senyam08)\
**Post date:** [February 23, 2023, 4:23pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/9 "2023-02-23T16:23:07Z")

</div>

Thanks Marc. After adding APM agent tuning config changes and increased Elasticsearch data/ingest instances and APM server, we are still seeing following error( we are @ 8.6.1 version)

2023-02-23 16:18:38,959 [elastic-apm-server-reporter] INFO co.elastic.apm.agent.report.IntakeV2ReportingEventHandler - Backing off for 4 seconds (+/-10%)  
2023-02-23 16:18:38,960 [elastic-apm-server-reporter] ERROR co.elastic.apm.agent.report.IntakeV2ReportingEventHandler - Error sending data to APM server: Read timed out, response code is -1  
2023-02-23 16:18:38,960 [elastic-apm-server-reporter] WARN co.elastic.apm.agent.report.IntakeV2ReportingEventHandler - null

---

<div class="post-metadata">

**Author:** ![marclop](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marclop/32/119339_2.png) [@marclop](https://discuss.elastic.co/u/marclop)\
**Post date:** [February 24, 2023, 12:10am UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/10 "2023-02-24T00:10:43Z")

</div>

It is really hard and nearly impossible for me to give any further guidance or suggestion without knowing any of the specifics of your set up. Could you please provide detailed metrics for:

- APM & Elasticsearch CPU utilization.
- APM & Elasticsearch sizing and topology (CPU, RAM).

You can obtain these metrics by enabling stack monitoring and collecting screenshots of at least a 12-24h time window.

---

<div class="post-metadata">

**Author:** ![senyam08](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/senyam08/32/106230_2.png) [@senyam08](https://discuss.elastic.co/u/senyam08)\
**Post date:** [February 27, 2023, 5:57pm UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/11 "2023-02-27T17:57:29Z")

</div>

We got the following topology

APM server running 5 instances in Kubernetes - 2 CPU request and no limit - CPU usage is less than 60% during read timeout error  
Elasticsearch datanodes running 6 instances in Kubernetes - 2 CPU request and no limit - CPU usage is less than 40% during read timeout error  
Elasticsearch ingest nodes running 3 instances in Kubernetes - 2 CPU request and no limit - CPU usage is less than 80% during read timeout error  
Elasticsearch master running 3 instances in Kubernetes - 1 CPU request and no limit - CPU usage is less than 20% during read timeout error

Do i need to update any of the tuning params for APM server as i see read timeout and queue full error in APM agents logs?

---

<div class="post-metadata">

**Author:** ![lahsivjar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lahsivjar/32/101451_2.png) [@lahsivjar](https://discuss.elastic.co/u/lahsivjar)\
**Post date:** [March 2, 2023, 5:07am UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/12 "2023-03-02T05:07:14Z")

</div>

Thanks for sharing more details. From what I can glean, it seems that all the pods are running without CPU limit, can you share details on how you are calculating the CPU usage? Also, I would suggest looking at Kubernetes node's CPU utilization to make sure that the node's are not overwhelmed.

> [@senyam08](#):
>
> Do i need to update any of the tuning params for APM server as i see read timeout and queue full error in APM agents logs?

It shouldn't be needed. If I understand correctly, the APM-Server version is still 8.4(?). If yes can you upgrade it to the latest `8.6.x` version. As Marc mentioned earlier, APM-Server versions \>= 8.6.0 ship with major performance improvements and autoscaling of internal indexers.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 23, 2023, 1:07am UTC](https://discuss.elastic.co/t/apm-server-tuning-for-heavy-workload/324455/13 "2023-03-23T01:07:20Z")

</div>

This topic was automatically closed 20 days after the last reply. New replies are no longer allowed.
