# APM Failed to publish events: temporary bulk send failure / Queue is full 503 error

**URL:** <https://discuss.elastic.co/t/apm-failed-to-publish-events-temporary-bulk-send-failure-queue-is-full-503-error/202884>\
**Category:** APM\
**Tags:** server\
**Created:** [October 9, 2019, 5:16pm UTC](https://discuss.elastic.co/t/apm-failed-to-publish-events-temporary-bulk-send-failure-queue-is-full-503-error/202884 "2019-10-09T17:16:02Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![tmihaldinec](https://avatars.discourse-cdn.com/v4/letter/t/ccd318/32.png) [@tmihaldinec](https://discuss.elastic.co/u/tmihaldinec)\
**Post date:** [October 11, 2019, 11:58am UTC](https://discuss.elastic.co/t/apm-failed-to-publish-events-temporary-bulk-send-failure-queue-is-full-503-error/202884/3 "2019-10-11T11:58:59Z")

</div>

Hi Juan,

i can remember that when we started with APM in version 7.1 i think i had similar issue but it was primary related to huge number of request and only one server.. than.. we added additional APM systems and removed unneeded APM systems and it was working with same or even higher throughput last 2, 3 months.. on 7.2 and 7.3

so yes.. i am pretty sure that it started with 7.4 version..

i know that this temporary bulk send failure is common issue which i happening on beats and because of this issue it puzzles me even more.. I have pretty extensive knowledge of designing and sizing systems and to me there is no system metric which would lead me to conclusion that something is undersized on ELK side (we have logstash on two machines which are doing way bigger load on same ELK and i dont see bulk send failure there

on other hand.. maybe i am doing something wrong and maybe APM is not ready for such load (6xAPM servers x 1000-1500 events/s and i need to add additional APM and/or elastic nodes (processing only)

20y of some experiences teach me that with horizontal sizing i might solve problem from user perspective but problem will still be there.. and 1500 events/sec is not such a huge load

load on servers is never bigger than 2 (8 core server CPU E5-2690 v4, IBM svc Tier 1 storage, 40+GB RAM per node)..

is there something i am missing? like some additional statistics or monitoring which would give me info why this problem is happening in first place (queue full) like where problem on elastic node is..

also.. i am reading today posts here and i see new post from today similar to mine with more or less same problem..

> [@APM: 503 Queue is full, server sleeping, nothing helps](https://discuss.elastic.co/t/apm-503-queue-is-full-server-sleeping-nothing-helps/203180):
>
> Hello, We are trying to use APM to monitor our website but so far APM starts producing 503 Queue is full error after some time. After this happens it won't get back to normal, only restart of the APM service helps. The server is literally sleeping, CPU usage was around 15% and memory only 50% full. When I enabled it today at night, I also did performance tests and there was no problem with 900rpm but then it crashed at around 50rpm... All the performance settings seems to be useless. I don't th…

i also have this problem that i need to restart APM.. process is up but it is not recovering..

maybe as a shot term solution to add some restart procedure if queue is full

example of state when it dies:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/5/f5e5aca8b4f68107490422e4bf00dda3dab7bdda.png)

Thanks in advance  
tomislav

---

_[View the full topic](https://discuss.elastic.co/t/apm-failed-to-publish-events-temporary-bulk-send-failure-queue-is-full-503-error/202884)._
