# Random appear io timeout error: \`Unable to reach APM Server(): Read timed out\`

**URL:** <https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488>\
**Category:** APM\
**Tags:** python, server\
**Created:** [January 19, 2021, 3:17am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488 "2021-01-19T03:17:12Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ackerr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ackerr/32/82484_2.png) [@Ackerr](https://discuss.elastic.co/u/Ackerr)\
**Post date:** [January 19, 2021, 3:17am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/1 "2021-01-19T03:17:12Z")

</div>

**Kibana version** : 7.10.1

**Elasticsearch version** : 7.10.1

**APM Server version** :7.10.1

**APM Agent language and version** : apm-agent-python 5.10.1

**Browser version** :

**Original install method (e.g. download page, yum, deb, from source, etc.) and version**: docker

**Fresh install or upgraded from other version?**

**Is there anything special in your setup?**

Use k8s deploy apm-server, kibana，elasticsearch. Ingress use nginx-ingress.

**apm-server configmap**

```auto
apm-server.yml:
  host: 0.0.0.0:8200
  max_event_size: 10485760
  rum:
    enabled: true
    allow_origins: ['*']
queue.mem.events: 12287
output.elasticsearch:
  worker: 3
  bulk_max_size: 4096
  hosts: ["elasticsearch:9200"]
logging.level: error
apm-server.ilm: 
  enabled: "auto"

```

**apm-agent-python config**

```auto
ELASTIC_APM_SERVER_URL="host:80"
ELASTIC_APM_TRANSACTIONS_IGNORE_PATTERNS="'^OPTIONS '"
ELASTIC_APM_CENTRAL_CONFIG="False"
ELASTIC_APM_CAPTURE_BODY="transactions"
ELASTIC_APM_TRANSACTION_SAMPLE_RATE="0.2"
ELASTIC_APM_DJANGO_TRANSACTION_NAME_FROM_ROUTE="True"

```

**Description of the problem including expected versus actual behavior. Please include screenshots (if relevant)**:  
Every once in a while，all project use `apm-agent-python` will report timeout error. but `apm-server` is normal and dont found related error report。

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/d/a/dad208eaba5082caa787e46b630734c539caf5f6.png)

**Steps to reproduce** :

1. restart apm-server (i think this because apm-server is not graceful shutdown)
2. random appear

**Errors in browser console (if relevant)**:  
apm-agent-python report error

```auto
TransportException("Unable to reach APM Server: HTTPConnectionPool(host='host', port=80): Read timed out. (read timeout=5) (url: http://host:80/intake/v2/events)")

```

---

<div class="post-metadata">

**Author:** ![axw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/axw/32/28197_2.png) [@axw](https://discuss.elastic.co/u/axw)\
**Post date:** [January 19, 2021, 3:58am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/2 "2021-01-19T03:58:52Z")

</div>

@Ackerr welcome to the forum!

> **Steps to reproduce** :
> 
> 1. restart apm-server (i think this because apm-server is not graceful shutdown)
> 2. random appear

Can you please clarify: does this issue only occur after restarting apm-server? Is it just for a short time, or does it continue happening indefinitely?

And when you say restarting the server, which exact commands are you using? If you can provide a k8s manifest and list the kubectl commands you're running, that would be ideal.

---

<div class="post-metadata">

**Author:** ![Ackerr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ackerr/32/82484_2.png) [@Ackerr](https://discuss.elastic.co/u/Ackerr)\
**Post date:** [January 19, 2021, 4:20am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/3 "2021-01-19T04:20:50Z")

</div>

Restart apm-server, every python-agent will report the error once.

command like this:

```auto
kubectl rollout restart deploy/apm -n elk
kubectl rollout status deploy/apm -n elk

```

k8s apm-server.yml:

```auto
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: apm
  namespace: elk
spec:
  selector:
    matchLabels:
      app: apm
  replicas: 4
  revisionHistoryLimit: 0
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0
      maxSurge: 1
  template:
    metadata:
      labels:
        app: apm
    spec:
      restartPolicy: Always
      containers:
        - name: apm
          image: docker.elastic.co/apm/apm-server:7.10.1
          ports:
            - containerPort: 8200
          resources:
            limits:
              cpu: 0.2
              memory: 800Mi
            requests:
              cpu: 0.1
              memory: 500Mi
          livenessProbe:
            tcpSocket:
              port: 8200
            initialDelaySeconds: 10
            periodSeconds: 10
          readinessProbe:
            httpGet:
              scheme: HTTP
              path: /
              port: 8200
            initialDelaySeconds: 10
            periodSeconds: 10
          startupProbe:
            httpGet:
              path: /
              port: 8200
              scheme: HTTP
            periodSeconds: 10
            failureThreshold: 3
          volumeMounts:
            - name: apm-data
              mountPath: /usr/share/apm-server/apm-server.yml
              subPath: apm-server.yml
      volumes:
        - name: apm-data
          configMap:
            name: apm-configmap

```

---

<div class="post-metadata">

**Author:** ![Ackerr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ackerr/32/82484_2.png) [@Ackerr](https://discuss.elastic.co/u/Ackerr)\
**Post date:** [January 19, 2021, 4:26am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/4 "2021-01-19T04:26:11Z")

</div>

And agent will report error once in a while( a few hours) , but apm-server is normal and only have `forbidden request` log

```auto
ERROR	[request]	middleware/log_middleware.go:99	forbidden request	{"request_id": "660bcb35-ab3a-4189-ab12-2760358cbcae", "method": "POST", "URL": "/config/v1/agents", "content_length": 61, "remote_address": "10.0.27.12", "user-agent": "elasticapm-python/5.10.0", "event.duration": 119703, "response_code": 403, "error": "forbidden request: Agent remote configuration is disabled. Configure the `apm-server.kibana` section in apm-server.yml to enable it. If you are using a RUM agent, you also need to configure the `apm-server.rum` section. If you are not using remote configuration, you can safely ignore this error."}

```

---

<div class="post-metadata">

**Author:** ![Ackerr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ackerr/32/82484_2.png) [@Ackerr](https://discuss.elastic.co/u/Ackerr)\
**Post date:** [January 20, 2021, 6:07am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/5 "2021-01-20T06:07:45Z")

</div>

I try add lifecycle in apm-server.yml. Now, Then `rollout restart` apm-server，no more errors report.

```auto
lifecycle:
  preStop:
    exec:
      command: ["sh", "-c", "sleep 5 && kill -s HUP 1"]

```

Random appear the error, not sure if it's because apm-server `queue.mem.events` set too large.

```auto
    queue.mem.events: 12288
    output.elasticsearch:
      worker: 3
      bulk_max_size: 4096

```

---

<div class="post-metadata">

**Author:** ![axw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/axw/32/28197_2.png) [@axw](https://discuss.elastic.co/u/axw)\
**Post date:** [January 20, 2021, 7:15am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/6 "2021-01-20T07:15:59Z")

</div>

> [@Ackerr](#):
>
> `command: ["sh", "-c", "sleep 5 && kill -s HUP 1"]`

I'm not a Kubernetes expert (I'm hoping someone else can chime in on this topic), but if I understand the docs correctly then `SIGTERM` is sent to the process by default (as I would expect). It's very surprising to me if your change fixes things; APM Server should perform a graceful shutdown on either `SIGTERM` or `SIGHUP`. I guess the sleep is somehow helping.

If nobody else responds with an answer, I'll try to reproduce the issue soon.

---

<div class="post-metadata">

**Author:** ![Ackerr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ackerr/32/82484_2.png) [@Ackerr](https://discuss.elastic.co/u/Ackerr)\
**Post date:** [January 20, 2021, 7:47am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/7 "2021-01-20T07:47:34Z")

</div>

Thanks for your reply !

Add some context

Before this error, Sentry often reported another error `queue is full`, Refer to [the document](https://www.elastic.co/guide/en/apm/server/master/common-problems.html#queue-full)

I add this config in apm-server config to fix it.

```auto
    queue.mem.events: 12288
    output.elasticsearch:
      worker: 3
      bulk_max_size: 4096

```

This config really fixed the `queue is full` error, but often report the `Read time out` error.

Now, I try to trun down the `queue.mem.events`，I don't know if it worked.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 10, 2021, 3:47am UTC](https://discuss.elastic.co/t/random-appear-io-timeout-error-unable-to-reach-apm-server-read-timed-out/261488/8 "2021-02-10T03:47:53Z")

</div>

This topic was automatically closed 20 days after the last reply. New replies are no longer allowed.
