# Fleet server stuck upgrading 8.5.3 -\> 8.6.0

**URL:** https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204
**Category:** Elastic Agent
**Tags:** fleet
**Created:** [January 16, 2023, 12:41am UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204 "2023-01-16T00:41:38Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![rossw](https://avatars.discourse-cdn.com/v4/letter/r/ecb155/32.png) [@rossw](https://discuss.elastic.co/u/rossw)
#### Post date: [January 16, 2023, 12:41am UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204/1 "2023-01-16T00:41:38Z")

</div>

Our fleet server runs on our kibana server, with about 20 agents connected.  
We upgrade our Elastic cluster from 8.5.3 to 8.6.0 and the upgrade popped up in the fleet UI just fine. Selected the fleet server for the upgrade, and the status changed to upgrading; and its been sitting like that for four days now.  
The agent is still running 8.5.3, and none of the logs show any reference to attempting to download the 8.6.0 code from the elastic repository. I have rebooted the kibana server, and the agent on its own. and it still just sits there running 8.5.3, and the UI status is updating.  
Elastic-agent status gives this response:

```auto
 elastic-agent status
Status: HEALTHY
Message: (no message)
Applications:
  * fleet-server (HEALTHY)
                           Running on default policy with Fleet Server integration
  * filebeat_monitoring (HEALTHY)
                           Running
  * metricbeat_monitoring (HEALTHY)
                           Running

```

Anyone got any ideas?

Ross

---

<div class="post-metadata">

### Author: ![AndersonQ](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andersonq/32/112214_2.png) [@AndersonQ](https://discuss.elastic.co/u/AndersonQ)
#### Post date: [January 17, 2023, 1:59pm UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204/2 "2023-01-17T13:59:43Z")

</div>

Hello,

It's hard to give an assertive answer without investigating the Elastic Agent and Fleet Server logs.

If the agent didn't update, I'd expect to have some error or warn message in the logs.

What you can do is to check the agent document on ES to see if there is a `upgrade_started_at` field, but `upgraded_at` is either missing or is `null`.

You can query the agent with:  
`GET .fleet-agents/_doc/AGENT_ID`

You can also re-trigger the upgrade using the Kibana Fleet API:

```auto
curl --request POST \
  --url https://<KIBANA_HOST>/api/fleet/agents/<AGENT_ID>/upgrade \
  --user "<SUPERUSER_NAME>:<SUPERUSER_PASSWORD>" \
  --header 'Content-Type: application/json' \
  --header 'kbn-xsrf: as' \
  --data '{"version": "<VERSION>","force": true}'

```

---

<div class="post-metadata">

### Author: ![rossw](https://avatars.discourse-cdn.com/v4/letter/r/ecb155/32.png) [@rossw](https://discuss.elastic.co/u/rossw)
#### Post date: [January 19, 2023, 3:55am UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204/3 "2023-01-19T03:55:08Z")

</div>

Thanks for that. I re-triggered the update, and it seemed to work in that the fleet-server is now running 8.6.0, but it is unhealthy, and the logs are full of:

```auto
{"log.level":"error","@timestamp":"2023-01-19T03:50:21.594Z","message":"Error fetching data for metricset beat.state: error making http request: Get \"http://unix/state\": dial unix /opt/Elastic/Agent/data/tmp/fleet-server-default.sock: connect: no such file or directory","component":{"binary":"metricbeat","dataset":"elastic_agent.metricbeat","id":"beat/metrics-monitoring","type":"beat/metrics"},"log.origin":{"file.line":256,"file.name":"module/wrapper.go"},"service.name":"metricbeat","ecs.version":"1.6.0","ecs.version":"1.6.0"}

```

Looking in the data/tmp directory I can see that that socket does not exist so it seems the fleet agent is not starting properly, as seen here:

```auto
./elastic-agent status
State: DEGRADED
Message: 1 or more components/units in a failed state
Components:
  * fleet-server (HEALTHY)
                  Healthy: communicating with pid '1700'
  * http/metrics (HEALTHY)
                  Healthy: communicating with pid '1710'
  * filestream (HEALTHY)
                  Healthy: communicating with pid '1719'
  * beat/metrics (HEALTHY)
                  Healthy: communicating with pid '1729'

```

Unfortunately, the logs don't indicate any issues when starting up, or why the sock is not created.  
Also, nothing is listening on 8220 so none of the agents can check in etc.

---

<div class="post-metadata">

### Author: ![rossw](https://avatars.discourse-cdn.com/v4/letter/r/ecb155/32.png) [@rossw](https://discuss.elastic.co/u/rossw)
#### Post date: [January 19, 2023, 4:28am UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204/4 "2023-01-19T04:28:12Z")

</div>

Hmm, just saw this in the logs:

```auto
{"log.level":"error","@timestamp":"2023-01-19T04:23:17.502Z","log.origin":{"file.name":"coordinator/coordinator.go","file.line":833},"message":"Unit state changed fleet-server-default (STARTING->FAILED): invalid log level; must be one of: trace, debug, info, warning, error accessing 'fleet.agent.logging'","component":{"id":"fleet-server-default","state":"HEALTHY"},"unit":{"id":"fleet-server-default","type":"output","state":"FAILED","old_state":"STARTING"},"ecs.version":"1.6.0"}

```

Any idea where this is set, and how I can fix it? Because its the fleet server, I can't use the fleet/kibana UI to do anything

---

<div class="post-metadata">

### Author: ![blaker](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/blaker/32/65621_2.png) [@blaker](https://discuss.elastic.co/u/blaker)
#### Post date: [January 19, 2023, 6:27pm UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204/5 "2023-01-19T18:27:16Z")

</div>

Please use the command `./elastic-agent status --output=yaml` that will povide more detail on the status of each unit. It will show which unit is in a failed state.

---

<div class="post-metadata">

### Author: ![rossw](https://avatars.discourse-cdn.com/v4/letter/r/ecb155/32.png) [@rossw](https://discuss.elastic.co/u/rossw)
#### Post date: [January 19, 2023, 8:01pm UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204/6 "2023-01-19T20:01:33Z")

</div>

So here is that output, and it again points to an invalid logging setting somewhere. Is there someway from the command line or editing a file on the local server that this can be over-ridden? I have grepped the Agent directories and can't find where this is configured in a yaml file anywhere.

```auto
 elastic-agent status --output=yaml
info:
  id: a8470c55-c2d8-44e3-b922-ad371f18fcfe
  version: 8.6.0
  commit: b79a5db77b5d6ffab9855234f8371d9e53978a24
  build_time: 2023-01-04 22:53:22 +0000 UTC
  snapshot: false
state: 3
message: 1 or more components/units in a failed state
components:
- id: fleet-server-default
  name: fleet-server
  state: 2
  message: 'Healthy: communicating with pid ''2048'''
  units:
  - unit_id: fleet-server-default-fleet-server-fleet_server-d4d1e18a-2ff0-41ae-b9ee-7005baa4d068
    unit_type: 0
    state: 0
    message: waiting for output unit
  - unit_id: fleet-server-default
    unit_type: 1
    state: 4
    message: 'invalid log level; must be one of: trace, debug, info, warning, error
      accessing ''fleet.agent.logging'''
  version_info:
    name: fleet-server
    version: 8.6.0
    meta:
      build_time: 2023-01-04 19:26:24 +0000 UTC
      commit: 05088c13
- id: http/metrics-monitoring
  name: http/metrics
  state: 2
  message: 'Healthy: communicating with pid ''2060'''
  units:
  - unit_id: http/metrics-monitoring
    unit_type: 1
    state: 2
    message: Healthy
  - unit_id: http/metrics-monitoring-metrics-monitoring-agent
    unit_type: 0
    state: 2
    message: Healthy
  version_info:
    name: beat-v2-client
    version: 8.6.0
    meta:
      build_time: 2023-01-04 01:30:07 +0000 UTC
      commit: 561a3e1839f1a50ce832e8e114de399b2bee2542
- id: filestream-monitoring
  name: filestream
  state: 2
  message: 'Healthy: communicating with pid ''2070'''
  units:
  - unit_id: filestream-monitoring
    unit_type: 1
    state: 2
    message: Healthy
  - unit_id: filestream-monitoring-filestream-monitoring-agent
    unit_type: 0
    state: 2
    message: Healthy
  version_info:
    name: beat-v2-client
    version: 8.6.0
    meta:
      build_time: 2023-01-04 01:28:13 +0000 UTC
      commit: 561a3e1839f1a50ce832e8e114de399b2bee2542
- id: beat/metrics-monitoring
  name: beat/metrics
  state: 2
  message: 'Healthy: communicating with pid ''2081'''
  units:
  - unit_id: beat/metrics-monitoring-metrics-monitoring-beats
    unit_type: 0
    state: 2
    message: Healthy
  - unit_id: beat/metrics-monitoring
    unit_type: 1
    state: 2
    message: Healthy
  version_info:
    name: beat-v2-client
    version: 8.6.0
    meta:
      build_time: 2023-01-04 01:30:07 +0000 UTC
      commit: 561a3e1839f1a50ce832e8e114de399b2bee2542

```

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [February 16, 2023, 8:02pm UTC](https://discuss.elastic.co/t/fleet-server-stuck-upgrading-8-5-3-8-6-0/323204/7 "2023-02-16T20:02:01Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
