# Fleet Server is unstable. Can't connect new hosts but status is 'healthy'

**URL:** <https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009>\
**Category:** Elastic Security\
**Tags:** fleet\
**Created:** [March 7, 2022, 7:37pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009 "2022-03-07T19:37:34Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [March 7, 2022, 7:37pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/1 "2022-03-07T19:37:35Z")

</div>

Hello,

since end of February I was not able to add new hosts to my fleet using my fleet server. Also I do not receive data from my Windows Domain Controller anymore for some reason.  
When I try to add another Windows Host it just says that the remote server 'is not ready to accept connections yet', but it does that over and over with no success. Tried the same with an Ubuntu container with the same result.  
If I do `curl -f http://fleet-server-ip:8220/api/status` from the host it sometimes does not respond at all and sometimes it says: `{"name":"fleet-server","status":"HEALTHY"}%`.  
Other hosts that were already added before end of February are fine and 'healthy' (besides the above mentioned DC controller) but I just can't add new ones.

ELK and all agents etc. are running version 8.0.0

I would appreciate any help, thank you

---

<div class="post-metadata">

**Author:** ![Nima\_Rezainia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nima_rezainia/32/88626_2.png) [@Nima\_Rezainia](https://discuss.elastic.co/u/Nima_Rezainia)\
**Post date:** [March 7, 2022, 10:05pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/2 "2022-03-07T22:05:47Z")

</div>

would you be able to show us the configuration on the fleet-server from the UI?  
(navigate integrations--\>fleet server --\> fleet server settings) you will see Fleet Server and an expansion. That will show you "Max Connections" and a yaml box for other config.

How many agents do you have?

thanks

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [March 7, 2022, 11:04pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/3 "2022-03-07T23:04:25Z")

</div>

You mean that?

 ![Bildschirmfoto 2022-03-07 um 23.57.50](https://us1.discourse-cdn.com/elastic/original/3X/3/a/3a217969626992933af3a261b29fac8798257398.png)

I have 10 agents enrolled:

 ![Bildschirmfoto 2022-03-08 um 00.02.15](https://us1.discourse-cdn.com/elastic/original/3X/6/3/63153a4982f9fa929261da503a3e1bb674da5fff.png)

pFleet was just added recently by myself in an effort to just use another fleet server but this second fleet server doesn't work ("`Error: failed to communicate with Elastic Agent daemon: rpc error: code = Unavailable desc = connection error: desc = "transport: Error while dialing dial unix /run/elastic-agent.sock: connect: no such file or directory`").

pMinecraft 'never checked in' (that's the Ubuntu container I tried to enroll), the Windows Server PVE-WINSRV-DC2 is the DC I talked about which, as can be seen has not sent logs since end of February. The other Windows host I tried to add doesn't show up here as I couldn't enroll it in the first place.

---

<div class="post-metadata">

**Author:** ![Nima\_Rezainia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nima_rezainia/32/88626_2.png) [@Nima\_Rezainia](https://discuss.elastic.co/u/Nima_Rezainia)\
**Post date:** [March 7, 2022, 11:25pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/4 "2022-03-07T23:25:51Z")

</div>

yes that page. The default config is satisfactory for your use here. ignoring pfleet for now, what does "elastic-agent status" command give you for the PVE-WINSRV-DC2 (execute as super user on that host).

What integrations have yo installed on the default policy?

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [March 7, 2022, 11:49pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/5 "2022-03-07T23:49:51Z")

</div>

Ok did that. Seems 'HEALTHY':

 ![Bildschirmfoto 2022-03-08 um 00.45.32](https://us1.discourse-cdn.com/elastic/original/3X/a/8/a89b2559c3ad18ba968b5240d6a4188c76dea6c8.jpeg)

This is the policy that I use for this DC (or all Windows hosts to be precise):

 ![Bildschirmfoto 2022-03-08 um 00.47.34](https://us1.discourse-cdn.com/elastic/original/3X/1/6/163adb102d5deb1b2be767628ec75bb9997eefc5.png)

Do you think upgrading Endpoint Security would be worth it?  
(I didn't want to risk it yet, as I have another host running this policy (Windows Host, not Server) which works just fine last time I checked)

---

<div class="post-metadata">

**Author:** ![Nima\_Rezainia](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nima_rezainia/32/88626_2.png) [@Nima\_Rezainia](https://discuss.elastic.co/u/Nima_Rezainia)\
**Post date:** [March 8, 2022, 12:06am UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/6 "2022-03-08T00:06:36Z")

</div>

can't tell tbh. all the agents in the "Default policy (windows)" are either offline or unhealthy. But others in Default policy are fine. What does endpoint security look like in the Default Policy?

we are working on adding more status reporting on the integrations to help with diagnostics.

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [March 8, 2022, 12:36am UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/7 "2022-03-08T00:36:14Z")

</div>

The default policy is just the Windows policy minus the Windows-Logger. It also recommends to upgrade Endpoint Security there. Also pMinecraft is in the Default Policy but not fine (It's stuck at "Updating", saying it never 'checked in')

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [March 8, 2022, 12:40am UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/8 "2022-03-08T00:40:43Z")

</div>

It btw seems like fleet is not sending any data besides empty TCP packets if requesting `curl http://10.24.1.7:8220/api/status`

10.24.1.5 is ES, 10.24.1.7 is Logstash and Fleet Server and 10.24.1.3 is the DC:

```auto
01:38:27.835299 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.7.8220 > 10.24.1.3.56625: Flags [S.], cksum 0x9f02 (correct), seq 748443443, ack 1512235924, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:28.835551 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.7.8220 > 10.24.1.3.56625: Flags [S.], cksum 0x9f02 (correct), seq 748443443, ack 1512235924, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:29.298891 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56626: Flags [S.], cksum 0x7399 (correct), seq 3175754204, ack 2289793402, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:29.731539 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56623: Flags [S.], cksum 0x39ca (correct), seq 773930180, ack 2040741478, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:30.307490 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56626: Flags [S.], cksum 0x7399 (correct), seq 3175754204, ack 2289793402, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:30.823353 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.7.8220 > 10.24.1.3.56625: Flags [S.], cksum 0x9f02 (correct), seq 748443443, ack 1512235924, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:31.015485 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56619: Flags [S.], cksum 0xf2fd (correct), seq 1619032043, ack 3708073038, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:32.298632 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56626: Flags [S.], cksum 0x7399 (correct), seq 3175754204, ack 2289793402, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:32.835490 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.7.8220 > 10.24.1.3.56625: Flags [S.], cksum 0x9f02 (correct), seq 748443443, ack 1512235924, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:34.307466 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56626: Flags [S.], cksum 0x7399 (correct), seq 3175754204, ack 2289793402, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:34.508482 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56627: Flags [S.], cksum 0x3295 (correct), seq 1381733823, ack 801968696, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0
01:38:35.523462 IP (tos 0x0, ttl 64, id 0, offset 0, flags [DF], proto TCP (6), length 52)
    10.24.1.5.9200 > 10.24.1.3.56627: Flags [S.], cksum 0x3295 (correct), seq 1381733823, ack 801968696, win 64240, options [mss 1460,nop,nop,sackOK,nop,wscale 7], length 0

```

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [March 12, 2022, 8:05pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/9 "2022-03-12T20:05:05Z")

</div>

Any update?

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [March 28, 2022, 1:45pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/10 "2022-03-28T13:45:54Z")

</div>

A friend of mine just installed the whole ELK stack from scratch and he has the same problem.

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [April 9, 2022, 2:55pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/11 "2022-04-09T14:55:09Z")

</div>

Update: Completely reinstalled the system on another machine (v8.1.2 now) Still doesn't work!

@Nima_Rezainia We are still struggling with this problem. Can't enroll any agent.

---

<div class="post-metadata">

**Author:** ![zx8086](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zx8086/32/94917_2.png) [@zx8086](https://discuss.elastic.co/u/zx8086)\
**Post date:** [April 9, 2022, 3:11pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/12 "2022-04-09T15:11:50Z")

</div>

What is the current state ? Are you still running on http or https ?

What are your Fleet Settings ?

 ![Screenshot 2022-04-09 at 17.11.20](https://us1.discourse-cdn.com/elastic/original/3X/a/8/a8aee2033a68d4eba9c114b78c9c4aa03e8d4ea2.png)

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [April 9, 2022, 3:36pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/13 "2022-04-09T15:36:09Z")

</div>

Fleet still reports its healthy:

```auto
[root@pfleet ~]# curl 10.20.1.8:8220/api/status

{"name":"fleet-server","status":"HEALTHY"}

```

 ![Bildschirmfoto 2022-04-09 um 17.24.00](https://us1.discourse-cdn.com/elastic/original/3X/0/4/04f71ce0a07f47b590199d20ac1518d723ec54ee.png)

---

<div class="post-metadata">

**Author:** ![zx8086](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zx8086/32/94917_2.png) [@zx8086](https://discuss.elastic.co/u/zx8086)\
**Post date:** [April 9, 2022, 4:08pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/14 "2022-04-09T16:08:01Z")

</div>

So for the Fleet Server you are using http, but for the Elasticsearch server you are using https ?

Also what is the command-line argument you are using to enroll ?

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [April 10, 2022, 5:04pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/16 "2022-04-10T17:04:01Z")

</div>

@zx8086 I used plain http for everything for my first setup this post was about. For the now used setup (that has the same issue) I use the default installation setting: https for Elasticsearch but not for Fleet or Kibana. I use the command suggested by the Kibana UI for adding host:  
`.\elastic-agent.exe install --url=http://10.20.1.8:8220 --enrollment-token=mytoken`

---

<div class="post-metadata">

**Author:** ![zx8086](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zx8086/32/94917_2.png) [@zx8086](https://discuss.elastic.co/u/zx8086)\
**Post date:** [April 10, 2022, 11:07pm UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/17 "2022-04-10T23:07:41Z")

</div>

@maof97

So both the fleet and elasticsearch servers are http and set as so in the fleet setting ?

You can curl to both of them and get correct responses?

```auto
curl 10.20.1.8:8220/api/status

```

```auto
curl 10.20.1.9200/_cluster/health?pretty

```

> [@maof97](#):
>
> .\elastic-agent.exe install --url=[http://10.20.1.8:8220](http://10.20.1.8:8220) --enrollment-token=mytoken

Also if you are using only http, you should use the --insecure flag

---

<div class="post-metadata">

**Author:** ![maof97](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/maof97/32/101433_2.png) [@maof97](https://discuss.elastic.co/u/maof97)\
**Post date:** [April 11, 2022, 12:21am UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/18 "2022-04-11T00:21:55Z")

</div>

> [@zx8086](#):
>
> @maof97
> 
> So both the fleet and elasticsearch servers are http and set as so in the fleet setting ?
> 
> You can curl to both of them and get correct responses?
> 
> ```auto
> curl 10.20.1.8:8220/api/status
> 
> ```

Sometimes it's _healthy_ sometimes I get "_Connection Refused_". That the problem, here ...

> [@zx8086](#):
>
> ```auto
> curl 10.20.1.9200/_cluster/health?pretty
> 
> ```

I think you mean curl _http **s** ://10.20.1.6:9200/\_cluster/health?pretty_ because as I stated before, Elasticsearch is using https and the rest http:

```auto
{
  "cluster_name" : "elasticsearch",
  "status" : "yellow",
  "timed_out" : false,
  "number_of_nodes" : 1,
  "number_of_data_nodes" : 1,
  "active_primary_shards" : 44,
  "active_shards" : 44,
  "relocating_shards" : 0,
  "initializing_shards" : 0,
  "unassigned_shards" : 23,
  "delayed_unassigned_shards" : 0,
  "number_of_pending_tasks" : 0,
  "number_of_in_flight_fetch" : 0,
  "task_max_waiting_in_queue_millis" : 0,
  "active_shards_percent_as_number" : 65.67164179104478
}

```

> [@zx8086](#):
>
> Also if you are using only http, you should use the --insecure flag

Yes I used the `--insecure` flag, forgot to copy that.

---

<div class="post-metadata">

**Author:** ![zx8086](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zx8086/32/94917_2.png) [@zx8086](https://discuss.elastic.co/u/zx8086)\
**Post date:** [April 11, 2022, 12:34am UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/19 "2022-04-11T00:34:02Z")

</div>

> [@maof97](#):
>
> Sometimes it's _healthy_ sometimes I get " _Connection Refused_ ". That the problem, here ...

This sounds more like the service itself (resources - memory, cpu) or infra / connectivity.

Are you running a single client to rule out load ?

If you can monitor the endpoint you should see more of the root cause, like if you put it behind a reverse proxy load balancer and monitor the upstream.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 9, 2022, 12:34am UTC](https://discuss.elastic.co/t/fleet-server-is-unstable-cant-connect-new-hosts-but-status-is-healthy/299009/20 "2022-05-09T00:34:46Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
