# Cluster not forming

**URL:** <https://discuss.elastic.co/t/cluster-not-forming/374581>\
**Category:** Elasticsearch\
**Tags:** docker\
**Created:** [February 14, 2025, 10:45pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581 "2025-02-14T22:45:52Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![joelatgrayv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joelatgrayv/32/141394_2.png) [@joelatgrayv](https://discuss.elastic.co/u/joelatgrayv)\
**Post date:** [February 14, 2025, 10:45pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/1 "2025-02-14T22:45:52Z")

</div>

I’m trying to get an ES cluster working with a cloud formation. I have it all up and running but the cluster is not forming correctly.

This is the error that I’m getting:

{"@timestamp":"2025-02-14T22:02:57.598Z", "log.level": "INFO", "message":"close connection exception caught on transport layer [Netty4TcpChannel{localAddress=/127.0.0.1:37302, remoteAddress=es02.elasticsearch.local/127.255.0.2:9300, profile=default}], disconnecting from relevant node: Connection reset", "ecs.version": "1.2.0","service.name":"ES\_ECS","event.dataset":"elasticsearch.server","process.thread.name":"elasticsearch[es01][transport\_worker][T#1]","log.logger":"org.elasticsearch.transport.TcpTransport","elasticsearch.node.name":"es01","elasticsearch.cluster.name":"docker-cluster”}

This is the template that I am using. You just need to provide the PrivateSubnetIds, PublicSubnetIds and ECSTaskExecutionRoleArn

> **[elasticsearch - StackBlitz](https://stackblitz.com/edit/vitejs-vite-3bh1ihyf?file=elasticsearch.yml)**
>
> A Node.js project

Any ideas?

---

<div class="post-metadata">

**Author:** ![ahmed\_charafouddine](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ahmed_charafouddine/32/45129_2.png) [@ahmed\_charafouddine](https://discuss.elastic.co/u/ahmed_charafouddine)\
**Post date:** [February 17, 2025, 11:06am UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/2 "2025-02-17T11:06:29Z")

</div>

I feel like **Cluster name setting** is missing

> **[Important Elasticsearch configuration | Elasticsearch Guide \[7.17\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/important-settings.html#cluster-name)**

---

<div class="post-metadata">

**Author:** ![joelatgrayv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joelatgrayv/32/141394_2.png) [@joelatgrayv](https://discuss.elastic.co/u/joelatgrayv)\
**Post date:** [February 17, 2025, 3:52pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/3 "2025-02-17T15:52:55Z")

</div>

The container defaults to "docker-cluster". I overrode this just to test it out with "cluster01" and didn't fix the issue.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [February 17, 2025, 4:46pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/4 "2025-02-17T16:46:35Z")

</div>

I have zero idea on cloud formation, but the elastic docs say:

> You must set `cluster.initial_master_nodes` to the same list of nodes on each node on which it is set in order to be sure that only a single cluster forms during bootstrapping. If `cluster.initial_master_nodes` varies across the nodes on which it is set then you may bootstrap multiple clusters.

In your config, this setting varies.

Also, I dont know if after your cluster fails to come up properly, do all the nodes stay up? if so, then can you login and do some troubleshooting from there? i.e. does any (how many?) of the 3 nodes think it's part of a cluster, and if so what does it think its cluster is composed of?

---

<div class="post-metadata">

**Author:** ![joelatgrayv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joelatgrayv/32/141394_2.png) [@joelatgrayv](https://discuss.elastic.co/u/joelatgrayv)\
**Post date:** [February 17, 2025, 6:13pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/5 "2025-02-17T18:13:39Z")

</div>

I fixed it and tried with just one node, I get the same. The nodes stay up and running but they just log the connection errors. I confirmed that they can communicate with each other:

nc -zv es02.elasticsearch.local 9200  
Connection to es02.elasticsearch.local 9200 port [tcp/_] succeeded!  
nc -zv es02.elasticsearch.local 9300  
Connection to es02.elasticsearch.local 9300 port [tcp/_] succeeded!

```
  ContainerDefinitions:
    - Name: es01
      Cpu: !Ref ContainerCpu
      Memory: !Ref ContainerMemory
      Image: !Ref ImageUrl
      PortMappings:
        - ContainerPort: !Ref ContainerPort
          HostPort: !Ref ContainerPort
          Protocol: tcp
          Name: "api"
      LogConfiguration:
        LogDriver: awslogs
        Options:
          mode: non-blocking
          max-buffer-size: 25m
          awslogs-group: !Ref LogGroup
          awslogs-region: !Ref AWS::Region
          awslogs-stream-prefix: es01
      Ulimits:
        - Name: memlock
          SoftLimit: -1
          HardLimit: -1
      Environment:
        - Name: discovery.type
          Value: multi-node
        - Name: cluster.name
          value: cluster01
        - Name: node.name
          Value: "es01"
        - Name: cluster.initial_master_nodes
          Value: "es01"
        - Name: discovery.seed_hosts
          Value: "es01.elasticsearch.local"
        - Name: discovery.cluster_formation_warning_timeout
          Value: "10m"
        - Name: ES_JAVA_OPTS
          Value: "-Xms6g -Xmx6g"
        - Name: xpack.security.enabled
          Value: false
        - Name: xpack.security.transport.ssl.enabled
          Value: false
        - Name: xpack.security.http.ssl.enabled
          Value: false
        - Name: xpack.security.authc.api_key.enabled
          Value: false
        - Name: xpack.security.authc.realms.native.native1.enabled
          Value: true
        - Name: xpack.security.authc.realms.native.native1.order
          Value: 0
        - Name: action.destructive_requires_name
          Value: false

```

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [February 17, 2025, 6:41pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/6 "2025-02-17T18:41:09Z")

</div>

When you start elasticsearch on one of the nodes, it should spit out all kinds of startup logs either to console, or log file. Share them here.

I had meant to use, eg, curl and do

curl -X GET -s -k [http://localhost:9200/\_cluster/health](http://localhost:9200/_cluster/health)

(or the local IP address)

on all 3 hosts.

But in end, for me anyways, you have a cloud formation issue (I dont really understand the syntax and how it uses it). For my 3-node docker cluster, I have the following in the compose file:

```auto
    environment:
      - node.name=es01
      - cluster.name=${CLUSTER_NAME}
      - cluster.initial_master_nodes=es01,es02,es03
      - discovery.seed_hosts=es02,es03

    environment:
      - node.name=es02
      - cluster.name=${CLUSTER_NAME}
      - cluster.initial_master_nodes=es01,es02,es03
      - discovery.seed_hosts=es01,es03

    environment:
      - node.name=es03
      - cluster.name=${CLUSTER_NAME}
      - cluster.initial_master_nodes=es01,es02,es03
      - discovery.seed_hosts=es01,es02

```

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [February 17, 2025, 7:09pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/7 "2025-02-17T19:09:05Z")

</div>

couple of little things

```auto
AWSTemplateFormatVersion: 2010-09-09

```

not important, but just the 15 years have passed since then . I thought AWS moved things along faster than that ...

2nd, why use 8.15.1? This reads to me like a new installations, why not use 8.17.2 aka latest ?

3rd, no security, not SSL, in-the-clear-HTTP, ... ? It its just POC then ... but so many times I've seen POC become production in a heartbeat.

---

<div class="post-metadata">

**Author:** ![joelatgrayv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joelatgrayv/32/141394_2.png) [@joelatgrayv](https://discuss.elastic.co/u/joelatgrayv)\
**Post date:** [February 18, 2025, 10:31pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/8 "2025-02-18T22:31:01Z")

</div>

```auto
            - Name: node.name
              Value: "es01"
            - Name: cluster.name
              value: cluster01
            - Name: cluster.initial_master_nodes
              Value: "es01,es02,es03"
            - Name: discovery.seed_hosts
              Value: "es02.elasticsearch.local,es03.elasticsearch.local"

            - Name: node.name
              Value: "es02"
            - Name: cluster.name
              value: cluster01
            - Name: cluster.initial_master_nodes
              Value: "es01,es02,es03"
            - Name: discovery.seed_hosts
              Value: "es01.elasticsearch.local,es03.elasticsearch.local"

            - Name: cluster.name
              value: cluster01
            - Name: node.name
              Value: "es03"
            - Name: cluster.initial_master_nodes
              Value: "es01,es02,es03"
            - Name: discovery.seed_hosts
              Value: "es01.elasticsearch.local,es02.elasticsearch.local"

```

For the config, I have similar. The containers run on different servers and need hostnames that ES will get the IPs.

I will add the extra security stuff if I get this working. I'm trying to reduce other potential issues. Thanks for the input.

I can't attach the logs here but they are in the SB link earlier. (the forum flagged it as spam so I can't link it again")

The nodes are on different EC2 instances. I'm going to try the network.host setting.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [February 18, 2025, 11:21pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/9 "2025-02-18T23:21:38Z")

</div>

> [@joelatgrayv](#):
>
> For the config, I have similar

Actually, you had different combinations.

For the logs, elasticsearch is verbose when starting up first time. Share those logs please.

---

<div class="post-metadata">

**Author:** ![joelatgrayv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joelatgrayv/32/141394_2.png) [@joelatgrayv](https://discuss.elastic.co/u/joelatgrayv)\
**Post date:** [February 18, 2025, 11:50pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/10 "2025-02-18T23:50:35Z")

</div>

I did. Can you access them?

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [February 19, 2025, 12:24am UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/11 "2025-02-19T00:24:17Z")

</div>

If you mean the post that shows as “This post was flagged by the community and is temporarily hidden.”, then no.

Double check your logs don’t contain anything dodgy, and try again. Or put on pastebin or similar.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 19, 2025, 8:27am UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/12 "2025-02-19T08:27:48Z")

</div>

The logs are there in the link from the OP, and say stuff like this:

```auto
1739909923928,"{""@timestamp"":""2025-02-18T20:18:43.927Z"", ""log.level"": ""WARN"", ""message"":""address [127.255.0.3:9300], node [unknown discovery result: [][127.255.0.3:9300] general node connection failure: handshake failed because connection reset; for summary, see logs from org.elasticsearch.cluster.coordination.ClusterFormationFailureHelper; for troubleshooting guidance, see https://www.elastic.co/guide/en/elasticsearch/reference/8.15/discovery-troubleshooting.html"", ""ecs.version"": ""1.2.0"",""service.name"":""ES_ECS"",""event.dataset"":""elasticsearch.server"",""process.thread.name"":""elasticsearch[es02][generic][T#2]"",""log.logger"":""org.elasticsearch.discovery.PeerFinder"",""elasticsearch.node.name"":""es02"",""elasticsearch.cluster.name"":""cluster01""}"
1739909924927,"{""@timestamp"":""2025-02-18T20:18:44.927Z"", ""log.level"": ""INFO"", ""message"":""close connection exception caught on transport layer [Netty4TcpChannel{localAddress=/127.0.0.1:55376, remoteAddress=es03.elasticsearch.local/127.255.0.3:9300, profile=default}], disconnecting from relevant node: Connection reset"", ""ecs.version"": ""1.2.0"",""service.name"":""ES_ECS"",""event.dataset"":""elasticsearch.server"",""process.thread.name"":""elasticsearch[es02][transport_worker][T#2]"",""log.logger"":""org.elasticsearch.transport.TcpTransport"",""elasticsearch.node.name"":""es02"",""elasticsearch.cluster.name"":""cluster01""}"

```

`Connection reset` means that ES opened a connection and then something outside of ES forced it to close with a RST packet. Quite possibly the RST happens in reaction to some data being sent between the nodes. That is consistent with `nc -zv es02.elasticsearch.local 9300` reporting success because this command sends no data, so it's not getting as far as the problematic step.

You need to look at your network infra to find out what is sending those RST packets to abort these inter-node connections.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [February 19, 2025, 10:50am UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/14 "2025-02-19T10:50:13Z")

</div>

> [@DavidTurner](#):
>
> The logs are there in the link from the OP

Now yes, last night no, blocked for some (likely spurious) reason.

Noting these long were from some days ago, but ...

es01 gets to the point where it logs

```auto
{"@timestamp":"2025-02-18T20:54:51.069Z","log.level":"INFO","message":"publish_address {172.30.5.223:9300}, bound_addresses {0.0.0.0:9300}","ecs.version":"1.2.0","service.name":"ES_ECS","event.dataset":"elasticsearch.server","process.thread.name":"main","log.logger":"org.elasticsearch.transport.TransportService","elasticsearch.node.name":"es01","elasticsearch.cluster.name":"cluster01"}

```

Notice the publish address is 172.30.5.223, on a private network.

But the other addresses in the 3 log files are

127.0.0.1  
and  
127.255.0.1/.2/.3 which appear to be what es01/2/3 are resolving to?

e.g. from logs like

localAddress=/127.0.0.1:57402  
remoteAddress=es03.elasticsearch.local/127.255.0.3:9300

Now I'm not a network specialist, but should traffic to 127.x.y.z even leave the host at all?

On my linux host

```auto
$ sudo lsof -i :9200
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NODE NAME
java 18591 elasticsearch 557u IPv6 1300989 0t0 TCP *:9200 (LISTEN)

$ nc -rz 127.0.0.1 9200
Connection to 127.0.0.1 9200 port [tcp/*] succeeded!

$ nc -rz 127.23.53.2 9200
Connection to 127.23.53.2 9200 port [tcp/*] succeeded!

```

All those "connections" are internal within the local host, I can change to any 127.x.y.z and I'll get exactly same.

I think part of issue here is that es01/2/3 should probably resolve to addresses like 172.30.5.223, and not 127.x.y,z, on all 3 hosts. Just a guess of course.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [February 19, 2025, 11:44am UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/15 "2025-02-19T11:44:39Z")

</div>

> [@RainTown](#):
>
> Now I'm not a network specialist, but should traffic to 127.x.y.z even leave the host at all?

You are right that these are unusual addresses, but IME cloud/container environments can do arbitrarily weird stuff to the network config so it's best not to make any assumptions about such things. The quickest and most reliable way to resolve this will be to break out `tcpdump` or Wireshark or similar and locate the source of those RST packets. Everything else is just guesswork.

It may also help to read [this section of the docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-network.html#modules-network-binding-publishing) particularly since ES seems to be running in a multi-homed context:

> It is usually a mistake to use `0.0.0.0` as a publish address on hosts with more than one network interface.

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [February 19, 2025, 12:10pm UTC](https://discuss.elastic.co/t/cluster-not-forming/374581/16 "2025-02-19T12:10:01Z")

</div>

David, I didn't make any assumptions, I merely (implicitly) asked @Joel to check the name resolutions are what _he_ expects.

And honestly, I think if they are resolving to 127.x.y.z addresses it's not right, but that isn't an assumption, it's speculation, though RFC5735 says (other RFCs have similar)

addresses within the entire 127.0.0.0/8 block _do not legitimately appear on any network anywhere_

Troubleshooting on limited information is hard, it's pretty hard with all the information to hand too. If we need take into account "arbitrarily weird stuff" it gets ...
