# ElasticSearch 5.4 Nodes unable to join cluster - Troubleshooting

**URL:** <https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536>\
**Category:** Elasticsearch\
**Created:** [May 20, 2017, 6:42pm UTC](https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536 "2017-05-20T18:42:33Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![techpanga](https://avatars.discourse-cdn.com/v4/letter/t/3ab097/32.png) [@techpanga](https://discuss.elastic.co/u/techpanga)\
**Post date:** [May 20, 2017, 6:42pm UTC](https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536/1 "2017-05-20T18:42:33Z")

</div>

Hi There,

We are running into some issues post upgrade from 2.3.x to 5.4.

OS : RHEL 7  
Java : JDK1.8.0\_111

# Server (VM) #1:

```
 node.name: node-1
network.host: xxx.xxx.197.14
cluster.name: dsinke3
node.master: true
node.data: true
path.data: /elkstore/elasticsearch
path.logs: /elkstore/logs
bootstrap.memory_lock: true
discovery.zen.ping.unicast.hosts: ["xxx.xxx.197.14:9200", "xxx.xxx.197.15:9200", "xxx.xxx.197.16:9200", "xxx.xxx.197.17:9200", "xxx.xxx.197.18:9200"]
discovery.zen.minimum_master_nodes: 2
discovery.zen.ping_timeout: 100s
discovery.zen.fd.ping_timeout: 100s
http.cors.enabled: true
http.cors.allow-origin: "*"

```

# Server (VM) #2:

```
 node.name: node-2
network.host: xxx.xxx.197.15
cluster.name: dsinke3
node.master: true
node.data: true
path.data: /elkstore/elasticsearch
path.logs: /elkstore/logs
bootstrap.memory_lock: true
discovery.zen.ping.unicast.hosts: ["xxx.xxx.197.14:9200", "xxx.xxx.197.15:9200", "xxx.xxx.197.16:9200", "xxx.xxx.197.17:9200", "xxx.xxx.197.18:9200"]
discovery.zen.minimum_master_nodes: 2
discovery.zen.ping_timeout: 100s
discovery.zen.fd.ping_timeout: 100s
http.cors.enabled: true
http.cors.allow-origin: "*"

```

# Server (VM) #3:

```
 node.name: node-3
network.host: xxx.xxx.197.16
cluster.name: dsinke3
node.master: true
node.data: true
path.data: /elkstore/elasticsearch
path.logs: /elkstore/logs
bootstrap.memory_lock: true
discovery.zen.ping.unicast.hosts: ["xxx.xxx.197.14:9200", "xxx.xxx.197.15:9200", "xxx.xxx.197.16:9200", "xxx.xxx.197.17:9200", "xxx.xxx.197.18:9200"]
discovery.zen.minimum_master_nodes: 2
discovery.zen.ping_timeout: 100s
discovery.zen.fd.ping_timeout: 100s
http.cors.enabled: true
http.cors.allow-origin: "*"

```

# Server (VM) #4:

```
 node.name: node-4
network.host: xxx.xxx.197.17
cluster.name: dsinke3
node.master: true
node.data: true
path.data: /elkstore/elasticsearch
path.logs: /elkstore/logs
bootstrap.memory_lock: true
discovery.zen.ping.unicast.hosts: ["xxx.xxx.197.14:9200", "xxx.xxx.197.15:9200", "xxx.xxx.197.16:9200", "xxx.xxx.197.17:9200", "xxx.xxx.197.18:9200"]
discovery.zen.minimum_master_nodes: 2
discovery.zen.ping_timeout: 100s
discovery.zen.fd.ping_timeout: 100s
http.cors.enabled: true
http.cors.allow-origin: "*"

```

# Server (VM) #5:

```
 node.name: node-5
network.host: xxx.xxx.197.18
cluster.name: dsinke3
node.master: true
node.data: true
path.data: /elkstore/elasticsearch
path.logs: /elkstore/logs
bootstrap.memory_lock: true
discovery.zen.ping.unicast.hosts: ["xxx.xxx.197.14:9200", "xxx.xxx.197.15:9200", "xxx.xxx.197.16:9200", "xxx.xxx.197.17:9200", "xxx.xxx.197.18:9200"]
discovery.zen.minimum_master_nodes: 2
discovery.zen.ping_timeout: 100s
discovery.zen.fd.ping_timeout: 100s
http.cors.enabled: true
http.cors.allow-origin: "*"

```

What would be the smooth starting order for this 5 nodes to be in cluster dsinke3?

I am running into these issues...

1. org.elasticsearch.cluster.block.ClusterBlockException: blocked by: [SERVICE\_UNAVAILABLE/1/state not recovered / initialized];
2. org.elasticsearch.transport.ConnectTransportException: [][xxx.xxx.197.15:9200] handshake\_timeout[1.6m]
3. [o.e.d.z.UnicastZenPing] [node-3] [6] failed to ping

Please help.

Thanks in advance, dp

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 20, 2017, 7:45pm UTC](https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536/2 "2017-05-20T19:45:47Z")

</div>

As you have 5 master eligible nodes, `minimum_master_nodes` should be set to 3, not 2. With the current configuration you could end up with a split cluster.

---

<div class="post-metadata">

**Author:** ![techpanga](https://avatars.discourse-cdn.com/v4/letter/t/3ab097/32.png) [@techpanga](https://discuss.elastic.co/u/techpanga)\
**Post date:** [May 20, 2017, 8:02pm UTC](https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536/3 "2017-05-20T20:02:42Z")

</div>

Hi Christian,

Thanks for response.

I updated the discovery zen minimum master nodes to 3.

and reduced the timeouts to 10s. I am getting the below error.

[node-4] [13] failed to ping {#zen\_unicast\_xxx.xxx.197.18:9200\_0#}{z5D73ZMaQbSqgfJcrq9t3A}{xxx.xxx.197.18}{xxx.xxx.197.18:9200}  
org.elasticsearch.transport.ConnectTransportException: [][xxx.xxx.197.18:9200] handshake\_timeout[10s]

I also see the below exception in logs...

[2017-05-20T19:55:35,600][WARN][o.e.t.n.Netty4Transport] [node-5] exception caught on transport layer [[id: 0x9445bdb5, L:/10.156.197.18:34094 - R:/10.156.197.15:9200]], closing connection  
io.netty.handler.codec.DecoderException: java.io.StreamCorruptedException: invalid internal transport message format, got (48,54,54,50)

The below log statement repeating every few sec.  
[o.e.d.z.ZenDiscovery] [node-3] not enough master nodes discovered during pinging (found [[Candidate{node={node-3}{KtGoZr7YSmeTuMHTMquQJQ}{nso4kJ5qSzag5zE2sRZ3Sw}{xxx.xxx.197.16}{xxx.xxx.197.16:9300}, clusterStateVersion=-1}]], but needed [3]), pinging again

I do see telnet at port 9200 && 9300 are good.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 20, 2017, 8:22pm UTC](https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536/4 "2017-05-20T20:22:48Z")

</div>

Unicast port should be 9300, not 9200, as this is the HTTP port.

---

<div class="post-metadata">

**Author:** ![techpanga](https://avatars.discourse-cdn.com/v4/letter/t/3ab097/32.png) [@techpanga](https://discuss.elastic.co/u/techpanga)\
**Post date:** [May 20, 2017, 8:36pm UTC](https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536/5 "2017-05-20T20:36:05Z")

</div>

Hi Christian,

I removed :9200 from

discovery.zen.ping.unicast.hosts:

Its started working.

Thanks for your help.  
-dp

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 17, 2017, 8:36pm UTC](https://discuss.elastic.co/t/elasticsearch-5-4-nodes-unable-to-join-cluster-troubleshooting/86536/6 "2017-06-17T20:36:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
