# Cluster stopped working but was working fine

**URL:** <https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635>\
**Category:** Elasticsearch\
**Created:** [October 23, 2018, 3:03pm UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635 "2018-10-23T15:03:47Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![grant\_donovan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/grant_donovan/32/42314_2.png) [@grant\_donovan](https://discuss.elastic.co/u/grant_donovan)\
**Post date:** [October 23, 2018, 3:03pm UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/1 "2018-10-23T15:03:48Z")

</div>

I have an Elasticsearch cluster. This was all working but now seems to have stopped working properly

The head no longer works and gives the message: cluster health: not connected

This command: [http://localhost:9200/\_cat/health](http://localhost:9200/_cat/health)  
gives this response

{"error":{"root\_cause":[{"type":"master\_not\_discovered\_exception","reason":null}],"type":"master\_not\_discovered\_exception","reason":null},"status":503}

When I try to write I get: ReadTimeoutError(HTTPConnectionPool(...

Timeout is set to 10

Reads work fine

What would be causing the issue? No changes have been made to the servers

Your help would be appreciated

Thanks

Grant

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [October 23, 2018, 5:39pm UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/2 "2018-10-23T17:39:56Z")

</div>

Hi,

Is your cluster up and running ? Is there any master ? Is it Red, Yellow or Green ?

bye,  
Xavier

---

<div class="post-metadata">

**Author:** ![grant\_donovan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/grant_donovan/32/42314_2.png) [@grant\_donovan](https://discuss.elastic.co/u/grant_donovan)\
**Post date:** [October 24, 2018, 10:03am UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/3 "2018-10-24T10:03:07Z")

</div>

When I try this, I get the following

curl [http://localhost:9200/\_cluster/state](http://localhost:9200/_cluster/state)  
{"error":{"root\_cause":[{"type":"master\_not\_discovered\_exception","reason":null}],"type":"master\_not\_discovered\_exception","reason":null},"status":503}

I have inherited the cluster so I am trying to troubleshoot why the cluster isn't working,I at least want to report to my infrastructure team what should work.

One thing I am presuming should work is

In here: /etc/elasticsearch/elasticsearch.yml

There are references to IP addresses

discovery.zen.ping.unicast.hosts: ['XX.XXX.XXX.XXX', 'YY.YYY.YYY.YYY']

My expectation is that from the client that these should work, as in get a response

telnet XX.XXX.XXX.XXX 9200  
curl -XGET "XX.XXX.XXX.XXX:9200"

Currently I don't, and I get a unable to connect: Connection Refused message

Thanks

Grant

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [October 24, 2018, 11:33am UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/4 "2018-10-24T11:33:03Z")

</div>

Hi,

It seems that node is not connecter to the cluster. There is no master available. Your local node cannot contact masters listed in the discovery.zen.ping.unicast.hosts ? If this is the case you should read this doc:

[https://www.elastic.co/guide/en/elasticsearch/reference/current/discovery-settings.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/discovery-settings.html)

Maybe the cluster is not joinable because of a firewall or something like that.

bye,  
Xavier

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [October 24, 2018, 12:14pm UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/5 "2018-10-24T12:14:41Z")

</div>

The node's logs will contain messages (including stack traces) describing in a bit more detail why it can't find the master.

---

<div class="post-metadata">

**Author:** ![grant\_donovan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/grant_donovan/32/42314_2.png) [@grant\_donovan](https://discuss.elastic.co/u/grant_donovan)\
**Post date:** [October 25, 2018, 2:42pm UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/6 "2018-10-25T14:42:15Z")

</div>

Thanks. I had a look and there were no logs. I restarted the service and saw some errors while I was restarting (in syslog)

I corrected that and I now have some logs (good start)

{#zen\_unicast\_6\_A\_ZhFv6mT3i65uDeJUdyjA#}{10.197.163.236}{XX.XXX.XXX.XXX:9300}{master=true}]  
Oct 25 10:38:37 hcukazprocatap03 elasticsearch[28877]: [2018-10-25 10:38:37,007][WARN][transport.netty] [es-client-01] exception caught on transport layer [[id: 0x5a1fc5a8]], closing connection  
Oct 25 10:38:37 hcukazprocatap03 elasticsearch[28877]: java.net.NoRouteToHostException: No route to host

I get a No route to host message, and I presume that relate to this ip: XX.XXX.XXX.XXX:9300

When I telnet to that (telnet XX.XXX.XXX.XXX 9300) I get a connection refused.

Could you confirm that I am approaching this the right way (telnet). My current thoughts are that the port is being blocked.

At the moment I am in the position where I need to tell our infrastructure team what the issue is

Any guidance would be appreciated

Thanks

Grant

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [October 25, 2018, 5:48pm UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/7 "2018-10-25T17:48:40Z")

</div>

Yes, this sounds like connectivity issues. `telnet` is a reasonable way to test basic connectivity to an Elasticsearch node's transport port (which defaults to 9300). If you manage to establish a connection, hitting `<Enter>` a few times should close the connection and yield the following sort of log messages on Elasticsearch's side which lets you see that you've actually connected to Elasticsearch and not to something else.

```auto
[2018-10-25T18:44:23,948][WARN][o.e.x.s.t.n.SecurityNetty4ServerTransport] [p6N7aBv] exception caught on transport layer [NettyTcpChannel{localAddress=/0:0:0:0:0:0:0:1:9300, remoteAddress=/0:0:0:0:0:0:0:1:53900}], closing connection
io.netty.handler.codec.DecoderException: java.io.StreamCorruptedException: invalid internal transport message format, got (d,a,d,a)

```

---

<div class="post-metadata">

**Author:** ![grant\_donovan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/grant_donovan/32/42314_2.png) [@grant\_donovan](https://discuss.elastic.co/u/grant_donovan)\
**Post date:** [October 26, 2018, 10:50am UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/8 "2018-10-26T10:50:43Z")

</div>

Thanks for your help with this.

Having looked into this further it looked like the Elasticsearch service on the other boxes had stopped. I don't know why, I need to add some extra monitoring onto that.

All looks to be sorted now

Thanks again

Grant

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 23, 2018, 10:50am UTC](https://discuss.elastic.co/t/cluster-stopped-working-but-was-working-fine/153635/9 "2018-11-23T10:50:44Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
