# Elastic cluster always down wihtout apparent cause

**URL:** <https://discuss.elastic.co/t/elastic-cluster-always-down-wihtout-apparent-cause/145077>\
**Category:** Elasticsearch\
**Created:** [August 20, 2018, 1:49am UTC](https://discuss.elastic.co/t/elastic-cluster-always-down-wihtout-apparent-cause/145077 "2018-08-20T01:49:29Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![haiyuancheng](https://avatars.discourse-cdn.com/v4/letter/h/5f9b8f/32.png) [@haiyuancheng](https://discuss.elastic.co/u/haiyuancheng)\
**Post date:** [August 20, 2018, 1:49am UTC](https://discuss.elastic.co/t/elastic-cluster-always-down-wihtout-apparent-cause/145077/1 "2018-08-20T01:49:29Z")

</div>

018-08-20T00:35:03,367][DEBUG][o.e.a.a.c.h.TransportClusterHealthAction] [plat-ecloud01-db-es01] timed out while retrying [cluster:monitor/health] after failure (timeout [30s])  
[2018-08-20T00:35:03,368][DEBUG][o.e.a.a.c.h.TransportClusterHealthAction] [plat-ecloud01-db-es01] no known master node, scheduling a retry  
[2018-08-20T00:35:03,368][WARN][r.suppressed] path: /\_cluster/health, params: {}  
org.elasticsearch.discovery.MasterNotDiscoveredException: null  
at org.elasticsearch.action.support.master.TransportMasterNodeAction$AsyncSingleAction$4.onTimeout(TransportMasterNodeAction.java:223) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.cluster.ClusterStateObserver$ContextPreservingListener.onTimeout(ClusterStateObserver.java:317) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.cluster.ClusterStateObserver$ObserverClusterStateListener.onTimeout(ClusterStateObserver.java:244) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.cluster.service.ClusterApplierService$NotifyTimeout.run(ClusterApplierService.java:576) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingRunnable.run(ThreadContext.java:625) [elasticsearch-6.3.0.jar:6.3.0]  
at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [?:1.8.0\_121]  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [?:1.8.0\_121]  
at java.lang.Thread.run(Thread.java:745) [?:1.8.0\_121]  
[2018-08-20T00:35:18,367][DEBUG][o.e.a.a.c.h.TransportClusterHealthAction] [plat-ecloud01-db-es01] no known master node, scheduling a retry  
[2018-08-20T00:35:18,367][DEBUG][o.e.a.a.c.h.TransportClusterHealthAction] [plat-ecloud01-db-es01] timed out while retrying [cluster:monitor/health] after failure (timeout [30s])  
[2018-08-20T00:35:18,367][WARN][r.suppressed] path: /\_cluster/health, params: {}  
org.elasticsearch.discovery.MasterNotDiscoveredException: null  
at org.elasticsearch.action.support.master.TransportMasterNodeAction$AsyncSingleAction$4.onTimeout(TransportMasterNodeAction.java:223) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.cluster.ClusterStateObserver$ContextPreservingListener.onTimeout(ClusterStateObserver.java:317) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.cluster.ClusterStateObserver$ObserverClusterStateListener.onTimeout(ClusterStateObserver.java:244) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.cluster.service.ClusterApplierService$NotifyTimeout.run(ClusterApplierService.java:576) [elasticsearch-6.3.0.jar:6.3.0]  
at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingRunnable.run(ThreadContext.java:625) [elasticsearch-6.3.0.jar:6.3.0]  
at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [?:1.8.0\_121]  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [?:1.8.0\_121]  
at java.lang.Thread.run(Thread.java:745) [?:1.8.0\_121]  
[2018-08-20T00:35:29,205][INFO][o.e.d.z.ZenDiscovery] [plat-ecloud01-db-es01] failed to send join request to master [{plat-ecloud01-db-es03}{Uj3hM1l6TIurc1x5eoM7XA}{ZD4pdJnLSkecmY  
XgsjDcuw}{10.176.140.58}{10.176.140.58:9300}{ml.machine\_memory=16658382848, ml.max\_open\_jobs=20, xpack.installed=true, ml.enabled=true}], reason [RemoteTransportException[[plat-ecloud01  
-db-es03][10.176.140.58:9300][internal:discovery/zen/join]]; nested: FailedToCommitClusterStateException[timed out while waiting for enough masters to ack sent cluster state. [1] left];  
]  
[2018-08-20T00:35:33,366][DEBUG][o.e.a.a.c.h.TransportClusterHealthAction] [plat-ecloud01-db-es01] no known master node, scheduling a retry  
[2018-08-20T00:35:33,368][DEBUG][o.e.a.a.c.h.TransportClusterHealthAction] [plat-ecloud01-db-es01] timed out

The es cluster run a perios of time, the service not available, is there any help for this ?  
Below is my configurations?  
node01:  
cluster.name: elasticsearch-group\_es\_01  
node.name: plat-ecloud01-db-es01  
node.master: true  
node.data: true  
path.data: /var/elasticsearch  
path.logs: /var/log/elasticsearch  
bootstrap.memory\_lock: True  
bootstrap.system\_call\_filter: false  
network.host: 0.0.0.0  
http.port: 9200  
http.enabled: true  
discovery.zen.ping.unicast.hosts: ["10.176.140.60", "10.176.140.61", "10.176.140.58"]  
discovery.zen.minimum\_master\_nodes: 2  
gateway.recover\_after\_nodes: 2  
gateway.expected\_nodes: 3  
gateway.recover\_after\_time: 1m  
discovery.zen.no\_master\_block: write  
discovery.zen.fd.ping\_timeout: 10s  
http.cors.enabled: true  
http.cors.allow-origin: "\*"  
http.max\_content\_length: 500mb  
indices.recovery.max\_bytes\_per\_sec: 200mb  
indices.memory.index\_buffer\_size: 20%  
xpack.security.enabled: false

node02:  
cluster.name: elasticsearch-group\_es\_01  
node.name: plat-ecloud01-db-es02  
node.master: true  
node.data: true  
path.data: /var/elasticsearch  
path.logs: /var/log/elasticsearch  
bootstrap.memory\_lock: True  
bootstrap.system\_call\_filter: false  
network.host: 0.0.0.0  
http.port: 9200  
http.enabled: true  
discovery.zen.ping.unicast.hosts: ["10.176.140.60", "10.176.140.61", "10.176.140.58"]  
discovery.zen.minimum\_master\_nodes: 2  
gateway.recover\_after\_nodes: 2  
gateway.expected\_nodes: 3  
gateway.recover\_after\_time: 1m  
discovery.zen.no\_master\_block: write  
discovery.zen.fd.ping\_timeout: 10s  
http.cors.enabled: true  
http.cors.allow-origin: "\*"  
http.max\_content\_length: 500mb  
indices.recovery.max\_bytes\_per\_sec: 200mb  
indices.memory.index\_buffer\_size: 20%  
xpack.security.enabled: false

node03:  
cluster.name: elasticsearch-group\_es\_01  
node.name: plat-ecloud01-db-es03  
node.master: true  
node.data: true  
path.data: /var/elasticsearch  
path.logs: /var/log/elasticsearch  
bootstrap.memory\_lock: true  
bootstrap.system\_call\_filter: false  
network.host: 0.0.0.0  
http.port: 9200  
http.enabled: true  
discovery.zen.ping.unicast.hosts: ["10.176.140.60", "10.176.140.61", "10.176.140.58"]  
discovery.zen.minimum\_master\_nodes: 2  
gateway.recover\_after\_nodes: 2  
gateway.expected\_nodes: 3  
gateway.recover\_after\_time: 1m  
discovery.zen.no\_master\_block: write  
discovery.zen.fd.ping\_timeout: 10s  
http.cors.enabled: true  
http.cors.allow-origin: "\*"  
http.max\_content\_length: 500mb  
indices.recovery.max\_bytes\_per\_sec: 200mb  
indices.memory.index\_buffer\_size: 20%  
xpack.security.enabled: false

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 17, 2018, 1:49am UTC](https://discuss.elastic.co/t/elastic-cluster-always-down-wihtout-apparent-cause/145077/2 "2018-09-17T01:49:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
