# My cluster always yellow

**URL:** <https://discuss.elastic.co/t/my-cluster-always-yellow/118551>\
**Category:** Elasticsearch\
**Created:** [February 6, 2018, 3:14am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551 "2018-02-06T03:14:09Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![sekaiga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sekaiga/32/27415_2.png) [@sekaiga](https://discuss.elastic.co/u/sekaiga)\
**Post date:** [February 6, 2018, 3:14am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/1 "2018-02-06T03:14:10Z")

</div>

hi,  
version: 5.1.1  
my cluster always yellow, a lot of unassinged shards. logs show lot of "master left"

[2018-02-06T05:16:04,873][WARN][o.e.i.c.IndicesClusterStateService] [node-15] [[xxx\_log\_201802][150]] marking and sending shard failed due to [shard failure, reason [primary shard [[xxx\_log\_201802][150], node[llX2oMHOT6aMFTpfjvXilg], [P], s[STARTED], a[id=GnaqByImR5y06I6bn5G0MQ]] was demoted while failing replica shard]]  
org.elasticsearch.cluster.action.shard.ShardStateAction$NoLongerPrimaryShardException: primary term [10] did not match current primary term [11]  
at org.elasticsearch.cluster.action.shard.ShardStateAction$ShardFailedClusterStateTaskExecutor.execute(ShardStateAction.java:280) ~[elasticsearch-5.1.1.jar:5.1.1]  
at org.elasticsearch.cluster.service.ClusterService.runTasksForExecutor(ClusterService.java:581) ~[elasticsearch-5.1.1.jar:5.1.1]  
at org.elasticsearch.cluster.service.ClusterService$UpdateTask.run(ClusterService.java:920) ~[elasticsearch-5.1.1.jar:5.1.1]  
at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingRunnable.run(ThreadContext.java:458) [elasticsearch-5.1.1.jar:5.1.1]  
at org.elasticsearch.common.util.concurrent.PrioritizedEsThreadPoolExecutor$TieBreakingPrioritizedRunnable.runAndClean(PrioritizedEsThreadPoolExecutor.java:238) ~[elasticsearch-5.1.1.jar:5.1.1]  
at org.elasticsearch.common.util.concurrent.PrioritizedEsThreadPoolExecutor$TieBreakingPrioritizedRunnable.run(PrioritizedEsThreadPoolExecutor.java:201) ~[elasticsearch-5.1.1.jar:5.1.1]  
at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [?:1.8.0\_92]  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [?:1.8.0\_92]  
at java.lang.Thread.run(Thread.java:745) [?:1.8.0\_92]  
[2018-02-06T05:16:05,556][INFO][o.e.d.z.ZenDiscovery] [node-15] master\_left [{node-1}{zzJVtCQDQFaNm\_jNx2YrjA}{3vrBZme0R\_aCZwqZHuoJDQ}{10.4.71.30}{10.4.71.30:9300}], reason [failed to ping, tried [3] times, each with maximum [30s] timeout]  
[2018-02-06T05:16:05,557][WARN][o.e.d.z.ZenDiscovery] [node-15] master left (reason = failed to ping, tried [3] times, each with maximum [30s] timeout), current nodes: nodes:  
{node-9}{UsskhSP7StKKpuD\_GcAcGw}{MYFs7ktOQ5KZsuaFpjW\_0g}{10.4.71.67}{10.4.71.67:9300}  
{node-10}{MaoCFIexSneolK-a9DTtgQ}{DdsgyhqbQ7ipE4rlV2LHyQ}{10.4.71.68}{10.4.71.68:9300}  
{node-4}{NPO\_-oC7RGikhqcmwc6JxA}{wGKtE55VT8uhp1mI5QaOlA}{10.4.71.33}{10.4.71.33:9300}  
{node-5}{qUDCyodwSq63x3Z3s5LlUw}{StzL-uBsTIug6gWmhzB7qA}{10.4.71.34}{10.4.71.34:9300}  
{node-8}{lTyb\_Y1CQRapvztkK-Uz1g}{wLHuYpgDQMaHisDWWslO\_Q}{10.4.71.66}{10.4.71.66:9300}  
{node-3}{KrYOYlfmRC6yS0uvNIBSHA}{X-glH3VmT-2c9\_XMQsTtpw}{10.4.71.32}{10.4.71.32:9300}  
{node-7}{ABmukxbzSQGgLJ29DBUnjA}{ppo9n1vgSw-LYq-0BO\_Dmw}{10.4.71.36}{10.4.71.36:9300}  
{node-11}{oSYHx9hgTZqNoI0gIqO6Rw}{BJa9EDLpSje4865hn0bXlw}{10.4.71.69}{10.4.71.69:9300}  
{node-6}{T7J4E5H4RqmM2DpJl4J3bA}{22uVp89gTvewPzhR6buhAA}{10.4.71.35}{10.4.71.35:9300}  
{node-14}{dPEClHXZR2iFpaiLw-I-nQ}{jkBVFbbiRs2d48jjoDAdAQ}{10.4.71.72}{10.4.71.72:9300}  
{node-13}{w82I1FcSQm6bqZAypk1P-g}{JZEc57UTQ6q90SsgV9IJRg}{10.4.71.71}{10.4.71.71:9300}  
{node-15}{llX2oMHOT6aMFTpfjvXilg}{VXe3JphuT6aCbbAVRoB9CQ}{10.4.71.73}{10.4.71.73:9300}, local  
{node-12}{VnLoFQVhTragjMlTo\_3TzA}{aAhZplHLSISNczETVnimPw}{10.4.71.70}{10.4.71.70:9300}  
{node-2}{qdhLZ9OPREiXE7L-G1GWlg}{QtCiGLrHS5eVuSYaIyhiGw}{10.4.71.31}{10.4.71.31:9300}

[2018-02-06T05:16:05,557][INFO][o.e.c.s.ClusterService] [node-15] removed {{node-1}{zzJVtCQDQFaNm\_jNx2YrjA}{3vrBZme0R\_aCZwqZHuoJDQ}{10.4.71.30}{10.4.71.30:9300},}, reason: master\_failed ({node-1}{zzJVtCQDQFaNm\_jNx2YrjA}{3vrBZme0R\_aCZwqZHuoJDQ}{10.4.71.30}{10.4.71.30:9300})  
[2018-02-06T05:16:08,830][INFO][o.e.c.s.ClusterService] [node-15] detected\_master {node-1}{zzJVtCQDQFaNm\_jNx2YrjA}{3vrBZme0R\_aCZwqZHuoJDQ}{10.4.71.30}{10.4.71.30:9300}, added {{node-1}{zzJVtCQDQFaNm\_jNx2YrjA}{3vrBZme0R\_aCZwqZHuoJDQ}{10.4.71.30}{10.4.71.30:9300},}, reason: zen-disco-receive(from master [master {node-1}{zzJVtCQDQFaNm\_jNx2YrjA}{3vrBZme0R\_aCZwqZHuoJDQ}{10.4.71.30}{10.4.71.30:9300} committed version [77351]])  
[2018-02-06T05:16:58,138][INFO][o.e.m.j.JvmGcMonitorService] [node-15] [gc][1861024] overhead, spent [433ms] collecting in the last [1s]

my cluster health :  
{  
"cluster\_name": "xxxxxxx",  
"status": "yellow",  
"timed\_out": false,  
"number\_of\_nodes": 15,  
"number\_of\_data\_nodes": 15,  
"active\_primary\_shards": 1465,  
"active\_shards": 2524,  
"relocating\_shards": 0,  
"initializing\_shards": 5,  
"unassigned\_shards": 164,  
"delayed\_unassigned\_shards": 0,  
"number\_of\_pending\_tasks": 0,  
"number\_of\_in\_flight\_fetch": 0,  
"task\_max\_waiting\_in\_queue\_millis": 0,  
"active\_shards\_percent\_as\_number": 93.72447085035277  
}

my config :  
cluster.name: xxxxxx  
node.name: node-15  
path.data: /data1/elasticsearch/data  
network.host: 0.0.0.0  
discovery.zen.ping.unicast.hosts: ["10.4.71.30", "10.4.71.31", "10.4.71.32", "10.4.71.33", "10.4.71.34", "10.4.71.35", "10.4.71.36", "10.4.71.66", "10.4.71.67", "10.4.71.68", "10.4.71.69", "10.4.71.70","10.4.71.71","10.4.71.72","10.4.71.73"]  
reindex.remote.whitelist: ["10.5.24.139:9200"]  
node.master: true  
node.data: true

bootstrap.memory\_lock: true

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 6, 2018, 4:01am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/2 "2018-02-06T04:01:38Z")

</div>

Please don't post pictures of text, they are difficult to read and some people may not be even able to see them 🙂

---

<div class="post-metadata">

**Author:** ![sekaiga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sekaiga/32/27415_2.png) [@sekaiga](https://discuss.elastic.co/u/sekaiga)\
**Post date:** [February 6, 2018, 4:28am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/3 "2018-02-06T04:28:34Z")

</div>

oh sorry . i cannot take text easy. i will post text soon.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 6, 2018, 6:49am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/4 "2018-02-06T06:49:57Z")

</div>

How much heap do you have assigned to the nodes? How many of your nodes are master eligible? Are you using the default value for the `discovery.zen.minimum_master_nodes` (should be set [as described here](https://www.elastic.co/guide/en/elasticsearch/reference/6.1/modules-node.html#split-brain))?

---

<div class="post-metadata">

**Author:** ![sekaiga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sekaiga/32/27415_2.png) [@sekaiga](https://discuss.elastic.co/u/sekaiga)\
**Post date:** [February 6, 2018, 7:02am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/5 "2018-02-06T07:02:55Z")

</div>

> How much heap do you have assigned to the nodes?

30GB heap size at each node  
i set  
-Xms30g  
-Xmx30g  
in jvm.options

> How many of your nodes are master eligible?

every node can be master. i had set this config (in my question) to all nodes except node.name is different.

> Are you using the default value for the discovery.zen.minimum\_master\_nodes

yes, i has not set discovery.zen.minimum\_master\_nodes.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 6, 2018, 7:09am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/6 "2018-02-06T07:09:06Z")

</div>

> [@sekaiga](#):
>
> yes, i has not set discovery.zen.minimum\_master\_nodes.

That is not good. Not good at all. You need to set this to the correct value as you can otherwise suffer from network partitions and data loss.

---

<div class="post-metadata">

**Author:** ![sekaiga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sekaiga/32/27415_2.png) [@sekaiga](https://discuss.elastic.co/u/sekaiga)\
**Post date:** [February 6, 2018, 7:11am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/7 "2018-02-06T07:11:38Z")

</div>

ok. i will do that immediately. thank you ~~~~ :grinning:

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 6, 2018, 7:28am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/8 "2018-02-06T07:28:22Z")

</div>

You may already have network partitions and inconsistencies within your cluster, so could potentially have conflicts and lose data when fixing this.

---

<div class="post-metadata">

**Author:** ![sekaiga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sekaiga/32/27415_2.png) [@sekaiga](https://discuss.elastic.co/u/sekaiga)\
**Post date:** [February 6, 2018, 8:01am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/9 "2018-02-06T08:01:14Z")

</div>

it doesn't matter.  
but. i had set discovery.zen.minimum\_master\_nodes : 8 (i have 15 nodes). "master left" problem is still happening.  
in 30 minutes there is 3 nodes logged "master left" and lots of unassigned shards appeared.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 6, 2018, 8:05am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/10 "2018-02-06T08:05:02Z")

</div>

If your cluster is under reasonably heavy load and suffering from long GC, you are probably better off introducing 3 smaller, dedicated master nodes as that will provide better stability and make it easier to scale out the cluster as minimum\_master\_nodes will not need to be adjusted.

---

<div class="post-metadata">

**Author:** ![sekaiga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sekaiga/32/27415_2.png) [@sekaiga](https://discuss.elastic.co/u/sekaiga)\
**Post date:** [February 6, 2018, 8:08am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/11 "2018-02-06T08:08:14Z")

</div>

ok,i will follow your suggestion , thanks a lot. 😀

---

<div class="post-metadata">

**Author:** ![sekaiga](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sekaiga/32/27415_2.png) [@sekaiga](https://discuss.elastic.co/u/sekaiga)\
**Post date:** [February 8, 2018, 8:42am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/12 "2018-02-08T08:42:06Z")

</div>

i found root cause: [ES 5.1.1: Cluster loses a node randomly every few hours. Error: Message not fully read (response) for requestId](https://discuss.elastic.co/t/es-5-1-1-cluster-loses-a-node-randomly-every-few-hours-error-message-not-fully-read-response-for-requestid/71208)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 8, 2018, 8:42am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/13 "2018-03-08T08:42:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 8, 2018, 8:42am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/14 "2018-03-08T08:42:43Z")

</div>



---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 4, 2022, 5:13am UTC](https://discuss.elastic.co/t/my-cluster-always-yellow/118551/15 "2022-11-04T05:13:00Z")

</div>


