# Cluster goes yellow abruptly

**URL:** <https://discuss.elastic.co/t/cluster-goes-yellow-abruptly/231669>\
**Category:** Elasticsearch\
**Created:** [May 8, 2020, 7:34am UTC](https://discuss.elastic.co/t/cluster-goes-yellow-abruptly/231669 "2020-05-08T07:34:43Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Faiz\_Ahmed\_Mushtak\_H](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/faiz_ahmed_mushtak_h/32/47387_2.png) [@Faiz\_Ahmed\_Mushtak\_H](https://discuss.elastic.co/u/Faiz_Ahmed_Mushtak_H)\
**Post date:** [May 8, 2020, 7:34am UTC](https://discuss.elastic.co/t/cluster-goes-yellow-abruptly/231669/1 "2020-05-08T07:34:44Z")

</div>

We're using elasticsearch 7.2 in production and lately we've been observing our cluster going yellow quite often even though none of the nodes left the cluster!

Whenever the cluster have gone yellow, we've seen a sudden drop in the `indices.store.size_in_bytes` on the problematic node. So far it has always been a single node that behaved bad. At the same time, there were a couple of 429 rejection requests (parent circuit breaker trips). Not sure if a destabilized cluster is a cause of circuit tripping or circuit tripping is the cause of node being inaccessible (note that it doesn't look like all the data is being deleted, it just drops by 300gb or so)

Regarding the cluster setup  
We have a 8 core, 64GB machine, JVM heap size is 30GB. We have taken care of [https://github.com/elastic/elasticsearch/pull/46169](https://github.com/elastic/elasticsearch/pull/46169) as well. We make use of AWS NVME SSD's (which means we lose data if the instance is stopped, but in this case the instance was up, the node never left)

I really doubt if our ingestion rate is the problem (around 8k updates per minute). Last week we ran our indexing job which was ingesting around 1M per minute but that didn't destabilize the cluster

Our usecase involves a lot of regular updates & periodic batched deletes

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 10, 2020, 11:36pm UTC](https://discuss.elastic.co/t/cluster-goes-yellow-abruptly/231669/2 "2020-05-10T23:36:05Z")

</div>

What does the logs from the master show around that time?

---

<div class="post-metadata">

**Author:** ![Faiz\_Ahmed\_Mushtak\_H](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/faiz_ahmed_mushtak_h/32/47387_2.png) [@Faiz\_Ahmed\_Mushtak\_H](https://discuss.elastic.co/u/Faiz_Ahmed_Mushtak_H)\
**Post date:** [May 11, 2020, 5:20am UTC](https://discuss.elastic.co/t/cluster-goes-yellow-abruptly/231669/3 "2020-05-11T05:20:11Z")

</div>

@warkolm i think its a duplicate of [Shards getting marked as stale frequently causing cluster to go yellow](https://discuss.elastic.co/t/shards-getting-marked-as-stale-frequently-causing-cluster-to-go-yellow/231835)

I have been able to correlate it with the time when we are indexing / updating huge documents around 5-10mb. In GC logs I've been seeing humongous allocations

[https://github.com/elastic/elasticsearch/pull/46169](https://github.com/elastic/elasticsearch/pull/46169) is taken care of but the IHOP is still adaptive. So isn't it possible that `InitiatingHeapOccupancyPercent` may grow back to 70% and we face the same issue again?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 8, 2020, 5:34am UTC](https://discuss.elastic.co/t/cluster-goes-yellow-abruptly/231669/4 "2020-06-08T05:34:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
