# Elasticsearch cluster is crashing often

**URL:** <https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640>\
**Category:** Elasticsearch\
**Created:** [November 7, 2020, 10:10pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640 "2020-11-07T22:10:18Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![alx2020](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alx2020/32/78621_2.png) [@alx2020](https://discuss.elastic.co/u/alx2020)\
**Post date:** [November 7, 2020, 10:10pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/1 "2020-11-07T22:10:18Z")

</div>

We have several large clusters v1.7. Please don't ask why ☹ - we are on the way to upgrade to v7 but need some hints to make it stable.  
Recently one of the clusters started crashing several times per day and we see some strange heap behavior.  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/e/bec64fa4041c89b16cf02c59c809b1332316dbb2.png)  
Other clusters work as expected with a similar write rate.

We have:

- 2 masters
- 2 query nodes
- 50 data nodes

Every node has 30G heap (50% of the VM), ConcMarkSweepGC  
Write rate is ~200k doc/min and small number or queries.  
There are ~1500 indices with up to 16 shards.

Tried many things and nothing helped so far (different merge policies and index refresh settings). At some point, GC could not free memory and the heap start growing.

As an additional observation - `_cat` API is very slow and it's mostly impossible to get even `_cat/nodes` (on other clusters it works as well).

Thanks in advance

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 8, 2020, 10:46pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/2 "2020-11-08T22:46:31Z")

</div>

1.7 is well into [EOL](https://www.elastic.co/support/eol).

My only advice for such an old version would be to add more nodes until such point that you can upgrade.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 9, 2020, 4:43am UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/3 "2020-11-09T04:43:17Z")

</div>

As already mentioned version 1.7 is very old. I have not used it in years so do not remember it very well apart from that it was hard to control heap usage and that propagating cluster state was slow and inefficient making running large clusters with lots of shards harder than it is nowadays. I do have a few comments though:

> [@alx2020](#):
>
> 2 masters

You should always have 3 master eligible nodes as this is required for a highly available cluster. You also need to make sure that `discovery.zen.minimum_master_nodes` is set to `2` in order to avoid split brain scenarios and data loss.

> [@alx2020](#):
>
> There are ~1500 indices with up to 16 shards.

How many shards do you have in the cluster? Are these time-based? How often do you create and/or delete indices in the cluster? Are you using dynamic mappings?

If you have time-based indices and some are no longer written to it makes sense to forcemerge these down to a single segment in order to reduce heap usage.

> [@alx2020](#):
>
> As an additional observation - `_cat` API is very slow and it's mostly impossible to get even `_cat/nodes` (on other clusters it works as well).

I believe this often involves the master nodes, which could very well be overloaded given the size of the cluster and the shard count.

> [@alx2020](#):
>
> Every node has 30G heap (50% of the VM), ConcMarkSweepGC

This sounds good, but make sure you are using compressed pointers.

---

<div class="post-metadata">

**Author:** ![alx2020](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alx2020/32/78621_2.png) [@alx2020](https://discuss.elastic.co/u/alx2020)\
**Post date:** [November 9, 2020, 4:21pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/4 "2020-11-09T16:21:03Z")

</div>

We are on the way to migrate the cluster to the latest version but it takes time and we need stable cluster to do it. As mentioned there are 6 similar clusters that behave as expected and we can't figure out what happened with this one recently.

At some point, GC could not free memory, and the heap starts growing.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/9/599e3dfae77126841f15ce04600d7060e3a099ad.png)

On a healthy cluster it looks like

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/9/39744659f1c1dbaeb9fc531762a9c2ea334bc880.png)

@Christian_Dahlqvist there are ~20k shards, with ~450 shards per node.  
We have 2 dedicated masters and an additional eligible node.

> Are these time-based?

Some of them are immutable time-based and some are mutable and documents could be updated over time.

> How often do you create and/or delete indices in the cluster?

Not very often. This is a multi-tenant configuration with 13 indices per tenant and they are created/deleted when we on-board/delete tenant.

> Are you using dynamic mappings?

Yes. Some large indices use dynamic mapping.

What is strange is that even on query nodes there is a similar issue with the heap

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/9/f9c9dfb0e05d4fc73d84556a99570bac3f57d729.png)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 9, 2020, 8:42pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/5 "2020-11-09T20:42:42Z")

</div>

Just to reiterate what Christian said, this is a super old version with known issues around efficiency of handling large shard/mapping sizes.

Your best bet is to add more nodes till it stabilises and then upgrade.

---

<div class="post-metadata">

**Author:** ![alx2020](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alx2020/32/78621_2.png) [@alx2020](https://discuss.elastic.co/u/alx2020)\
**Post date:** [November 10, 2020, 6:26pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/6 "2020-11-10T18:26:44Z")

</div>

@warkolm We know that we need to upgrade but this is what we have now and we need to stabilize it first. We keep adding nodes but it's still unstable and we see the same heap behavior.  
What would be your suggestion for troubleshooting this issue to understand what is causing such GC behavior even on query nodes?  
Thanks

---

<div class="post-metadata">

**Author:** ![wangqinghuan](https://avatars.discourse-cdn.com/v4/letter/w/d26b3c/32.png) [@wangqinghuan](https://discuss.elastic.co/u/wangqinghuan)\
**Post date:** [November 14, 2020, 3:20pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/7 "2020-11-14T15:20:55Z")

</div>

We encounterred some memory issue when using version 1.7 is making some sort or aggregation operation on some fields without setting doc-values structure, which could cause memory issue at earlier version.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 12, 2020, 3:21pm UTC](https://discuss.elastic.co/t/elasticsearch-cluster-is-crashing-often/254640/8 "2020-12-12T15:21:18Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
