# Please help - ES 2.1.1 cluster randomly crashing

**URL:** <https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861>\
**Category:** Elasticsearch\
**Created:** [January 11, 2016, 10:02am UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861 "2016-01-11T10:02:55Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 11, 2016, 10:02am UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/1 "2016-01-11T10:02:55Z")

</div>

Hi,

I have an Elasticsearch cluster of 2 nodes, both running ES2.1.1, but I have had random crashes where the master node will fail, then sometime after the second node will also fail.

All nodes are clean ES2.1.1 installations, on Amazon Linux, using the RPM installer, with plugins `cloud-aws`, `kopf` and `head` installed. There is minimum configuration of the elasticsearch.yml file to enable cluster replication within AWS.

If I start the elasticsearch process back, the system runs fine for 4+ days fine then crashes again.

Honestly, I've looked through the various ES logs and I cannot see any obvious exceptions.

I'm hoping someone here can help me determine why this is happening?

---

<div class="post-metadata">

**Author:** ![msimos](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@msimos](https://discuss.elastic.co/u/msimos)\
**Post date:** [January 12, 2016, 12:40am UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/2 "2016-01-12T00:40:30Z")

</div>

Hi,

How much memory have you given Elasticsearch in /etc/sysconfig/elasticsearch in ES\_HEAP\_SIZE?

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 12, 2016, 8:54am UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/3 "2016-01-12T08:54:13Z")

</div>

@msimos I have not configured that setting at all, so it is the default. I  
have not checked yet, but I will soon and let you know.

If you know what the unconfigured default is, I think it would be that  
though.

Thanks.

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 12, 2016, 9:16am UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/4 "2016-01-12T09:16:42Z")

</div>

@msimos both the nodes crashed again.  
I set the `ES_HEAP_SIZE` to `1g`, as the nodes have 2GB RAM, and restarted both of them.

I would be thankful for any advice on how to determine why these nodes are crashing so frequently.

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 12, 2016, 3:35pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/5 "2016-01-12T15:35:06Z")

</div>

Please anyone who might be able to help, this keeps happening.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 12, 2016, 4:08pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/6 "2016-01-12T16:08:39Z")

</div>

Impossible to know. We have absolutely no information about what you are doing.

- number of indices
- number of shards
- volume
- queries (using aggs, sorting?)

Is it a demo? Or a real production server?

`1g` seems pretty low to me.

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 12, 2016, 5:08pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/7 "2016-01-12T17:08:03Z")

</div>

@dadoonet I'm aware that there is little information here, but rather than dump every log and config into here, I thought I would ask if there is a good place to start.

I'm asking for some assistance is all, on why a simple setup should fail so quietly.

What I am doing is running a fresh-installation cluster with 3 plugins installed to enable replication within the AWS environment. Replication was working, but it randomly crashes.

To answer your questions:  
-131 indicies  
-1302 shards  
-volume? its about 1,201,000 documents taking up around 1GB space on disk  
-queries? I have localised kibana 4.3.1 installations on the ES2.1.1 nodes, which are the only entities using the cluster.

I'm putting this together as pre-production at the moment, but its been extremely unstable since I started using 2.x some months ago.

This is all using fresh installations, no upgrades.

`1g` is specified per the recommendation in documentation that states ES\_HEAP\_SIZE should be 50% of available RAM, as I mentioned above.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 12, 2016, 5:43pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/8 "2016-01-12T17:43:02Z")

</div>

Too many shards for sure.

It's like running 1000 databases on a single machine with 1gb RAM.

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 12, 2016, 6:15pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/9 "2016-01-12T18:15:57Z")

</div>

Ok, I've been judging by memory and CPU usage, which has been pretty  
nominal (CPU ~20%, memory around 60-70%).

Still, I've decided to once again rebuild the cluster from scratch. I  
understand that there may have been load on the systems, but I don't see  
why they should just crash like that, I can understand them running slow or  
throwing up errors.

Would you have a recommendation for node sizing within AWS? What do you  
use? An instance type would be useful.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 12, 2016, 6:51pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/10 "2016-01-12T18:51:03Z")

</div>

I don't know why it's crashing. I assume you had some errors in logs or GC warnings...

Reduce the number of shards (1 shard is often enough) and may be reduce the number of indices. But I don't know about your use case so I don't know if it's doable of not.

If you don't want to suffer from noisy neighbors x-large instances are interesting.

Also, if you don't want to manage that by yourself, I'd recommend looking at found (elasticsearch as a service by elastic).

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 12, 2016, 6:53pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/11 "2016-01-12T18:53:45Z")

</div>

Given the amount of data you have the instance type you are using may very well be sufficient, as long as you drastically reduce the number of shards. Each shard is as David pointed out a separate Lucene index and carries with it some overhead in terms of file descriptors and memory.

I would recommend reducing both the number of indices as well as the number of shards per index. Given your heap size and data volume, aiming to have tens rather than thousands of shards in the cluster might be a good target.

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 14, 2016, 10:47am UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/12 "2016-01-14T10:47:15Z")

</div>

@dadoonet @Christian_Dahlqvist  
This is a fairly out-of-box installation, so I have not been manually administering the shards, they may have simply grown out of control for some reason.

I'm going to investigate how to limit shards, but I would appreciate if you had any advice on config settings related to this.

Thanks.

**EDIT:** I found the `index.number_of_shards` setting, I will be rebuilding with this set to `1` in the `elasticsearch.yml` file. Found the details [here](https://www.elastic.co/guide/en/elasticsearch/reference/2.1/index-modules.html#index-modules-settings).

(For anyone else wondering about this, the default is set to `5`, so if there are 131 indicies, replicated to a single mirror node, that would explain why there are 1302 shards. The last 2 are `.kibana`, I think.)  
==\> _`131x5x2=1310`_

**EDIT2:** I'm going to reduce the shard count primarily by setting the `index.number_of_shards` setting to `1`, but in addition, I'm going to investigate using less indices, as recommended by @dadoonet and @Christian_Dahlqvist.  
This second part I think I need to do on the Logstash side, by setting the `index` setting in my output config, as according to [this docs page](https://www.elastic.co/guide/en/logstash/2.0/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-index) , it defaults to `"logstash-%{+YYYY.MM.dd}"`, creating an index for each day as standard.  
I think setting this to a definite single index such as `"logstash-someindexname"`, would allow me to leave the `index.number_of_shards` at the unconfigured default of `5`.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 14, 2016, 1:29pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/13 "2016-01-14T13:29:10Z")

</div>

It is quite easy to change from daily to e.g. monthly indices in Logstash just by removing the date part of the date pattern (to "logstash-%{+[YYYY.MM](http://YYYY.MM)}"). You can also modify the template it has uploaded to Elasticsearch and specify that each new index should use 1 shard and 1 replica there or change the default values in elasticsearch.yml.

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 14, 2016, 2:47pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/14 "2016-01-14T14:47:15Z")

</div>

@dadoonet and @Christian_Dahlqvist you've helped me a lot. **_Thanks!_**

I've rebuilt my cluster, and it's now running much faster (Although I cant speak for stability yet).  
I did this by setting my elasticsearch output `index => "logstash-%{+YYYY.MM}"` to a single index, rather than let it build a new index for each day.

My updated status:  
-3 indices  
-22 shards  
-490,000 documents (currently)  
-Heap usage is between 20% and 50%

So far it looks good, but I've discovered a strange issue of duplicate events.  
I don't want to hijack this thread, so I created another [here](https://discuss.elastic.co/t/duplicate-logs-logstash-input-s3/39231) to explain the issue.

Maybe you might be able to have a look in there if you're interested.

Thanks.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [January 14, 2016, 3:03pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/15 "2016-01-14T15:03:47Z")

</div>

22 shards is still a lot on a single machine with few memory IMO.

I can inject for example 1m docs in a single index with 5 shards.  
May be you could imagine having one index (1 shard - 0 replica as you have a single node) per year instead of per month?

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 14, 2016, 3:16pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/16 "2016-01-14T15:16:59Z")

</div>

Well, I think its more like 11 shards per node at this point, split across 3 indices (One of those indices is the `.kibana` index, so its more like 10 shards per node across 2 indices.).

So its like 10 shards per machine, if I'm thinking about this correct.

Its running much faster at this point, and I think i'll want to experiment with reducing the shard count also, but I'm going with 1 change at a time to test.

---

<div class="post-metadata">

**Author:** ![anhlqn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anhlqn/32/5454_2.png) [@anhlqn](https://discuss.elastic.co/u/anhlqn)\
**Post date:** [January 14, 2016, 11:46pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/17 "2016-01-14T23:46:46Z")

</div>

Two threads with useful information about number of shards in an ES cluster

> **[Six Ways to Crash Elasticsearch
	  	 | Elastic](https://www.elastic.co/blog/found-crash-elasticsearch#too-many-shards-or-the-gazillion-shards-problem)**
>
> As much as we love Elasticsearch, at Found we've seen customers crash their clusters in numerous ways. Mostly due to simple misunderstandings and usually the fixes are fairly straightforward. In our quest to enlighten new adopters and entertain the...

  

> **[Optimizing Elasticsearch: How Many Shards per Index?](https://qbox.io/blog/optimizing-elasticsearch-how-many-shards-per-index)**
>
> A key question in the minds of most Elasticsearch users when they create an index is “How many shards should I use?" In this article, we explain the design tradeoffs and performance consequences of choosing different values for the number of shards....

---

<div class="post-metadata">

**Author:** ![plonka2000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/plonka2000/32/5077_2.png) [@plonka2000](https://discuss.elastic.co/u/plonka2000)\
**Post date:** [January 15, 2016, 10:04am UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/18 "2016-01-15T10:04:12Z")

</div>

Thanks @anhlqn, this is helpful reading.

I've taken note on some of this info.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:24pm UTC](https://discuss.elastic.co/t/please-help-es-2-1-1-cluster-randomly-crashing/38861/19 "2017-07-05T23:24:12Z")

</div>


