# Production Cluster Suddenly Crashed Last Night

**URL:** <https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691>\
**Category:** Elasticsearch\
**Created:** [July 1, 2015, 8:18am UTC](https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691 "2015-07-01T08:18:00Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Yosi\_Haran](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yosi_haran/32/44914_2.png) [@Yosi\_Haran](https://discuss.elastic.co/u/Yosi_Haran)\
**Post date:** [July 1, 2015, 8:18am UTC](https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691/1 "2015-07-01T08:18:00Z")

</div>

Hi Guys,

We've been running a production cluster of elasticsearch 1.0.0 with 3 nodes in 3 regions in AWS for about a year now, with about ~5K requests per minute on average.  
Last night, although there was no traffic spike, the elasticsearch log started to fill up with the following exception:

```
org.elasticsearch.common.util.concurrent.EsRejectedExecutionException: rejected execution (queue capacity 1000)

```

Shortly after that the CPU of the 3 machines went up to 100% and they became inaccessible.

We restarted them and all is well for now, but we are trying to understand what happened and would greatly appreciate any insight we can get from the members of this forum.  
Just to re-iterate, there was no traffic spike.  
(We are also waiting to see if there was some network error with amazon, but that wouldn't fully explain the CPU spike).

Thanks!

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [July 1, 2015, 11:29am UTC](https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691/2 "2015-07-01T11:29:25Z")

</div>

Do you know if you got this exception as part of a search or indexing operation (the full stack trace could help figure it out)?

---

<div class="post-metadata">

**Author:** ![Yosi\_Haran](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yosi_haran/32/44914_2.png) [@Yosi\_Haran](https://discuss.elastic.co/u/Yosi_Haran)\
**Post date:** [July 1, 2015, 11:41am UTC](https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691/3 "2015-07-01T11:41:19Z")

</div>

It was a searching operation.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 2, 2015, 3:12am UTC](https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691/4 "2015-07-02T03:12:40Z")

</div>

How much data and indices + shards in your cluster.

---

<div class="post-metadata">

**Author:** ![Yosi\_Haran](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yosi_haran/32/44914_2.png) [@Yosi\_Haran](https://discuss.elastic.co/u/Yosi_Haran)\
**Post date:** [July 2, 2015, 3:05pm UTC](https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691/5 "2015-07-02T15:05:14Z")

</div>

Crash was almost 100% caused because of the AWS downtime: [http://mashable.com/2015/06/30/aws-disruption/](http://mashable.com/2015/06/30/aws-disruption/)...  
Thanks for the help anyway 😄

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:03am UTC](https://discuss.elastic.co/t/production-cluster-suddenly-crashed-last-night/24691/6 "2017-07-06T00:03:56Z")

</div>


