# Unresponsive cluster: weird fluctuating behavior

**URL:** <https://discuss.elastic.co/t/unresponsive-cluster-weird-fluctuating-behavior/37924>\
**Category:** Elasticsearch\
**Created:** [December 24, 2015, 9:14am UTC](https://discuss.elastic.co/t/unresponsive-cluster-weird-fluctuating-behavior/37924 "2015-12-24T09:14:27Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![vicvega](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vicvega/32/6825_2.png) [@vicvega](https://discuss.elastic.co/u/vicvega)\
**Post date:** [December 24, 2015, 9:14am UTC](https://discuss.elastic.co/t/unresponsive-cluster-weird-fluctuating-behavior/37924/1 "2015-12-24T09:14:27Z")

</div>

Hello  
Since a few days ago, my ES cluster is unresponsive and I noticed a weird fluctuating behavior.

If a periodically check the status, I see `number_of_pending_tasks` increasing (3 millions and more) and at the same time `unassigned_shards` decreasing (5k). And this is what I expected, but the weird thing is that, at a certain point it falls down and `unassigned_shards` goes back to 10k and `number_of_pending_tasks` back to 4k... and so on again and again...

15 minutes ago my cluster status was

```auto
{
  "cluster_name" : "sods",
  "status" : "red",
  "timed_out" : false,
  "number_of_nodes" : 5,
  "number_of_data_nodes" : 5,
  "active_primary_shards" : 17969,
  "active_shards" : 30613,
  "relocating_shards" : 0,
  "initializing_shards" : 2,
  "unassigned_shards" : 5333,
  "delayed_unassigned_shards" : 0,
  "number_of_pending_tasks" : 3018887,
  "number_of_in_flight_fetch" : 0,
  "task_max_waiting_in_queue_millis" : 19152497,
  "active_shards_percent_as_number" : 85.15911872705018
}

```

10 minutes ago

```auto
{
  "cluster_name" : "sods",
  "status" : "red",
  "timed_out" : false,
  "number_of_nodes" : 5,
  "number_of_data_nodes" : 5,
  "active_primary_shards" : 17969,
  "active_shards" : 30857,
  "relocating_shards" : 0,
  "initializing_shards" : 2,
  "unassigned_shards" : 5089, <==== decreasing, ok
  "delayed_unassigned_shards" : 0,
  "number_of_pending_tasks" : 3385989, <==== increasing, ok 
  "number_of_in_flight_fetch" : 0,
  "task_max_waiting_in_queue_millis" : 20548203,
  "active_shards_percent_as_number" : 85.83787693334817
}

```

And now

```auto
{
  "cluster_name" : "sods",
  "status" : "red",
  "timed_out" : false,
  "number_of_nodes" : 4,
  "number_of_data_nodes" : 4,
  "active_primary_shards" : 17969,
  "active_shards" : 24204,
  "relocating_shards" : 0,
  "initializing_shards" : 8,
  "unassigned_shards" : 11736, <======== up again!
  "delayed_unassigned_shards" : 0,
  "number_of_pending_tasks" : 4660, <======== fallen down!
  "number_of_in_flight_fetch" : 0,
  "task_max_waiting_in_queue_millis" : 96240,
  "active_shards_percent_as_number" : 67.33058862801825
}

```

The same thing occurred several times in the last days. Is it a right behavior?

I don't know what is going on. In the log files I see lots of `ProcessClusterEventTimeoutException`. I'm using Elasticsearch 2.0

Thanks for any advice

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 25, 2015, 9:46am UTC](https://discuss.elastic.co/t/unresponsive-cluster-weird-fluctuating-behavior/37924/2 "2015-12-25T09:46:12Z")

</div>

You have way too many shards for a cluster that size, which is most likely contributing a lot to your cluster issues. Each shard is an instance of a Lucene index and carries with it a certain amount of overhead. Having a very large number of shards can therefore be very inefficient as it ties up a lot of cluster resources.

I would recommend reducing the number of shards so that you have hundreds per node rather than thousands and see how that affects the stability and behaviour of the cluster. There is unfortunately no exact limit to the number of shards a node can handle and it will depend on your use case.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 25, 2015, 8:24pm UTC](https://discuss.elastic.co/t/unresponsive-cluster-weird-fluctuating-behavior/37924/3 "2015-12-25T20:24:38Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> You have way too many shards for a cluster that size

That's an understatement!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:28pm UTC](https://discuss.elastic.co/t/unresponsive-cluster-weird-fluctuating-behavior/37924/4 "2017-07-05T23:28:36Z")

</div>


