# Es 5.5 cluster unsearcable, node stats time-outs

**URL:** <https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204>\
**Category:** Elasticsearch\
**Created:** [September 2, 2017, 10:36pm UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204 "2017-09-02T22:36:33Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![martinrm77](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/martinrm77/32/21772_2.png) [@martinrm77](https://discuss.elastic.co/u/martinrm77)\
**Post date:** [September 2, 2017, 10:36pm UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204/1 "2017-09-02T22:36:33Z")

</div>

Lately our logging elasticsearch cluster started acting up, crashing a node once in a while because of heap out of memory issues. I started putting more resources into our 6 nodes, continued upgrading from 2.4 to 5.5 and now I'm stuck with cluster that does index data, but will not answer my queries.  
Right now its 10 nodes, 8 data, 2 master. around 1080 indices of 10-50gb in 11500 shards. Its gotten big, maybe too big.  
Now it keeps timing out requests and there is a lot of node stats errors in the logs on the master:  
`org.elasticsearch.transport.ReceiveTimeoutTransportException: [elasticdb02pl][10.77.168.41:9300][cluster:monitor/nodes/stats[n]] request_id [233363] timed out after [15000ms]`

I tried putting more resources into the cluster:  
\*gave the data nodes more memory, now at 32g and 20g for heap  
\*more cpu, now at 6 vcpu was 4  
\*more nodes, was at 2 masters + 6 data, now at 2 masters + 8 data.  
It still wont answer my prayers... errr requests...

I dont know where the bottleneck is. They are all in the same VLAN with no firewalls, HW is VMware 6.0, Storage is NFS, 1 mount of 9TB per node, apprx. 65% full.  
It ran fine, until it didnt...

Any suggestions please?

/Martin

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 2, 2017, 10:59pm UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204/2 "2017-09-02T22:59:36Z")

</div>

> [@martinrm77](#):
>
> 2 master

That's bad, there is no majority of 2, see [Important Configuration Changes | Elasticsearch: The Definitive Guide [2.x] | Elastic](https://www.elastic.co/guide/en/elasticsearch/guide/2.x/important-configuration-changes.html#_minimum_master_nodes)

> [@martinrm77](#):
>
> 1080 indices of 10-50gb in 11500 shards

That's too many shards, use `_shrink` to reduce that so each index is a single shard.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 3, 2017, 6:31am UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204/3 "2017-09-03T06:31:17Z")

</div>

Based on the data you have provided you have around 6TB of data per node with an average shard size of 4GB. In order to be able to hold as much data as possible per node, you should aim for a larger shard size, ideally somewhere around 20-30GB in size. Use the shrink API to reduce the number of primary shards to 1 as Mark suggested. Start with the smallest indices. You can also reduce overhead by running a [force merge down to a few segments](https://www.elastic.co/guide/en/elasticsearch/reference/5.5/indices-forcemerge.html) (is I/O intensive) once the shrink operation has completed. Only do this for indices no longer being indexed into.

I would also recommend installing X-Pack Monitoring if you have not already, as this will give you a better idea about what is going on.

Also add another master node as per Mark's suggestion. You always want 3 dedicated master nodes so that you can lose one and the remaining ones are able to form a majority and elect a master.

---

<div class="post-metadata">

**Author:** ![martinrm77](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/martinrm77/32/21772_2.png) [@martinrm77](https://discuss.elastic.co/u/martinrm77)\
**Post date:** [September 3, 2017, 6:46am UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204/4 "2017-09-03T06:46:53Z")

</div>

What about new indices? I have it set to 8 now, to split new indices over all 8 data nodes to get max indexing performance.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 3, 2017, 6:53am UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204/5 "2017-09-03T06:53:47Z")

</div>

If you need to have 8 primary shards I would recommend having each index cover a longer time period, e.g. a week or a month. It is also possible to index into 8 primary shards and then use the shrink API to reduce the number of primary shards when it is no longer indexed into.

You can also use the [rollover API](https://www.elastic.co/blog/managing-time-based-indices-efficiently) to cut indices based on size rather than time to ensure you don't end up with too many small shards, which is inefficient. If you have a long retention period, all indices does not necessarily need to cover the same amount of time.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 3, 2017, 9:06am UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204/6 "2017-09-03T09:06:14Z")

</div>

Given your volume size, you don't really need to worry about index performance at this point.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 1, 2017, 9:06am UTC](https://discuss.elastic.co/t/es-5-5-cluster-unsearcable-node-stats-time-outs/99204/7 "2017-10-01T09:06:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
