# Elasticsearch cpu spike, search thread pool queues explode

**URL:** <https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950>\
**Category:** Elasticsearch\
**Created:** [November 8, 2018, 7:43pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950 "2018-11-08T19:43:06Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![bitsofinfo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bitsofinfo/32/459_2.png) [@bitsofinfo](https://discuss.elastic.co/u/bitsofinfo)\
**Post date:** [November 8, 2018, 7:43pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/1 "2018-11-08T19:43:06Z")

</div>

Hi not looking for a definitive answer here but a bit puzzled as to what is going on or some pointers on config settings to look at

Running ES 6.3

10 node cluster, 5 warm, 5 hot

Indicies:

- logstash indexes, being rotated every few hours, indexing rate ~2500/s, 5 shards, 1 replica. Search rate ~1-3 searches a second

- one unrelated app index ~300mb total in size 5 shards, 1 replica. Indexing rate 0-1 a second, search rate 1-2 searches a second

All requests (read/write) hit hot nodes

Every once in a while, all search queues (14 threads, 1000 queue size) on all hot nodes fill up and start rejecting requests. During this period of time when looking at the elasticsearch overall and per node metrics we see CPU spiking well over 100% on all nodes. Only after we stop indexing and/or reduce the number of searches going on does the cluster eventually slow down

There is nothing in the slow search logs on any node during this time period

The only thing we can see is that this tends to occur 5-10m after logstash stops writing to a logstash index, creates a new one and starts writing there.... perhaps while concurrently one or more searches are in process (impossible to tell) as we have zero insight into the search queue contents.

The same metrics do NOT show any increase in the search rate during this period, yet the search queues are maxed out.

This feels like some sort of blocking operation preventing searches is going on across the cluster during this period of time but cannot figure out what that would be. None of the known searches being executed take more than 500ms to max 3s to return typically. During this same period, watchers fail with timeouts as well.

Not really sure where to be begin looking. Any ideas appreciated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 8, 2018, 7:49pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/2 "2018-11-08T19:49:51Z")

</div>

How many indices and shards do you have in the cluster? What is the average shard size? How many of these are on the hot nodes?

---

<div class="post-metadata">

**Author:** ![bitsofinfo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bitsofinfo/32/459_2.png) [@bitsofinfo](https://discuss.elastic.co/u/bitsofinfo)\
**Post date:** [November 8, 2018, 7:50pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/3 "2018-11-08T19:50:43Z")

</div>

> [@bitsofinfo](#):
>
> Indicies:
> 
> - logstash indexes, being rotated every few hours, indexing rate ~2500/s, 5 shards, 1 replica. Search rate ~1-3 searches a second
> - one unrelated app index ~300mb total in size 5 shards, 1 replica. Indexing rate 0-1 a second, search rate 1-2 searches a second

All of the above indicies have one shard on each of the nodes described. Logstash indicies are many, but one one is actively written to.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 8, 2018, 7:52pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/4 "2018-11-08T19:52:14Z")

</div>

That is not what I asked. What is the output of the [cluster health API](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-health.html)?

---

<div class="post-metadata">

**Author:** ![bitsofinfo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bitsofinfo/32/459_2.png) [@bitsofinfo](https://discuss.elastic.co/u/bitsofinfo)\
**Post date:** [November 8, 2018, 7:55pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/5 "2018-11-08T19:55:54Z")

</div>

6.1TB total data  
231 indexes  
about 195 per hot node

{  
"cluster\_name" : "jiblet8",  
"status" : "green",  
"timed\_out" : false,  
"number\_of\_nodes" : 13,  
"number\_of\_data\_nodes" : 10,  
"active\_primary\_shards" : 955,  
"active\_shards" : 1940,  
"relocating\_shards" : 0,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 0,  
"delayed\_unassigned\_shards" : 0,  
"number\_of\_pending\_tasks" : 0,  
"number\_of\_in\_flight\_fetch" : 0,  
"task\_max\_waiting\_in\_queue\_millis" : 0,  
"active\_shards\_percent\_as\_number" : 100.0  
}

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 8, 2018, 8:02pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/6 "2018-11-08T20:02:37Z")

</div>

How many of these 1940 shards are on the hot nodes?

---

<div class="post-metadata">

**Author:** ![bitsofinfo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bitsofinfo/32/459_2.png) [@bitsofinfo](https://discuss.elastic.co/u/bitsofinfo)\
**Post date:** [November 8, 2018, 8:02pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/7 "2018-11-08T20:02:59Z")

</div>

~193 per node

---

<div class="post-metadata">

**Author:** ![bitsofinfo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bitsofinfo/32/459_2.png) [@bitsofinfo](https://discuss.elastic.co/u/bitsofinfo)\
**Post date:** [November 9, 2018, 3:19pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/8 "2018-11-09T15:19:10Z")

</div>

Any thoughts?

---

<div class="post-metadata">

**Author:** ![haifidelity](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/haifidelity/32/33436_2.png) [@haifidelity](https://discuss.elastic.co/u/haifidelity)\
**Post date:** [November 13, 2018, 3:30am UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/9 "2018-11-13T03:30:40Z")

</div>

For what it's worth, we are experiencing the same thing from a cluster upgraded from 5.6.x to 6.3.1. No issues in 5.6, but start seeing the active search queue fill up then start rejecting subsequent queries after a few days. The health api responds, but all other API's and cluster responsiveness times out, requiring us to do a complete cluster restart.

---

<div class="post-metadata">

**Author:** ![wrinehart](https://avatars.discourse-cdn.com/v4/letter/w/ba9def/32.png) [@wrinehart](https://discuss.elastic.co/u/wrinehart)\
**Post date:** [November 14, 2018, 4:06pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/10 "2018-11-14T16:06:04Z")

</div>

We are also seeing this same behavior on a cluster that was recently upgraded to 6.4.2. The search queue grows and grows until it hits the queue capacity and then the cluster locks up. The cluster responds to health checks but is otherwise unsearchable/indexable. Only a full cluster restart fixes this.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 12, 2018, 4:18pm UTC](https://discuss.elastic.co/t/elasticsearch-cpu-spike-search-thread-pool-queues-explode/155950/11 "2018-12-12T16:18:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
