# Timeouts under spiky load

**URL:** <https://discuss.elastic.co/t/timeouts-under-spiky-load/8302>\
**Category:** Elasticsearch\
**Created:** [July 3, 2012, 8:51pm UTC](https://discuss.elastic.co/t/timeouts-under-spiky-load/8302 "2012-07-03T20:51:30Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Erik\_Rose](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/erik_rose/32/2825_2.png) [@Erik\_Rose](https://discuss.elastic.co/u/Erik_Rose)\
**Post date:** [July 3, 2012, 8:51pm UTC](https://discuss.elastic.co/t/timeouts-under-spiky-load/8302/1 "2012-07-03T20:51:30Z")

</div>

How does everybody deal with timeouts under spiky load?

I have a 2-node, 5-shard ES 0.18.7 setup with a 45MB corpus. Each node has 70GB RAM and a 19GB heap size for ES. I'm using the mmapfs store.

Our workload is such that we inundate ES with hundreds of queries over a few seconds. Due to...

- the number of pending TCP connections that build up at the ESes and
- the 15-second GC pauses on such big heaps,

...a lot of those requests time out. Worse, the large number of pending connections sometimes causes the JVM to become unresponsive, similar to the situation in [https://groups.google.com/forum/#!msg/elasticsearch/fxxJG6iSVrM/3fynSV7xPyYJ](https://groups.google.com/forum/#!msg/elasticsearch/fxxJG6iSVrM/3fynSV7xPyYJ).

I can hardly be unique here. What does everyone else do? My current avenue of exploration is messing with the search threadpool configuration, making it use a fixed-size queue:

threadpool:  
search:  
type: fixed  
queue\_size: 70  
reject\_policy: caller

I'm a little fuzzy on what the "caller" reject\_policy does, but "abort" would at least return an HTTP 503, which I could catch in my app to trigger a back-off.

Is this the typical approach?

Cheers,  
Erik

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 4, 2012, 10:25am UTC](https://discuss.elastic.co/t/timeouts-under-spiky-load/8302/2 "2012-07-04T10:25:59Z")

</div>

On Tue, 2012-07-03 at 13:51 -0700, Erik Rose wrote:

> How does everybody deal with timeouts under spiky load?
> 
> I have a 2-node, 5-shard ES 0.18.7 setup with a 45MB corpus. Each node has 70GB RAM and a 19GB heap size for ES. I'm using the mmapfs store.
> 
> Our workload is such that we inundate ES with hundreds of queries over a few seconds. Due to...
> 
> - the number of pending TCP connections that build up at the ESes and
> - the 15-second GC pauses on such big heaps,

Have you assigned the user that is running elasticsearch the right to  
lock all 19GB (ulimit -l), and are you using bootstrap.mlockall?

clint

---

<div class="post-metadata">

**Author:** ![Erik\_Rose](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/erik_rose/32/2825_2.png) [@Erik\_Rose](https://discuss.elastic.co/u/Erik_Rose)\
**Post date:** [July 5, 2012, 4:28pm UTC](https://discuss.elastic.co/t/timeouts-under-spiky-load/8302/3 "2012-07-05T16:28:54Z")

</div>

> Have you assigned the user that is running elasticsearch the right to

lock all 19GB (ulimit -l), and are you using bootstrap.mlockall?

> 

Not only that, but swap isn't even enabled on that box.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [July 6, 2012, 11:13am UTC](https://discuss.elastic.co/t/timeouts-under-spiky-load/8302/4 "2012-07-06T11:13:34Z")

</div>

On Thu, 2012-07-05 at 09:28 -0700, Erik Rose wrote:

> ```
> Have you assigned the user that is running elasticsearch the
> right to  
> lock all 19GB (ulimit -l), and are you using
> bootstrap.mlockall? 
> 
> ```
> 
> Not only that, but swap isn't even enabled on that box.

Then I don't understand why you're seeing 15 second GC pauses. We have  
40GB of data in our indices, and two nodes with 36GB total, of which  
24GB is assigned to the ES heap.

Before we started using mlockall (or turning off swap), we had frequent  
long GC pauses.

Since turning off swap, we have none. It is super fast. Don't know what  
to suggest.

clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:21am UTC](https://discuss.elastic.co/t/timeouts-under-spiky-load/8302/5 "2017-07-06T03:21:17Z")

</div>


