# Need help to overcome 100% CPU

**URL:** <https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126>\
**Category:** Elasticsearch\
**Created:** [April 12, 2017, 9:29am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126 "2017-04-12T09:29:27Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 12, 2017, 9:29am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/1 "2017-04-12T09:29:27Z")

</div>

Hi, everybody I`m stuck on cluster performance tuning/scale and need help.

I'm implementing solution on top of Elasticsearch. Now I`ve started load testing and 5 parallel thread requests put my cluster down.  
In short, I have this configuration:

1 node - all roles, 2x6 cores, 64 GB RAM.  
ES\_HEAP\_SIZE=30g  
bootstrap.mlockall: true  
indices.fielddata.cache.size: 20%  
network.tcp.blocking: true  
Others by default

My main index is now 5+ million documents and about 80 GB.  
For the future expansion, it`s laid out for 12 shards.

The basic query is quite heavy. It filters nothing but 2 types (now is only 2 of them) but runs several (5-7) aggregations on the whole set of documents. With 1 thread query time is acceptable, about 350-700 mils. But in a multithreaded test mode CPU immediately flies up to 100%

In hot\_thread i see  
100.1% (500.4ms out of 500ms) cpu usage by thread 'elasticsearch[node-4][search][T#9]  
94.7% (473.6ms out of 500ms) cpu usage by thread 'elasticsearch[node-4][search][T#15]  
92.8% (463.9ms out of 500ms) cpu usage by thread 'elasticsearch[node-4][search][T#25]  
(Can provide more details if needed)  
And even EsRejectedExecutionException in els.log

If I profile the query, I see that most of the time (and apparently CPU) costs goes for the aggregations.  
"took": **462** ,  
.....  
"query": [  
{  
"query\_type": "ConstantScoreQuery",  
"lucene": "ConstantScore((ConstantScore(\_type:bidutp) ConstantScore(\_type:prgos))~1)",  
"time": **"81.2** 6062800ms",  
.....  
"name": "MultiCollector",  
"reason": "search\_multi",  
"time": " **348.6** 118540ms",  
(Can post full if needed)

So now we have came to questions.  
What am I doing wrong?  
Is it the meter of shards count, or i should reduce the heap, or monitor GC?  
Maybe investigate some more?  
I will for sure add some nodes to my cluster (2-3, i don't have tons of them in my pocket) , but I need to understand whether this will be enough.

Will be grateful for any advice to help

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 12, 2017, 3:00pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/2 "2017-04-12T15:00:29Z")

</div>

Guys, please give me some hint to dig further

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 14, 2017, 3:04pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/3 "2017-04-14T15:04:21Z")

</div>

I'v added 2 nodes with 8 cores per one. Now it handles 9 reqs/sec till 100% cpu. It seems i need tuning more then HW expansion. But i dont know what to tune (((

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [April 14, 2017, 3:12pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/4 "2017-04-14T15:12:37Z")

</div>

Can you share the output of the nodes hot threads API under load?

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 14, 2017, 3:23pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/5 "2017-04-14T15:23:32Z")

</div>

In 2 hours. On my way home now

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 14, 2017, 5:45pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/6 "2017-04-14T17:45:13Z")

</div>

Hello again! Gist of hot [https://gist.github.com/anonymous/8b0da9373cdf0a6e87be67fb46c9419c](https://gist.github.com/anonymous/8b0da9373cdf0a6e87be67fb46c9419c)

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 14, 2017, 5:52pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/7 "2017-04-14T17:52:16Z")

</div>

And query [https://gist.github.com/anonymous/919b5d2059d6543f9475ebd68dc2184b](https://gist.github.com/anonymous/919b5d2059d6543f9475ebd68dc2184b)

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 14, 2017, 5:57pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/8 "2017-04-14T17:57:46Z")

</div>

I'm on 2.4.4 if metters

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 17, 2017, 11:11am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/10 "2017-04-17T11:11:03Z")

</div>

Now i have nginx+post cache in front of elastic. It helps a bit against dummy F5-s on main page, but i still have trouble with cpu query cost. Can someone help?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 17, 2017, 11:15am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/11 "2017-04-17T11:15:15Z")

</div>

What does disk I/O and iowait look like? What type of storage do you have?

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 17, 2017, 11:31am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/12 "2017-04-17T11:31:13Z")

</div>

I do not see any changes in the disk load. It is less then 5% regardless of my tests. I think my 30GB/node cache prevents the load to reach the disk level.

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 17, 2017, 11:34am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/13 "2017-04-17T11:34:26Z")

</div>

![](https://us1.discourse-cdn.com/elastic/original/3X/f/5/f5960a7742964d926a6c67e25917e45c8e70e690.png)  
CPU under test (not most havy)

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 17, 2017, 11:36am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/14 "2017-04-17T11:36:11Z")

</div>

And Disk utilization at same time

 ![](https://us1.discourse-cdn.com/elastic/original/3X/7/3/73f710ebcc96713ecd479753c512e72390e31ec3.png)

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 17, 2017, 2:34pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/15 "2017-04-17T14:34:21Z")

</div>

Could hashed fields be handy for my cardinality aggregation?

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 20, 2017, 7:50am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/16 "2017-04-20T07:50:54Z")

</div>

If somebody interested Im still facing the problem. Any advice please?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [April 20, 2017, 1:50pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/17 "2017-04-20T13:50:19Z")

</div>

Hot threads suggest your node is just busy running aggregations, I can't think of ways to speed this up significantly. I think you would just need to add more processing power. One thing surprised me from the hot threads: you seem to be using the `niofs` directory, did you opt in for it explicitly? Switching to `mmapfs` might help read directly into the FS cache rather than copying memory from the FS cache to Java. But I don't expect it to bring significant speedups.

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 20, 2017, 4:00pm UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/18 "2017-04-20T16:00:12Z")

</div>

Adrien, thanks for your reply! Now after several days of research, I also think so. I hoped only that I missed something in the configurations. I will discuss your remark about FS with our OS administrators. Thanks again.

---

<div class="post-metadata">

**Author:** ![Gang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gang/32/15413_2.png) [@Gang](https://discuss.elastic.co/u/Gang)\
**Post date:** [April 27, 2017, 9:00am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/19 "2017-04-27T09:00:59Z")

</div>

Can someone tell about niofs to mmapfs switсh procedure. Is it dynamic or I`ll need to open/close or rebuild my index?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 25, 2017, 9:02am UTC](https://discuss.elastic.co/t/need-help-to-overcome-100-cpu/82126/20 "2017-05-25T09:02:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
