# Sudden Unexplained CPU Usage

**URL:** <https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071>\
**Category:** Elasticsearch\
**Created:** [December 13, 2012, 10:05pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071 "2012-12-13T22:05:05Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 13, 2012, 10:05pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/1 "2012-12-13T22:05:05Z")

</div>

Back with more issues. Periodically, and seemingly inexplicably, the whole  
cluster becomes essentially unusable for what's sometimes hours. There's  
nothing in the logs (or slow log) to indicate what's up, and `hot_threads`  
([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines has been more than a  
little vague. Each of the machines' CPU usage gets pegged at 100% of all  
cores, and then gradually the different machines back off. From paramedic:

[https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)

I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [December 14, 2012, 4:25am UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/2 "2012-12-14T04:25:06Z")

</div>

Ouch, that's a lot of green. How's your JVM/GC doing when this is  
happening? Have you tried looking at the thread dump? What about all the  
other system/ES metrics?

## Otis

ELASTICSEARCH Performance Monitoring -  
[Elasticsearch - Sematext Documentation](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)

On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:

> Back with more issues. Periodically, and seemingly inexplicably, the whole  
> cluster becomes essentially unusable for what's sometimes hours. There's  
> nothing in the logs (or slow log) to indicate what's up, and `hot_threads` (  
> [https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines has been more than a  
> little vague. Each of the machines' CPU usage gets pegged at 100% of all  
> cores, and then gradually the different machines back off. From paramedic:
> 
> [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> 
> I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 14, 2012, 10:34pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/3 "2012-12-14T22:34:24Z")

</div>

Thanks for the reply,

So, there is a certain amount of garbage collection going on, though not  
enough to seem to explain this. The stack traces seem to indicate much the  
same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)). Something that's  
puzzling to me is that it seems to be doing a fair amount of reading /  
writing from disk, but the partition with our data is a RAID0 across four  
ephemeral drives. `iotop` is claiming only a few MBps of read/write, but  
I've easily seen it hit 100MBps.

Here's also a full picture from paramedic:

[https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)

On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic wrote:

> Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> happening? Have you tried looking at the thread dump? What about all the  
> other system/ES metrics?
> 
> ## Otis
> 
> ELASTICSEARCH Performance Monitoring -  
> [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> 
> On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> 
> > Back with more issues. Periodically, and seemingly inexplicably, the  
> > whole cluster becomes essentially unusable for what's sometimes hours.  
> > There's nothing in the logs (or slow log) to indicate what's up, and  
> > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines has  
> > been more than a little vague. Each of the machines' CPU usage gets pegged  
> > at 100% of all cores, and then gradually the different machines back off.  
> > From paramedic:
> > 
> > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > 
> > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [December 15, 2012, 4:14am UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/4 "2012-12-15T04:14:41Z")

</div>

Hi,

I can't quite tell what is going on from the screenshot.... everything  
looks super small and uniform... ☹

If you like paramedic you may also like

> **[Elasticsearch - Sematext Documentation](https://sematext.com/docs/integration/elasticsearch-integration/)**
>
> Collect and monitor key Elasticsearch metrics such as request latency, indexing rate, and segment merges with built-in anomaly detection, threshold, and heartbeat alerts. Send notifications to email and various chatops messaging services, correlate...

But yes, I looked at those stack traces and it looks like a lot of disk  
reading. Maybe iotop is lying. dstat is nice, as is iostat and vmstat.  
How big is your index, how much RAM have you got? What sort of queries are  
you serving? Are they very diverse? Can you show a bit of vmstat 2 output  
or disk IO graph from SPM?  
Have you tried using MMapDirectory?  
See [Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)

## Otis

ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:

> Thanks for the reply,
> 
> So, there is a certain amount of garbage collection going on, though not  
> enough to seem to explain this. The stack traces seem to indicate much the  
> same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)). Something  
> that's puzzling to me is that it seems to be doing a fair amount of reading  
> / writing from disk, but the partition with our data is a RAID0 across four  
> ephemeral drives. `iotop` is claiming only a few MBps of read/write, but  
> I've easily seen it hit 100MBps.
> 
> Here's also a full picture from paramedic:
> 
> [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> 
> On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic wrote:
> 
> > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > happening? Have you tried looking at the thread dump? What about all the  
> > other system/ES metrics?
> > 
> > ## Otis
> > 
> > ELASTICSEARCH Performance Monitoring -  
> > [Elasticsearch - Sematext Documentation](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > 
> > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > 
> > > Back with more issues. Periodically, and seemingly inexplicably, the  
> > > whole cluster becomes essentially unusable for what's sometimes hours.  
> > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines has  
> > > been more than a little vague. Each of the machines' CPU usage gets pegged  
> > > at 100% of all cores, and then gradually the different machines back off.  
> > > From paramedic:
> > > 
> > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > 
> > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [December 17, 2012, 3:45pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/5 "2012-12-17T15:45:25Z")

</div>

Dan,

Which version of java are you using? What you are describing looks  
like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)

I would recommend upgrading to the latest java 7 release.

Igor

On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic wrote:

> Hi,
> 
> I can't quite tell what is going on from the screenshot.... everything  
> looks super small and uniform... ☹
> 
> If you like paramedic you may also like  
> [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> 
> But yes, I looked at those stack traces and it looks like a lot of disk  
> reading. Maybe iotop is lying. dstat is nice, as is iostat and vmstat.  
> How big is your index, how much RAM have you got? What sort of queries  
> are you serving? Are they very diverse? Can you show a bit of vmstat 2  
> output or disk IO graph from SPM?  
> Have you tried using MMapDirectory? See  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> 
> ## Otis
> 
> ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> 
> > Thanks for the reply,
> > 
> > So, there is a certain amount of garbage collection going on, though not  
> > enough to seem to explain this. The stack traces seem to indicate much the  
> > same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)). Something  
> > that's puzzling to me is that it seems to be doing a fair amount of reading  
> > / writing from disk, but the partition with our data is a RAID0 across four  
> > ephemeral drives. `iotop` is claiming only a few MBps of read/write, but  
> > I've easily seen it hit 100MBps.
> > 
> > Here's also a full picture from paramedic:
> > 
> > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > 
> > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic wrote:
> > 
> > > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > > happening? Have you tried looking at the thread dump? What about all the  
> > > other system/ES metrics?
> > > 
> > > ## Otis
> > > 
> > > ELASTICSEARCH Performance Monitoring -  
> > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > 
> > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > 
> > > > Back with more issues. Periodically, and seemingly inexplicably, the  
> > > > whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines has  
> > > > been more than a little vague. Each of the machines' CPU usage gets pegged  
> > > > at 100% of all cores, and then gradually the different machines back off.  
> > > > From paramedic:
> > > > 
> > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > 
> > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 17, 2012, 4:07pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/6 "2012-12-17T16:07:51Z")

</div>

Hey guys, thanks for the replies.

The queries are pretty straightforward, and relatively diverse in terms of  
the content they're after. That said, most of the queries are limited to  
query string, and then limiting date ranges. I've not tried MMapDirectory,  
though from that documentation it sounds like the best interface is  
automatically chosen?

WRT the java version, I can try 7, as we're apparently running 6.

On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:

> Dan,
> 
> Which version of java are you using? What you are describing looks  
> like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> 
> I would recommend upgrading to the latest java 7 release.
> 
> Igor
> 
> On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic wrote:
> 
> > Hi,
> > 
> > I can't quite tell what is going on from the screenshot.... everything  
> > looks super small and uniform... ☹
> > 
> > If you like paramedic you may also like  
> > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > 
> > But yes, I looked at those stack traces and it looks like a lot of disk  
> > reading. Maybe iotop is lying. dstat is nice, as is iostat and vmstat.  
> > How big is your index, how much RAM have you got? What sort of queries  
> > are you serving? Are they very diverse? Can you show a bit of vmstat 2  
> > output or disk IO graph from SPM?  
> > Have you tried using MMapDirectory? See  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > 
> > ## Otis
> > 
> > ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > 
> > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > 
> > > Thanks for the reply,
> > > 
> > > So, there is a certain amount of garbage collection going on, though not  
> > > enough to seem to explain this. The stack traces seem to indicate much the  
> > > same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)). Something  
> > > that's puzzling to me is that it seems to be doing a fair amount of reading  
> > > / writing from disk, but the partition with our data is a RAID0 across four  
> > > ephemeral drives. `iotop` is claiming only a few MBps of read/write, but  
> > > I've easily seen it hit 100MBps.
> > > 
> > > Here's also a full picture from paramedic:
> > > 
> > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > 
> > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic wrote:
> > > 
> > > > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > > > happening? Have you tried looking at the thread dump? What about all the  
> > > > other system/ES metrics?
> > > > 
> > > > ## Otis
> > > > 
> > > > ELASTICSEARCH Performance Monitoring -  
> > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > 
> > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > > 
> > > > > Back with more issues. Periodically, and seemingly inexplicably, the  
> > > > > whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines has  
> > > > > been more than a little vague. Each of the machines' CPU usage gets pegged  
> > > > > at 100% of all cores, and then gradually the different machines back off.  
> > > > > From paramedic:
> > > > > 
> > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > 
> > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 17, 2012, 11:26pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/7 "2012-12-17T23:26:01Z")

</div>

I've now switched to mmapfs, and java 7, and still no luck. At least, we've  
been able to reproduce the problem. It always happens when people query  
elasticsearch :-/

We've had limited success with a handful of concurrent users using our API  
internally, but after more than just a couple, it becomes almost completely  
unresponsive (multi-minute search times).

We have about 440M docs across about a dozen indexes, on 27 m1.xlarge  
instances, each configured with a RAID0 across 4 ephemeral drives. It's  
both confusing and frustrating to not understand why we're not getting  
better query performance.

On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:

> Hey guys, thanks for the replies.
> 
> The queries are pretty straightforward, and relatively diverse in terms of  
> the content they're after. That said, most of the queries are limited to  
> query string, and then limiting date ranges. I've not tried MMapDirectory,  
> though from that documentation it sounds like the best interface is  
> automatically chosen?
> 
> WRT the java version, I can try 7, as we're apparently running 6.
> 
> On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> 
> > Dan,
> > 
> > Which version of java are you using? What you are describing looks  
> > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > 
> > I would recommend upgrading to the latest java 7 release.
> > 
> > Igor
> > 
> > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic wrote:
> > 
> > > Hi,
> > > 
> > > I can't quite tell what is going on from the screenshot.... everything  
> > > looks super small and uniform... ☹
> > > 
> > > If you like paramedic you may also like  
> > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > 
> > > But yes, I looked at those stack traces and it looks like a lot of disk  
> > > reading. Maybe iotop is lying. dstat is nice, as is iostat and vmstat.  
> > > How big is your index, how much RAM have you got? What sort of queries  
> > > are you serving? Are they very diverse? Can you show a bit of vmstat 2  
> > > output or disk IO graph from SPM?  
> > > Have you tried using MMapDirectory? See  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > 
> > > ## Otis
> > > 
> > > ELASTICSEARCH Performance Monitoring -  
> > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > 
> > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > 
> > > > Thanks for the reply,
> > > > 
> > > > So, there is a certain amount of garbage collection going on, though  
> > > > not enough to seem to explain this. The stack traces seem to indicate much  
> > > > the same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)).  
> > > > Something that's puzzling to me is that it seems to be doing a fair amount  
> > > > of reading / writing from disk, but the partition with our data is a RAID0  
> > > > across four ephemeral drives. `iotop` is claiming only a few MBps of  
> > > > read/write, but I've easily seen it hit 100MBps.
> > > > 
> > > > Here's also a full picture from paramedic:
> > > > 
> > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > 
> > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic wrote:
> > > > 
> > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > > > > happening? Have you tried looking at the thread dump? What about all the  
> > > > > other system/ES metrics?
> > > > > 
> > > > > ## Otis
> > > > > 
> > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > 
> > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > > > 
> > > > > > Back with more issues. Periodically, and seemingly inexplicably, the  
> > > > > > whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines  
> > > > > > has been more than a little vague. Each of the machines' CPU usage gets  
> > > > > > pegged at 100% of all cores, and then gradually the different machines back  
> > > > > > off. From paramedic:
> > > > > > 
> > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > 
> > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [December 18, 2012, 7:18pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/8 "2012-12-18T19:18:36Z")

</div>

Are you getting the same stack traces in hot threads?

On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:

> I've now switched to mmapfs, and java 7, and still no luck. At least,  
> we've been able to reproduce the problem. It always happens when people  
> query elasticsearch :-/
> 
> We've had limited success with a handful of concurrent users using our API  
> internally, but after more than just a couple, it becomes almost completely  
> unresponsive (multi-minute search times).
> 
> We have about 440M docs across about a dozen indexes, on 27 m1.xlarge  
> instances, each configured with a RAID0 across 4 ephemeral drives. It's  
> both confusing and frustrating to not understand why we're not getting  
> better query performance.
> 
> On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> 
> > Hey guys, thanks for the replies.
> > 
> > The queries are pretty straightforward, and relatively diverse in terms  
> > of the content they're after. That said, most of the queries are limited to  
> > query string, and then limiting date ranges. I've not tried MMapDirectory,  
> > though from that documentation it sounds like the best interface is  
> > automatically chosen?
> > 
> > WRT the java version, I can try 7, as we're apparently running 6.
> > 
> > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > 
> > > Dan,
> > > 
> > > Which version of java are you using? What you are describing looks  
> > > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > 
> > > I would recommend upgrading to the latest java 7 release.
> > > 
> > > Igor
> > > 
> > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic wrote:
> > > 
> > > > Hi,
> > > > 
> > > > I can't quite tell what is going on from the screenshot.... everything  
> > > > looks super small and uniform... ☹
> > > > 
> > > > If you like paramedic you may also like  
> > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > 
> > > > But yes, I looked at those stack traces and it looks like a lot of disk  
> > > > reading. Maybe iotop is lying. dstat is nice, as is iostat and vmstat.  
> > > > How big is your index, how much RAM have you got? What sort of queries  
> > > > are you serving? Are they very diverse? Can you show a bit of vmstat 2  
> > > > output or disk IO graph from SPM?  
> > > > Have you tried using MMapDirectory? See  
> > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > 
> > > > ## Otis
> > > > 
> > > > ELASTICSEARCH Performance Monitoring -  
> > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > 
> > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > 
> > > > > Thanks for the reply,
> > > > > 
> > > > > So, there is a certain amount of garbage collection going on, though  
> > > > > not enough to seem to explain this. The stack traces seem to indicate much  
> > > > > the same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)).  
> > > > > Something that's puzzling to me is that it seems to be doing a fair amount  
> > > > > of reading / writing from disk, but the partition with our data is a RAID0  
> > > > > across four ephemeral drives. `iotop` is claiming only a few MBps of  
> > > > > read/write, but I've easily seen it hit 100MBps.
> > > > > 
> > > > > Here's also a full picture from paramedic:
> > > > > 
> > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > 
> > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic  
> > > > > wrote:
> > > > > 
> > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > > > > > happening? Have you tried looking at the thread dump? What about all the  
> > > > > > other system/ES metrics?
> > > > > > 
> > > > > > ## Otis
> > > > > > 
> > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > 
> > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > > > > 
> > > > > > > Back with more issues. Periodically, and seemingly inexplicably, the  
> > > > > > > whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines  
> > > > > > > has been more than a little vague. Each of the machines' CPU usage gets  
> > > > > > > pegged at 100% of all cores, and then gradually the different machines back  
> > > > > > > off. From paramedic:
> > > > > > > 
> > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > 
> > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 18, 2012, 7:27pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/9 "2012-12-18T19:27:13Z")

</div>

The stack traces from jstack and hot threads are essentially the same.  
jstack's quite a bit more verbose, but effectively the same.

The hot threads now (using mmapfs) no longer say anything about nio, but  
the profile is very similar in terms of how much CPU is listed as getting  
used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've been watching  
iostat, because even after we submit some basic queries and it gets into  
this state, it remains in a bad way for quite a while. What we're seeing is  
the io ops (included in gist) stay relatively high during these times  
despite the throughput (sounds a little like thrashing to me). Could this  
be symptomatic of the number of indexes we have? The number of shards?

On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:

> Are you getting the same stack traces in hot threads?
> 
> On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> 
> > I've now switched to mmapfs, and java 7, and still no luck. At least,  
> > we've been able to reproduce the problem. It always happens when people  
> > query elasticsearch :-/
> > 
> > We've had limited success with a handful of concurrent users using our  
> > API internally, but after more than just a couple, it becomes almost  
> > completely unresponsive (multi-minute search times).
> > 
> > We have about 440M docs across about a dozen indexes, on 27 m1.xlarge  
> > instances, each configured with a RAID0 across 4 ephemeral drives. It's  
> > both confusing and frustrating to not understand why we're not getting  
> > better query performance.
> > 
> > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > 
> > > Hey guys, thanks for the replies.
> > > 
> > > The queries are pretty straightforward, and relatively diverse in terms  
> > > of the content they're after. That said, most of the queries are limited to  
> > > query string, and then limiting date ranges. I've not tried MMapDirectory,  
> > > though from that documentation it sounds like the best interface is  
> > > automatically chosen?
> > > 
> > > WRT the java version, I can try 7, as we're apparently running 6.
> > > 
> > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > 
> > > > Dan,
> > > > 
> > > > Which version of java are you using? What you are describing looks  
> > > > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > 
> > > > I would recommend upgrading to the latest java 7 release.
> > > > 
> > > > Igor
> > > > 
> > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic wrote:
> > > > 
> > > > > Hi,
> > > > > 
> > > > > I can't quite tell what is going on from the screenshot.... everything  
> > > > > looks super small and uniform... ☹
> > > > > 
> > > > > If you like paramedic you may also like  
> > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > 
> > > > > But yes, I looked at those stack traces and it looks like a lot of  
> > > > > disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > vmstat.  
> > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > Have you tried using MMapDirectory? See  
> > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > 
> > > > > ## Otis
> > > > > 
> > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > 
> > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > 
> > > > > > Thanks for the reply,
> > > > > > 
> > > > > > So, there is a certain amount of garbage collection going on, though  
> > > > > > not enough to seem to explain this. The stack traces seem to indicate much  
> > > > > > the same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)).  
> > > > > > Something that's puzzling to me is that it seems to be doing a fair amount  
> > > > > > of reading / writing from disk, but the partition with our data is a RAID0  
> > > > > > across four ephemeral drives. `iotop` is claiming only a few MBps of  
> > > > > > read/write, but I've easily seen it hit 100MBps.
> > > > > > 
> > > > > > Here's also a full picture from paramedic:
> > > > > > 
> > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > 
> > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic  
> > > > > > wrote:
> > > > > > 
> > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > > > > > > happening? Have you tried looking at the thread dump? What about all the  
> > > > > > > other system/ES metrics?
> > > > > > > 
> > > > > > > ## Otis
> > > > > > > 
> > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > 
> > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > > > > > 
> > > > > > > > Back with more issues. Periodically, and seemingly inexplicably,  
> > > > > > > > the whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > > > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > > > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines  
> > > > > > > > has been more than a little vague. Each of the machines' CPU usage gets  
> > > > > > > > pegged at 100% of all cores, and then gradually the different machines back  
> > > > > > > > off. From paramedic:
> > > > > > > > 
> > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > 
> > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [December 18, 2012, 7:46pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/10 "2012-12-18T19:46:28Z")

</div>

Which file is the latest hot thread dump?

Could you also add the output of the following commands to your gist?

curl "localhost:9200/\_nodes/\_local/jvm?pretty=true"  
curl "localhost:9200/\_nodes/\_local/stats?jvm=true&pretty=true"

On Tuesday, December 18, 2012 2:27:13 PM UTC-5, Dan Lecocq wrote:

> The stack traces from jstack and hot threads are essentially the same.  
> jstack's quite a bit more verbose, but effectively the same.
> 
> The hot threads now (using mmapfs) no longer say anything about nio, but  
> the profile is very similar in terms of how much CPU is listed as getting  
> used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've been watching  
> iostat, because even after we submit some basic queries and it gets into  
> this state, it remains in a bad way for quite a while. What we're seeing is  
> the io ops (included in gist) stay relatively high during these times  
> despite the throughput (sounds a little like thrashing to me). Could this  
> be symptomatic of the number of indexes we have? The number of shards?
> 
> On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:
> 
> > Are you getting the same stack traces in hot threads?
> > 
> > On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> > 
> > > I've now switched to mmapfs, and java 7, and still no luck. At least,  
> > > we've been able to reproduce the problem. It always happens when people  
> > > query elasticsearch :-/
> > > 
> > > We've had limited success with a handful of concurrent users using our  
> > > API internally, but after more than just a couple, it becomes almost  
> > > completely unresponsive (multi-minute search times).
> > > 
> > > We have about 440M docs across about a dozen indexes, on 27 m1.xlarge  
> > > instances, each configured with a RAID0 across 4 ephemeral drives. It's  
> > > both confusing and frustrating to not understand why we're not getting  
> > > better query performance.
> > > 
> > > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > > 
> > > > Hey guys, thanks for the replies.
> > > > 
> > > > The queries are pretty straightforward, and relatively diverse in terms  
> > > > of the content they're after. That said, most of the queries are limited to  
> > > > query string, and then limiting date ranges. I've not tried MMapDirectory,  
> > > > though from that documentation it sounds like the best interface is  
> > > > automatically chosen?
> > > > 
> > > > WRT the java version, I can try 7, as we're apparently running 6.
> > > > 
> > > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > > 
> > > > > Dan,
> > > > > 
> > > > > Which version of java are you using? What you are describing looks  
> > > > > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > > 
> > > > > I would recommend upgrading to the latest java 7 release.
> > > > > 
> > > > > Igor
> > > > > 
> > > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic wrote:
> > > > > 
> > > > > > Hi,
> > > > > > 
> > > > > > I can't quite tell what is going on from the screenshot....  
> > > > > > everything looks super small and uniform... ☹
> > > > > > 
> > > > > > If you like paramedic you may also like  
> > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > > 
> > > > > > But yes, I looked at those stack traces and it looks like a lot of  
> > > > > > disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > > vmstat.  
> > > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > > Have you tried using MMapDirectory? See  
> > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > > 
> > > > > > ## Otis
> > > > > > 
> > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > > 
> > > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > > 
> > > > > > > Thanks for the reply,
> > > > > > > 
> > > > > > > So, there is a certain amount of garbage collection going on, though  
> > > > > > > not enough to seem to explain this. The stack traces seem to indicate much  
> > > > > > > the same as hot\_threads did ([https://gist.github.com/4280439](https://gist.github.com/4280439)).  
> > > > > > > Something that's puzzling to me is that it seems to be doing a fair amount  
> > > > > > > of reading / writing from disk, but the partition with our data is a RAID0  
> > > > > > > across four ephemeral drives. `iotop` is claiming only a few MBps of  
> > > > > > > read/write, but I've easily seen it hit 100MBps.
> > > > > > > 
> > > > > > > Here's also a full picture from paramedic:
> > > > > > > 
> > > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > > 
> > > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic  
> > > > > > > wrote:
> > > > > > > 
> > > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > > > > > > > happening? Have you tried looking at the thread dump? What about all the  
> > > > > > > > other system/ES metrics?
> > > > > > > > 
> > > > > > > > ## Otis
> > > > > > > > 
> > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > > 
> > > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > 
> > > > > > > > > Back with more issues. Periodically, and seemingly inexplicably,  
> > > > > > > > > the whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > > > > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > > > > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines  
> > > > > > > > > has been more than a little vague. Each of the machines' CPU usage gets  
> > > > > > > > > pegged at 100% of all cores, and then gradually the different machines back  
> > > > > > > > > off. From paramedic:
> > > > > > > > > 
> > > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > > 
> > > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 20, 2012, 9:48pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/11 "2012-12-20T21:48:59Z")

</div>

Sorry for the late reply :-/ We've been able to reproduce this condition by  
"flooding" the cluster with a few basic queries at a time. Most of these  
time out (and the cluster stays in this spinning-its-wheels state for as  
much as an hour), but they all generally take at least 10 seconds if they  
don't time out.

We compiled pylucene on one of the nodes so that we could try to query the  
individual lucene indexes (to verify what we found using luke), and we find  
that most of the shards on a node can be queries very quickly. In fact,  
that average query time there is about 3-5ms.

I updated this gist with the jvm and stats info. Thanks for your continued  
interest in this issue :-/ It's quickly becoming extremely disconcerting.

On Tuesday, December 18, 2012 11:46:28 AM UTC-8, Igor Motov wrote:

> Which file is the latest hot thread dump?
> 
> Could you also add the output of the following commands to your gist?
> 
> curl "localhost:9200/\_nodes/\_local/jvm?pretty=true"  
> curl "localhost:9200/\_nodes/\_local/stats?jvm=true&pretty=true"
> 
> On Tuesday, December 18, 2012 2:27:13 PM UTC-5, Dan Lecocq wrote:
> 
> > The stack traces from jstack and hot threads are essentially the same.  
> > jstack's quite a bit more verbose, but effectively the same.
> > 
> > The hot threads now (using mmapfs) no longer say anything about nio, but  
> > the profile is very similar in terms of how much CPU is listed as getting  
> > used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've been  
> > watching iostat, because even after we submit some basic queries and it  
> > gets into this state, it remains in a bad way for quite a while. What we're  
> > seeing is the io ops (included in gist) stay relatively high during these  
> > times despite the throughput (sounds a little like thrashing to me). Could  
> > this be symptomatic of the number of indexes we have? The number of shards?
> > 
> > On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:
> > 
> > > Are you getting the same stack traces in hot threads?
> > > 
> > > On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> > > 
> > > > I've now switched to mmapfs, and java 7, and still no luck. At least,  
> > > > we've been able to reproduce the problem. It always happens when people  
> > > > query elasticsearch :-/
> > > > 
> > > > We've had limited success with a handful of concurrent users using our  
> > > > API internally, but after more than just a couple, it becomes almost  
> > > > completely unresponsive (multi-minute search times).
> > > > 
> > > > We have about 440M docs across about a dozen indexes, on 27 m1.xlarge  
> > > > instances, each configured with a RAID0 across 4 ephemeral drives. It's  
> > > > both confusing and frustrating to not understand why we're not getting  
> > > > better query performance.
> > > > 
> > > > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > > > 
> > > > > Hey guys, thanks for the replies.
> > > > > 
> > > > > The queries are pretty straightforward, and relatively diverse in  
> > > > > terms of the content they're after. That said, most of the queries are  
> > > > > limited to query string, and then limiting date ranges. I've not tried  
> > > > > MMapDirectory, though from that documentation it sounds like the best  
> > > > > interface is automatically chosen?
> > > > > 
> > > > > WRT the java version, I can try 7, as we're apparently running 6.
> > > > > 
> > > > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > > > 
> > > > > > Dan,
> > > > > > 
> > > > > > Which version of java are you using? What you are describing looks  
> > > > > > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > > > 
> > > > > > I would recommend upgrading to the latest java 7 release.
> > > > > > 
> > > > > > Igor
> > > > > > 
> > > > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic wrote:
> > > > > > 
> > > > > > > Hi,
> > > > > > > 
> > > > > > > I can't quite tell what is going on from the screenshot....  
> > > > > > > everything looks super small and uniform... ☹
> > > > > > > 
> > > > > > > If you like paramedic you may also like  
> > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > > > 
> > > > > > > But yes, I looked at those stack traces and it looks like a lot of  
> > > > > > > disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > > > vmstat.  
> > > > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > > > Have you tried using MMapDirectory? See  
> > > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > > > 
> > > > > > > ## Otis
> > > > > > > 
> > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > > > 
> > > > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > > > 
> > > > > > > > Thanks for the reply,
> > > > > > > > 
> > > > > > > > So, there is a certain amount of garbage collection going on,  
> > > > > > > > though not enough to seem to explain this. The stack traces seem to  
> > > > > > > > indicate much the same as hot\_threads did (  
> > > > > > > > [https://gist.github.com/4280439](https://gist.github.com/4280439)). Something that's puzzling to me  
> > > > > > > > is that it seems to be doing a fair amount of reading / writing from disk,  
> > > > > > > > but the partition with our data is a RAID0 across four ephemeral drives.  
> > > > > > > > `iotop` is claiming only a few MBps of read/write, but I've easily seen it  
> > > > > > > > hit 100MBps.
> > > > > > > > 
> > > > > > > > Here's also a full picture from paramedic:
> > > > > > > > 
> > > > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > > > 
> > > > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic  
> > > > > > > > wrote:
> > > > > > > > 
> > > > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this is  
> > > > > > > > > happening? Have you tried looking at the thread dump? What about all the  
> > > > > > > > > other system/ES metrics?
> > > > > > > > > 
> > > > > > > > > ## Otis
> > > > > > > > > 
> > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > > > 
> > > > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > > 
> > > > > > > > > > Back with more issues. Periodically, and seemingly inexplicably,  
> > > > > > > > > > the whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > > > > > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > > > > > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these  
> > > > > > > > > > machines has been more than a little vague. Each of the machines' CPU usage  
> > > > > > > > > > gets pegged at 100% of all cores, and then gradually the different machines  
> > > > > > > > > > back off. From paramedic:
> > > > > > > > > > 
> > > > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > > > 
> > > > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [December 20, 2012, 10:27pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/12 "2012-12-20T22:27:01Z")

</div>

Are you using leading wildcards in these "basic" queries?

On Thursday, December 20, 2012 4:48:59 PM UTC-5, Dan Lecocq wrote:

> Sorry for the late reply :-/ We've been able to reproduce this condition  
> by "flooding" the cluster with a few basic queries at a time. Most of these  
> time out (and the cluster stays in this spinning-its-wheels state for as  
> much as an hour), but they all generally take at least 10 seconds if they  
> don't time out.
> 
> We compiled pylucene on one of the nodes so that we could try to query the  
> individual lucene indexes (to verify what we found using luke), and we find  
> that most of the shards on a node can be queries very quickly. In fact,  
> that average query time there is about 3-5ms.
> 
> I updated this gist with the jvm and stats info. Thanks for your continued  
> interest in this issue :-/ It's quickly becoming extremely disconcerting.
> 
> On Tuesday, December 18, 2012 11:46:28 AM UTC-8, Igor Motov wrote:
> 
> > Which file is the latest hot thread dump?
> > 
> > Could you also add the output of the following commands to your gist?
> > 
> > curl "localhost:9200/\_nodes/\_local/jvm?pretty=true"  
> > curl "localhost:9200/\_nodes/\_local/stats?jvm=true&pretty=true"
> > 
> > On Tuesday, December 18, 2012 2:27:13 PM UTC-5, Dan Lecocq wrote:
> > 
> > > The stack traces from jstack and hot threads are essentially the same.  
> > > jstack's quite a bit more verbose, but effectively the same.
> > > 
> > > The hot threads now (using mmapfs) no longer say anything about nio, but  
> > > the profile is very similar in terms of how much CPU is listed as getting  
> > > used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've been  
> > > watching iostat, because even after we submit some basic queries and it  
> > > gets into this state, it remains in a bad way for quite a while. What we're  
> > > seeing is the io ops (included in gist) stay relatively high during these  
> > > times despite the throughput (sounds a little like thrashing to me). Could  
> > > this be symptomatic of the number of indexes we have? The number of shards?
> > > 
> > > On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:
> > > 
> > > > Are you getting the same stack traces in hot threads?
> > > > 
> > > > On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> > > > 
> > > > > I've now switched to mmapfs, and java 7, and still no luck. At least,  
> > > > > we've been able to reproduce the problem. It always happens when people  
> > > > > query elasticsearch :-/
> > > > > 
> > > > > We've had limited success with a handful of concurrent users using our  
> > > > > API internally, but after more than just a couple, it becomes almost  
> > > > > completely unresponsive (multi-minute search times).
> > > > > 
> > > > > We have about 440M docs across about a dozen indexes, on 27 m1.xlarge  
> > > > > instances, each configured with a RAID0 across 4 ephemeral drives. It's  
> > > > > both confusing and frustrating to not understand why we're not getting  
> > > > > better query performance.
> > > > > 
> > > > > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > > > > 
> > > > > > Hey guys, thanks for the replies.
> > > > > > 
> > > > > > The queries are pretty straightforward, and relatively diverse in  
> > > > > > terms of the content they're after. That said, most of the queries are  
> > > > > > limited to query string, and then limiting date ranges. I've not tried  
> > > > > > MMapDirectory, though from that documentation it sounds like the best  
> > > > > > interface is automatically chosen?
> > > > > > 
> > > > > > WRT the java version, I can try 7, as we're apparently running 6.
> > > > > > 
> > > > > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > > > > 
> > > > > > > Dan,
> > > > > > > 
> > > > > > > Which version of java are you using? What you are describing looks  
> > > > > > > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > > > > 
> > > > > > > I would recommend upgrading to the latest java 7 release.
> > > > > > > 
> > > > > > > Igor
> > > > > > > 
> > > > > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic  
> > > > > > > wrote:
> > > > > > > 
> > > > > > > > Hi,
> > > > > > > > 
> > > > > > > > I can't quite tell what is going on from the screenshot....  
> > > > > > > > everything looks super small and uniform... ☹
> > > > > > > > 
> > > > > > > > If you like paramedic you may also like  
> > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > > > > 
> > > > > > > > But yes, I looked at those stack traces and it looks like a lot of  
> > > > > > > > disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > > > > vmstat.  
> > > > > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > > > > Have you tried using MMapDirectory? See  
> > > > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > > > > 
> > > > > > > > ## Otis
> > > > > > > > 
> > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > > > > 
> > > > > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > 
> > > > > > > > > Thanks for the reply,
> > > > > > > > > 
> > > > > > > > > So, there is a certain amount of garbage collection going on,  
> > > > > > > > > though not enough to seem to explain this. The stack traces seem to  
> > > > > > > > > indicate much the same as hot\_threads did (  
> > > > > > > > > [https://gist.github.com/4280439](https://gist.github.com/4280439)). Something that's puzzling to me  
> > > > > > > > > is that it seems to be doing a fair amount of reading / writing from disk,  
> > > > > > > > > but the partition with our data is a RAID0 across four ephemeral drives.  
> > > > > > > > > `iotop` is claiming only a few MBps of read/write, but I've easily seen it  
> > > > > > > > > hit 100MBps.
> > > > > > > > > 
> > > > > > > > > Here's also a full picture from paramedic:
> > > > > > > > > 
> > > > > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > > > > 
> > > > > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic  
> > > > > > > > > wrote:
> > > > > > > > > 
> > > > > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this  
> > > > > > > > > > is happening? Have you tried looking at the thread dump? What about all  
> > > > > > > > > > the other system/ES metrics?
> > > > > > > > > > 
> > > > > > > > > > ## Otis
> > > > > > > > > > 
> > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > > > > 
> > > > > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > > > 
> > > > > > > > > > > Back with more issues. Periodically, and seemingly inexplicably,  
> > > > > > > > > > > the whole cluster becomes essentially unusable for what's sometimes hours.  
> > > > > > > > > > > There's nothing in the logs (or slow log) to indicate what's up, and  
> > > > > > > > > > > `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439)) on these  
> > > > > > > > > > > machines has been more than a little vague. Each of the machines' CPU usage  
> > > > > > > > > > > gets pegged at 100% of all cores, and then gradually the different machines  
> > > > > > > > > > > back off. From paramedic:
> > > > > > > > > > > 
> > > > > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > > > > 
> > > > > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 20, 2012, 10:45pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/13 "2012-12-20T22:45:11Z")

</div>

Our initial tests included some wildcards in some queries, but they were  
never leading wildcards, and we've since removed them for subsequent tests  
and still see this issue.

On Thursday, December 20, 2012 2:27:01 PM UTC-8, Igor Motov wrote:

> Are you using leading wildcards in these "basic" queries?
> 
> On Thursday, December 20, 2012 4:48:59 PM UTC-5, Dan Lecocq wrote:
> 
> > Sorry for the late reply :-/ We've been able to reproduce this condition  
> > by "flooding" the cluster with a few basic queries at a time. Most of these  
> > time out (and the cluster stays in this spinning-its-wheels state for as  
> > much as an hour), but they all generally take at least 10 seconds if they  
> > don't time out.
> > 
> > We compiled pylucene on one of the nodes so that we could try to query  
> > the individual lucene indexes (to verify what we found using luke), and we  
> > find that most of the shards on a node can be queries very quickly. In  
> > fact, that average query time there is about 3-5ms.
> > 
> > I updated this gist with the jvm and stats info. Thanks for your  
> > continued interest in this issue :-/ It's quickly becoming extremely  
> > disconcerting.
> > 
> > On Tuesday, December 18, 2012 11:46:28 AM UTC-8, Igor Motov wrote:
> > 
> > > Which file is the latest hot thread dump?
> > > 
> > > Could you also add the output of the following commands to your gist?
> > > 
> > > curl "localhost:9200/\_nodes/\_local/jvm?pretty=true"  
> > > curl "localhost:9200/\_nodes/\_local/stats?jvm=true&pretty=true"
> > > 
> > > On Tuesday, December 18, 2012 2:27:13 PM UTC-5, Dan Lecocq wrote:
> > > 
> > > > The stack traces from jstack and hot threads are essentially the same.  
> > > > jstack's quite a bit more verbose, but effectively the same.
> > > > 
> > > > The hot threads now (using mmapfs) no longer say anything about nio,  
> > > > but the profile is very similar in terms of how much CPU is listed as  
> > > > getting used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've  
> > > > been watching iostat, because even after we submit some basic queries and  
> > > > it gets into this state, it remains in a bad way for quite a while. What  
> > > > we're seeing is the io ops (included in gist) stay relatively high during  
> > > > these times despite the throughput (sounds a little like thrashing to me).  
> > > > Could this be symptomatic of the number of indexes we have? The number of  
> > > > shards?
> > > > 
> > > > On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:
> > > > 
> > > > > Are you getting the same stack traces in hot threads?
> > > > > 
> > > > > On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> > > > > 
> > > > > > I've now switched to mmapfs, and java 7, and still no luck. At least,  
> > > > > > we've been able to reproduce the problem. It always happens when people  
> > > > > > query elasticsearch :-/
> > > > > > 
> > > > > > We've had limited success with a handful of concurrent users using  
> > > > > > our API internally, but after more than just a couple, it becomes almost  
> > > > > > completely unresponsive (multi-minute search times).
> > > > > > 
> > > > > > We have about 440M docs across about a dozen indexes, on 27 m1.xlarge  
> > > > > > instances, each configured with a RAID0 across 4 ephemeral drives. It's  
> > > > > > both confusing and frustrating to not understand why we're not getting  
> > > > > > better query performance.
> > > > > > 
> > > > > > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > > > > > 
> > > > > > > Hey guys, thanks for the replies.
> > > > > > > 
> > > > > > > The queries are pretty straightforward, and relatively diverse in  
> > > > > > > terms of the content they're after. That said, most of the queries are  
> > > > > > > limited to query string, and then limiting date ranges. I've not tried  
> > > > > > > MMapDirectory, though from that documentation it sounds like the best  
> > > > > > > interface is automatically chosen?
> > > > > > > 
> > > > > > > WRT the java version, I can try 7, as we're apparently running 6.
> > > > > > > 
> > > > > > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > > > > > 
> > > > > > > > Dan,
> > > > > > > > 
> > > > > > > > Which version of java are you using? What you are describing looks  
> > > > > > > > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > > > > > 
> > > > > > > > I would recommend upgrading to the latest java 7 release.
> > > > > > > > 
> > > > > > > > Igor
> > > > > > > > 
> > > > > > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic  
> > > > > > > > wrote:
> > > > > > > > 
> > > > > > > > > Hi,
> > > > > > > > > 
> > > > > > > > > I can't quite tell what is going on from the screenshot....  
> > > > > > > > > everything looks super small and uniform... ☹
> > > > > > > > > 
> > > > > > > > > If you like paramedic you may also like  
> > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > > > > > 
> > > > > > > > > But yes, I looked at those stack traces and it looks like a lot of  
> > > > > > > > > disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > > > > > vmstat.  
> > > > > > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > > > > > Have you tried using MMapDirectory? See  
> > > > > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > > > > > 
> > > > > > > > > ## Otis
> > > > > > > > > 
> > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > > > > > 
> > > > > > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > > 
> > > > > > > > > > Thanks for the reply,
> > > > > > > > > > 
> > > > > > > > > > So, there is a certain amount of garbage collection going on,  
> > > > > > > > > > though not enough to seem to explain this. The stack traces seem to  
> > > > > > > > > > indicate much the same as hot\_threads did (  
> > > > > > > > > > [https://gist.github.com/4280439](https://gist.github.com/4280439)). Something that's puzzling to  
> > > > > > > > > > me is that it seems to be doing a fair amount of reading / writing from  
> > > > > > > > > > disk, but the partition with our data is a RAID0 across four ephemeral  
> > > > > > > > > > drives. `iotop` is claiming only a few MBps of read/write, but I've easily  
> > > > > > > > > > seen it hit 100MBps.
> > > > > > > > > > 
> > > > > > > > > > Here's also a full picture from paramedic:
> > > > > > > > > > 
> > > > > > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > > > > > 
> > > > > > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis Gospodnetic  
> > > > > > > > > > wrote:
> > > > > > > > > > 
> > > > > > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this  
> > > > > > > > > > > is happening? Have you tried looking at the thread dump? What about all  
> > > > > > > > > > > the other system/ES metrics?
> > > > > > > > > > > 
> > > > > > > > > > > ## Otis
> > > > > > > > > > > 
> > > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > > > > > 
> > > > > > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq  
> > > > > > > > > > > wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > Back with more issues. Periodically, and seemingly  
> > > > > > > > > > > > inexplicably, the whole cluster becomes essentially unusable for what's  
> > > > > > > > > > > > sometimes hours. There's nothing in the logs (or slow log) to indicate  
> > > > > > > > > > > > what's up, and `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439))  
> > > > > > > > > > > > on these machines has been more than a little vague. Each of the machines'  
> > > > > > > > > > > > CPU usage gets pegged at 100% of all cores, and then gradually the  
> > > > > > > > > > > > different machines back off. From paramedic:
> > > > > > > > > > > > 
> > > > > > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > > > > > 
> > > > > > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 20, 2012, 10:46pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/14 "2012-12-20T22:46:35Z")

</div>

Oh, it's also worth mentioning that at this point, given our free RAM on  
these instances is extremely low, we think it might be an issue of  
thrashing the fs cache. We're giving it a try with instances with both  
higher IO and higher memory available.

On Thursday, December 20, 2012 2:45:11 PM UTC-8, Dan Lecocq wrote:

> Our initial tests included some wildcards in some queries, but they were  
> never leading wildcards, and we've since removed them for subsequent tests  
> and still see this issue.
> 
> On Thursday, December 20, 2012 2:27:01 PM UTC-8, Igor Motov wrote:
> 
> > Are you using leading wildcards in these "basic" queries?
> > 
> > On Thursday, December 20, 2012 4:48:59 PM UTC-5, Dan Lecocq wrote:
> > 
> > > Sorry for the late reply :-/ We've been able to reproduce this condition  
> > > by "flooding" the cluster with a few basic queries at a time. Most of these  
> > > time out (and the cluster stays in this spinning-its-wheels state for as  
> > > much as an hour), but they all generally take at least 10 seconds if they  
> > > don't time out.
> > > 
> > > We compiled pylucene on one of the nodes so that we could try to query  
> > > the individual lucene indexes (to verify what we found using luke), and we  
> > > find that most of the shards on a node can be queries very quickly. In  
> > > fact, that average query time there is about 3-5ms.
> > > 
> > > I updated this gist with the jvm and stats info. Thanks for your  
> > > continued interest in this issue :-/ It's quickly becoming extremely  
> > > disconcerting.
> > > 
> > > On Tuesday, December 18, 2012 11:46:28 AM UTC-8, Igor Motov wrote:
> > > 
> > > > Which file is the latest hot thread dump?
> > > > 
> > > > Could you also add the output of the following commands to your gist?
> > > > 
> > > > curl "localhost:9200/\_nodes/\_local/jvm?pretty=true"  
> > > > curl "localhost:9200/\_nodes/\_local/stats?jvm=true&pretty=true"
> > > > 
> > > > On Tuesday, December 18, 2012 2:27:13 PM UTC-5, Dan Lecocq wrote:
> > > > 
> > > > > The stack traces from jstack and hot threads are essentially the same.  
> > > > > jstack's quite a bit more verbose, but effectively the same.
> > > > > 
> > > > > The hot threads now (using mmapfs) no longer say anything about nio,  
> > > > > but the profile is very similar in terms of how much CPU is listed as  
> > > > > getting used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've  
> > > > > been watching iostat, because even after we submit some basic queries and  
> > > > > it gets into this state, it remains in a bad way for quite a while. What  
> > > > > we're seeing is the io ops (included in gist) stay relatively high during  
> > > > > these times despite the throughput (sounds a little like thrashing to me).  
> > > > > Could this be symptomatic of the number of indexes we have? The number of  
> > > > > shards?
> > > > > 
> > > > > On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:
> > > > > 
> > > > > > Are you getting the same stack traces in hot threads?
> > > > > > 
> > > > > > On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> > > > > > 
> > > > > > > I've now switched to mmapfs, and java 7, and still no luck. At  
> > > > > > > least, we've been able to reproduce the problem. It always happens when  
> > > > > > > people query elasticsearch :-/
> > > > > > > 
> > > > > > > We've had limited success with a handful of concurrent users using  
> > > > > > > our API internally, but after more than just a couple, it becomes almost  
> > > > > > > completely unresponsive (multi-minute search times).
> > > > > > > 
> > > > > > > We have about 440M docs across about a dozen indexes, on 27  
> > > > > > > m1.xlarge instances, each configured with a RAID0 across 4 ephemeral  
> > > > > > > drives. It's both confusing and frustrating to not understand why we're not  
> > > > > > > getting better query performance.
> > > > > > > 
> > > > > > > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > > > > > > 
> > > > > > > > Hey guys, thanks for the replies.
> > > > > > > > 
> > > > > > > > The queries are pretty straightforward, and relatively diverse in  
> > > > > > > > terms of the content they're after. That said, most of the queries are  
> > > > > > > > limited to query string, and then limiting date ranges. I've not tried  
> > > > > > > > MMapDirectory, though from that documentation it sounds like the best  
> > > > > > > > interface is automatically chosen?
> > > > > > > > 
> > > > > > > > WRT the java version, I can try 7, as we're apparently running 6.
> > > > > > > > 
> > > > > > > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > > > > > > 
> > > > > > > > > Dan,
> > > > > > > > > 
> > > > > > > > > Which version of java are you using? What you are describing looks  
> > > > > > > > > like this java bug: [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > > > > > > 
> > > > > > > > > I would recommend upgrading to the latest java 7 release.
> > > > > > > > > 
> > > > > > > > > Igor
> > > > > > > > > 
> > > > > > > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic  
> > > > > > > > > wrote:
> > > > > > > > > 
> > > > > > > > > > Hi,
> > > > > > > > > > 
> > > > > > > > > > I can't quite tell what is going on from the screenshot....  
> > > > > > > > > > everything looks super small and uniform... ☹
> > > > > > > > > > 
> > > > > > > > > > If you like paramedic you may also like  
> > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > > > > > > 
> > > > > > > > > > But yes, I looked at those stack traces and it looks like a lot  
> > > > > > > > > > of disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > > > > > > vmstat.  
> > > > > > > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > > > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > > > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > > > > > > Have you tried using MMapDirectory? See  
> > > > > > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > > > > > > 
> > > > > > > > > > ## Otis
> > > > > > > > > > 
> > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > > > > > > 
> > > > > > > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > > > 
> > > > > > > > > > > Thanks for the reply,
> > > > > > > > > > > 
> > > > > > > > > > > So, there is a certain amount of garbage collection going on,  
> > > > > > > > > > > though not enough to seem to explain this. The stack traces seem to  
> > > > > > > > > > > indicate much the same as hot\_threads did (  
> > > > > > > > > > > [https://gist.github.com/4280439](https://gist.github.com/4280439)). Something that's puzzling to  
> > > > > > > > > > > me is that it seems to be doing a fair amount of reading / writing from  
> > > > > > > > > > > disk, but the partition with our data is a RAID0 across four ephemeral  
> > > > > > > > > > > drives. `iotop` is claiming only a few MBps of read/write, but I've easily  
> > > > > > > > > > > seen it hit 100MBps.
> > > > > > > > > > > 
> > > > > > > > > > > Here's also a full picture from paramedic:
> > > > > > > > > > > 
> > > > > > > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > > > > > > 
> > > > > > > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis  
> > > > > > > > > > > Gospodnetic wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when this  
> > > > > > > > > > > > is happening? Have you tried looking at the thread dump? What about all  
> > > > > > > > > > > > the other system/ES metrics?
> > > > > > > > > > > > 
> > > > > > > > > > > > ## Otis
> > > > > > > > > > > > 
> > > > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > > > > > > 
> > > > > > > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq  
> > > > > > > > > > > > wrote:
> > > > > > > > > > > > 
> > > > > > > > > > > > > Back with more issues. Periodically, and seemingly  
> > > > > > > > > > > > > inexplicably, the whole cluster becomes essentially unusable for what's  
> > > > > > > > > > > > > sometimes hours. There's nothing in the logs (or slow log) to indicate  
> > > > > > > > > > > > > what's up, and `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439))  
> > > > > > > > > > > > > on these machines has been more than a little vague. Each of the machines'  
> > > > > > > > > > > > > CPU usage gets pegged at 100% of all cores, and then gradually the  
> > > > > > > > > > > > > different machines back off. From paramedic:
> > > > > > > > > > > > > 
> > > > > > > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > > > > > > 
> > > > > > > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [December 20, 2012, 10:51pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/15 "2012-12-20T22:51:49Z")

</div>

Judging from hot threads that you posted you are still running wild card  
queries and this is where most of the time is spent.

On Thursday, December 20, 2012 5:46:35 PM UTC-5, Dan Lecocq wrote:

> Oh, it's also worth mentioning that at this point, given our free RAM on  
> these instances is extremely low, we think it might be an issue of  
> thrashing the fs cache. We're giving it a try with instances with both  
> higher IO and higher memory available.
> 
> On Thursday, December 20, 2012 2:45:11 PM UTC-8, Dan Lecocq wrote:
> 
> > Our initial tests included some wildcards in some queries, but they were  
> > never leading wildcards, and we've since removed them for subsequent tests  
> > and still see this issue.
> > 
> > On Thursday, December 20, 2012 2:27:01 PM UTC-8, Igor Motov wrote:
> > 
> > > Are you using leading wildcards in these "basic" queries?
> > > 
> > > On Thursday, December 20, 2012 4:48:59 PM UTC-5, Dan Lecocq wrote:
> > > 
> > > > Sorry for the late reply :-/ We've been able to reproduce this  
> > > > condition by "flooding" the cluster with a few basic queries at a time.  
> > > > Most of these time out (and the cluster stays in this spinning-its-wheels  
> > > > state for as much as an hour), but they all generally take at least 10  
> > > > seconds if they don't time out.
> > > > 
> > > > We compiled pylucene on one of the nodes so that we could try to query  
> > > > the individual lucene indexes (to verify what we found using luke), and we  
> > > > find that most of the shards on a node can be queries very quickly. In  
> > > > fact, that average query time there is about 3-5ms.
> > > > 
> > > > I updated this gist with the jvm and stats info. Thanks for your  
> > > > continued interest in this issue :-/ It's quickly becoming extremely  
> > > > disconcerting.
> > > > 
> > > > On Tuesday, December 18, 2012 11:46:28 AM UTC-8, Igor Motov wrote:
> > > > 
> > > > > Which file is the latest hot thread dump?
> > > > > 
> > > > > Could you also add the output of the following commands to your gist?
> > > > > 
> > > > > curl "localhost:9200/\_nodes/\_local/jvm?pretty=true"  
> > > > > curl "localhost:9200/\_nodes/\_local/stats?jvm=true&pretty=true"
> > > > > 
> > > > > On Tuesday, December 18, 2012 2:27:13 PM UTC-5, Dan Lecocq wrote:
> > > > > 
> > > > > > The stack traces from jstack and hot threads are essentially the  
> > > > > > same. jstack's quite a bit more verbose, but effectively the same.
> > > > > > 
> > > > > > The hot threads now (using mmapfs) no longer say anything about nio,  
> > > > > > but the profile is very similar in terms of how much CPU is listed as  
> > > > > > getting used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've  
> > > > > > been watching iostat, because even after we submit some basic queries and  
> > > > > > it gets into this state, it remains in a bad way for quite a while. What  
> > > > > > we're seeing is the io ops (included in gist) stay relatively high during  
> > > > > > these times despite the throughput (sounds a little like thrashing to me).  
> > > > > > Could this be symptomatic of the number of indexes we have? The number of  
> > > > > > shards?
> > > > > > 
> > > > > > On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:
> > > > > > 
> > > > > > > Are you getting the same stack traces in hot threads?
> > > > > > > 
> > > > > > > On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> > > > > > > 
> > > > > > > > I've now switched to mmapfs, and java 7, and still no luck. At  
> > > > > > > > least, we've been able to reproduce the problem. It always happens when  
> > > > > > > > people query elasticsearch :-/
> > > > > > > > 
> > > > > > > > We've had limited success with a handful of concurrent users using  
> > > > > > > > our API internally, but after more than just a couple, it becomes almost  
> > > > > > > > completely unresponsive (multi-minute search times).
> > > > > > > > 
> > > > > > > > We have about 440M docs across about a dozen indexes, on 27  
> > > > > > > > m1.xlarge instances, each configured with a RAID0 across 4 ephemeral  
> > > > > > > > drives. It's both confusing and frustrating to not understand why we're not  
> > > > > > > > getting better query performance.
> > > > > > > > 
> > > > > > > > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > > > > > > > 
> > > > > > > > > Hey guys, thanks for the replies.
> > > > > > > > > 
> > > > > > > > > The queries are pretty straightforward, and relatively diverse in  
> > > > > > > > > terms of the content they're after. That said, most of the queries are  
> > > > > > > > > limited to query string, and then limiting date ranges. I've not tried  
> > > > > > > > > MMapDirectory, though from that documentation it sounds like the best  
> > > > > > > > > interface is automatically chosen?
> > > > > > > > > 
> > > > > > > > > WRT the java version, I can try 7, as we're apparently running 6.
> > > > > > > > > 
> > > > > > > > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > > > > > > > 
> > > > > > > > > > Dan,
> > > > > > > > > > 
> > > > > > > > > > Which version of java are you using? What you are describing  
> > > > > > > > > > looks like this java bug:  
> > > > > > > > > > [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > > > > > > > 
> > > > > > > > > > I would recommend upgrading to the latest java 7 release.
> > > > > > > > > > 
> > > > > > > > > > Igor
> > > > > > > > > > 
> > > > > > > > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic  
> > > > > > > > > > wrote:
> > > > > > > > > > 
> > > > > > > > > > > Hi,
> > > > > > > > > > > 
> > > > > > > > > > > I can't quite tell what is going on from the screenshot....  
> > > > > > > > > > > everything looks super small and uniform... ☹
> > > > > > > > > > > 
> > > > > > > > > > > If you like paramedic you may also like  
> > > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > > > > > > > 
> > > > > > > > > > > But yes, I looked at those stack traces and it looks like a lot  
> > > > > > > > > > > of disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > > > > > > > vmstat.  
> > > > > > > > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > > > > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > > > > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > > > > > > > Have you tried using MMapDirectory? See  
> > > > > > > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > > > > > > > 
> > > > > > > > > > > ## Otis
> > > > > > > > > > > 
> > > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > > > > > > > 
> > > > > > > > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > Thanks for the reply,
> > > > > > > > > > > > 
> > > > > > > > > > > > So, there is a certain amount of garbage collection going on,  
> > > > > > > > > > > > though not enough to seem to explain this. The stack traces seem to  
> > > > > > > > > > > > indicate much the same as hot\_threads did (  
> > > > > > > > > > > > [https://gist.github.com/4280439](https://gist.github.com/4280439)). Something that's puzzling to  
> > > > > > > > > > > > me is that it seems to be doing a fair amount of reading / writing from  
> > > > > > > > > > > > disk, but the partition with our data is a RAID0 across four ephemeral  
> > > > > > > > > > > > drives. `iotop` is claiming only a few MBps of read/write, but I've easily  
> > > > > > > > > > > > seen it hit 100MBps.
> > > > > > > > > > > > 
> > > > > > > > > > > > Here's also a full picture from paramedic:
> > > > > > > > > > > > 
> > > > > > > > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > > > > > > > 
> > > > > > > > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis  
> > > > > > > > > > > > Gospodnetic wrote:
> > > > > > > > > > > > 
> > > > > > > > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when  
> > > > > > > > > > > > > this is happening? Have you tried looking at the thread dump? What about  
> > > > > > > > > > > > > all the other system/ES metrics?
> > > > > > > > > > > > > 
> > > > > > > > > > > > > ## Otis
> > > > > > > > > > > > > 
> > > > > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > > > > > > > 
> > > > > > > > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq  
> > > > > > > > > > > > > wrote:
> > > > > > > > > > > > > 
> > > > > > > > > > > > > > Back with more issues. Periodically, and seemingly  
> > > > > > > > > > > > > > inexplicably, the whole cluster becomes essentially unusable for what's  
> > > > > > > > > > > > > > sometimes hours. There's nothing in the logs (or slow log) to indicate  
> > > > > > > > > > > > > > what's up, and `hot_threads` ([https://gist.github.com/4280439](https://gist.github.com/4280439))  
> > > > > > > > > > > > > > on these machines has been more than a little vague. Each of the machines'  
> > > > > > > > > > > > > > CPU usage gets pegged at 100% of all cores, and then gradually the  
> > > > > > > > > > > > > > different machines back off. From paramedic:
> > > > > > > > > > > > > > 
> > > > > > > > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > > > > > > > 
> > > > > > > > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![Dan\_Lecocq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_lecocq/32/2565_2.png) [@Dan\_Lecocq](https://discuss.elastic.co/u/Dan_Lecocq)\
**Post date:** [December 21, 2012, 2:21pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/16 "2012-12-21T14:21:52Z")

</div>

D'oh -- I think I deleted the wrong hot\_threads in that gist.

At any rate, we just ran a test using pylucene against the lucene indexes  
from one of our nodes, using an instance with more RAM. We suspect that we  
were thrashing the page-level cache, and judging from our results on a  
m2.4xlarge, we've largely confirmed our suspicions. Our next step is to  
spin up a small ES cluster with that instance type and confirm the same  
timing results through ES.

On Thursday, December 20, 2012 2:51:49 PM UTC-8, Igor Motov wrote:

> Judging from hot threads that you posted you are still running wild card  
> queries and this is where most of the time is spent.
> 
> On Thursday, December 20, 2012 5:46:35 PM UTC-5, Dan Lecocq wrote:
> 
> > Oh, it's also worth mentioning that at this point, given our free RAM on  
> > these instances is extremely low, we think it might be an issue of  
> > thrashing the fs cache. We're giving it a try with instances with both  
> > higher IO and higher memory available.
> > 
> > On Thursday, December 20, 2012 2:45:11 PM UTC-8, Dan Lecocq wrote:
> > 
> > > Our initial tests included some wildcards in some queries, but they were  
> > > never leading wildcards, and we've since removed them for subsequent tests  
> > > and still see this issue.
> > > 
> > > On Thursday, December 20, 2012 2:27:01 PM UTC-8, Igor Motov wrote:
> > > 
> > > > Are you using leading wildcards in these "basic" queries?
> > > > 
> > > > On Thursday, December 20, 2012 4:48:59 PM UTC-5, Dan Lecocq wrote:
> > > > 
> > > > > Sorry for the late reply :-/ We've been able to reproduce this  
> > > > > condition by "flooding" the cluster with a few basic queries at a time.  
> > > > > Most of these time out (and the cluster stays in this spinning-its-wheels  
> > > > > state for as much as an hour), but they all generally take at least 10  
> > > > > seconds if they don't time out.
> > > > > 
> > > > > We compiled pylucene on one of the nodes so that we could try to query  
> > > > > the individual lucene indexes (to verify what we found using luke), and we  
> > > > > find that most of the shards on a node can be queries very quickly. In  
> > > > > fact, that average query time there is about 3-5ms.
> > > > > 
> > > > > I updated this gist with the jvm and stats info. Thanks for your  
> > > > > continued interest in this issue :-/ It's quickly becoming extremely  
> > > > > disconcerting.
> > > > > 
> > > > > On Tuesday, December 18, 2012 11:46:28 AM UTC-8, Igor Motov wrote:
> > > > > 
> > > > > > Which file is the latest hot thread dump?
> > > > > > 
> > > > > > Could you also add the output of the following commands to your gist?
> > > > > > 
> > > > > > curl "localhost:9200/\_nodes/\_local/jvm?pretty=true"  
> > > > > > curl "localhost:9200/\_nodes/\_local/stats?jvm=true&pretty=true"
> > > > > > 
> > > > > > On Tuesday, December 18, 2012 2:27:13 PM UTC-5, Dan Lecocq wrote:
> > > > > > 
> > > > > > > The stack traces from jstack and hot threads are essentially the  
> > > > > > > same. jstack's quite a bit more verbose, but effectively the same.
> > > > > > > 
> > > > > > > The hot threads now (using mmapfs) no longer say anything about nio,  
> > > > > > > but the profile is very similar in terms of how much CPU is listed as  
> > > > > > > getting used (gist updated: [https://gist.github.com/4280439](https://gist.github.com/4280439)). We've  
> > > > > > > been watching iostat, because even after we submit some basic queries and  
> > > > > > > it gets into this state, it remains in a bad way for quite a while. What  
> > > > > > > we're seeing is the io ops (included in gist) stay relatively high during  
> > > > > > > these times despite the throughput (sounds a little like thrashing to me).  
> > > > > > > Could this be symptomatic of the number of indexes we have? The number of  
> > > > > > > shards?
> > > > > > > 
> > > > > > > On Tuesday, December 18, 2012 11:18:36 AM UTC-8, Igor Motov wrote:
> > > > > > > 
> > > > > > > > Are you getting the same stack traces in hot threads?
> > > > > > > > 
> > > > > > > > On Monday, December 17, 2012 6:26:01 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > 
> > > > > > > > > I've now switched to mmapfs, and java 7, and still no luck. At  
> > > > > > > > > least, we've been able to reproduce the problem. It always happens when  
> > > > > > > > > people query elasticsearch :-/
> > > > > > > > > 
> > > > > > > > > We've had limited success with a handful of concurrent users using  
> > > > > > > > > our API internally, but after more than just a couple, it becomes almost  
> > > > > > > > > completely unresponsive (multi-minute search times).
> > > > > > > > > 
> > > > > > > > > We have about 440M docs across about a dozen indexes, on 27  
> > > > > > > > > m1.xlarge instances, each configured with a RAID0 across 4 ephemeral  
> > > > > > > > > drives. It's both confusing and frustrating to not understand why we're not  
> > > > > > > > > getting better query performance.
> > > > > > > > > 
> > > > > > > > > On Monday, December 17, 2012 8:07:51 AM UTC-8, Dan Lecocq wrote:
> > > > > > > > > 
> > > > > > > > > > Hey guys, thanks for the replies.
> > > > > > > > > > 
> > > > > > > > > > The queries are pretty straightforward, and relatively diverse in  
> > > > > > > > > > terms of the content they're after. That said, most of the queries are  
> > > > > > > > > > limited to query string, and then limiting date ranges. I've not tried  
> > > > > > > > > > MMapDirectory, though from that documentation it sounds like the best  
> > > > > > > > > > interface is automatically chosen?
> > > > > > > > > > 
> > > > > > > > > > WRT the java version, I can try 7, as we're apparently running 6.
> > > > > > > > > > 
> > > > > > > > > > On Monday, December 17, 2012 7:45:25 AM UTC-8, Igor Motov wrote:
> > > > > > > > > > 
> > > > > > > > > > > Dan,
> > > > > > > > > > > 
> > > > > > > > > > > Which version of java are you using? What you are describing  
> > > > > > > > > > > looks like this java bug:  
> > > > > > > > > > > [http://bugs.sun.com/view\_bug.do?bug\_id=6919638](http://bugs.sun.com/view_bug.do?bug_id=6919638)
> > > > > > > > > > > 
> > > > > > > > > > > I would recommend upgrading to the latest java 7 release.
> > > > > > > > > > > 
> > > > > > > > > > > Igor
> > > > > > > > > > > 
> > > > > > > > > > > On Friday, December 14, 2012 8:14:41 PM UTC-8, Otis Gospodnetic  
> > > > > > > > > > > wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > Hi,
> > > > > > > > > > > > 
> > > > > > > > > > > > I can't quite tell what is going on from the screenshot....  
> > > > > > > > > > > > everything looks super small and uniform... ☹
> > > > > > > > > > > > 
> > > > > > > > > > > > If you like paramedic you may also like  
> > > > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)
> > > > > > > > > > > > 
> > > > > > > > > > > > But yes, I looked at those stack traces and it looks like a lot  
> > > > > > > > > > > > of disk reading. Maybe iotop is lying. dstat is nice, as is iostat and  
> > > > > > > > > > > > vmstat.  
> > > > > > > > > > > > How big is your index, how much RAM have you got? What sort of  
> > > > > > > > > > > > queries are you serving? Are they very diverse? Can you show a bit of  
> > > > > > > > > > > > vmstat 2 output or disk IO graph from SPM?  
> > > > > > > > > > > > Have you tried using MMapDirectory? See  
> > > > > > > > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html)
> > > > > > > > > > > > 
> > > > > > > > > > > > ## Otis
> > > > > > > > > > > > 
> > > > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > > > > > > > > > 
> > > > > > > > > > > > On Friday, December 14, 2012 5:34:24 PM UTC-5, Dan Lecocq wrote:
> > > > > > > > > > > > 
> > > > > > > > > > > > > Thanks for the reply,
> > > > > > > > > > > > > 
> > > > > > > > > > > > > So, there is a certain amount of garbage collection going on,  
> > > > > > > > > > > > > though not enough to seem to explain this. The stack traces seem to  
> > > > > > > > > > > > > indicate much the same as hot\_threads did (  
> > > > > > > > > > > > > [https://gist.github.com/4280439](https://gist.github.com/4280439)). Something that's puzzling  
> > > > > > > > > > > > > to me is that it seems to be doing a fair amount of reading / writing from  
> > > > > > > > > > > > > disk, but the partition with our data is a RAID0 across four ephemeral  
> > > > > > > > > > > > > drives. `iotop` is claiming only a few MBps of read/write, but I've easily  
> > > > > > > > > > > > > seen it hit 100MBps.
> > > > > > > > > > > > > 
> > > > > > > > > > > > > Here's also a full picture from paramedic:
> > > > > > > > > > > > > 
> > > > > > > > > > > > > [https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png](https://lh3.googleusercontent.com/-i-OXrFSbLH0/UMupZivxgWI/AAAAAAAAAAw/DOL1B2BJLoc/s1600/Paramedic+++freshscape+production+sudden+unexplained.png)
> > > > > > > > > > > > > 
> > > > > > > > > > > > > On Thursday, December 13, 2012 8:25:06 PM UTC-8, Otis  
> > > > > > > > > > > > > Gospodnetic wrote:
> > > > > > > > > > > > > 
> > > > > > > > > > > > > > Ouch, that's a lot of green. How's your JVM/GC doing when  
> > > > > > > > > > > > > > this is happening? Have you tried looking at the thread dump? What about  
> > > > > > > > > > > > > > all the other system/ES metrics?
> > > > > > > > > > > > > > 
> > > > > > > > > > > > > > ## Otis
> > > > > > > > > > > > > > 
> > > > > > > > > > > > > > ELASTICSEARCH Performance Monitoring -  
> > > > > > > > > > > > > > [Elasticsearch Monitoring](http://sematext.com/spm/elasticsearch-performance-monitoring/index.html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > > > > > > > > > > > > > 
> > > > > > > > > > > > > > On Thursday, December 13, 2012 5:05:05 PM UTC-5, Dan Lecocq  
> > > > > > > > > > > > > > wrote:
> > > > > > > > > > > > > > 
> > > > > > > > > > > > > > > Back with more issues. Periodically, and seemingly  
> > > > > > > > > > > > > > > inexplicably, the whole cluster becomes essentially unusable for what's  
> > > > > > > > > > > > > > > sometimes hours. There's nothing in the logs (or slow log) to indicate  
> > > > > > > > > > > > > > > what's up, and `hot_threads` (  
> > > > > > > > > > > > > > > [https://gist.github.com/4280439](https://gist.github.com/4280439)) on these machines has been  
> > > > > > > > > > > > > > > more than a little vague. Each of the machines' CPU usage gets pegged at  
> > > > > > > > > > > > > > > 100% of all cores, and then gradually the different machines back off. From  
> > > > > > > > > > > > > > > paramedic:
> > > > > > > > > > > > > > > 
> > > > > > > > > > > > > > > [https://lh6.googleusercontent.com/-Y1G\_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png](https://lh6.googleusercontent.com/-Y1G_abxvcEY/UMpQSKkrtJI/AAAAAAAAAAg/NryuOOLXEAM/s1600/Screen+Shot+2012-12-13+at+2.01.26+PM.png)
> > > > > > > > > > > > > > > 
> > > > > > > > > > > > > > > I'm at a loss trying to explaining why this is happening

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 21, 2012, 10:30pm UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/17 "2012-12-21T22:30:00Z")

</div>

Hi Dan,

would be interesting to know about what ratio between JVM memory and other  
memory for the machine you had before and after tuning, to help users who  
are experiencing similar CPU hogs.

Thanks,

Jörg

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:58am UTC](https://discuss.elastic.co/t/sudden-unexplained-cpu-usage/10071/18 "2017-07-06T02:58:42Z")

</div>


