# Constant High (~99%) CPU on 1 of 5 Nodes in Cluster

**URL:** <https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961>\
**Category:** Elasticsearch\
**Created:** [July 29, 2014, 8:59pm UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961 "2014-07-29T20:59:52Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![michael\_4](https://avatars.discourse-cdn.com/v4/letter/m/cc9497/32.png) [@michael\_4](https://discuss.elastic.co/u/michael_4)\
**Post date:** [July 29, 2014, 8:59pm UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/1 "2014-07-29T20:59:52Z")

</div>

Hey guys,

We've been running a 5 node cluster for our index (5 shards, 1 replica,  
evenly distributed on 5 nodes), and are running into a problem with one of  
the nodes in the cluster. It is not unique to any specific node, and can  
happen sporadically on any of the nodes.

One of the machines starts spiking up close to 100% CPU Load, and close to  
8 OS Load (which is amusing, considering there are only 4 CPU cores on the  
machine), while all the other machines operate normally way below those  
figures. Naturally, this behavior is accompanied by extremely high write  
times, and read times, as well.

Here's what Marvel looks like:

[https://lh5.googleusercontent.com/-bxUFPhqAnVk/U9gKg4c19nI/AAAAAAAAABE/S\_w68vZ63Uo/s1600/Marvel+-+Node+Statistics.png](https://lh5.googleusercontent.com/-bxUFPhqAnVk/U9gKg4c19nI/AAAAAAAAABE/S_w68vZ63Uo/s1600/Marvel+-+Node+Statistics.png)

Here's all the information we could gather:

- Full thread dump from while this  
occurred: [https://gist.github.com/danielschonfeld/ff6c3744197f2c748632](https://gist.github.com/danielschonfeld/ff6c3744197f2c748632)
- GET  
\_nodes/stats: [https://gist.github.com/schonfeld/693c8dbf0dd57e4cff7c](https://gist.github.com/schonfeld/693c8dbf0dd57e4cff7c)
- GET  
\_nodes/hot\_threads: [https://gist.github.com/schonfeld/766d771d211e452a7100](https://gist.github.com/schonfeld/766d771d211e452a7100)
- GET  
\_cluster/stats: [https://gist.github.com/schonfeld/d5395f97e3a87745cc1f](https://gist.github.com/schonfeld/d5395f97e3a87745cc1f)

Thoughts? insights? Any clues would be greatly appreciated.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/05b552dc-70fe-4b76-abfb-eb9db2a9dd34%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/05b552dc-70fe-4b76-abfb-eb9db2a9dd34%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Kireet\_Reddy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kireet_reddy/32/1071_2.png) [@Kireet\_Reddy](https://discuss.elastic.co/u/Kireet_Reddy)\
**Post date:** [July 30, 2014, 2:33am UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/2 "2014-07-30T02:33:09Z")

</div>

We've had a very similar issue, but haven't been able to figure out what  
the problem is. How do you "fix" the problem? Will a node restart fix the  
problem immediately or do you need to restart the whole machine?

On Tuesday, July 29, 2014 1:59:52 PM UTC-7, [mic...@modernmast.com](mailto:mic...@modernmast.com) wrote:

> Hey guys,
> 
> We've been running a 5 node cluster for our index (5 shards, 1 replica,  
> evenly distributed on 5 nodes), and are running into a problem with one of  
> the nodes in the cluster. It is not unique to any specific node, and can  
> happen sporadically on any of the nodes.
> 
> One of the machines starts spiking up close to 100% CPU Load, and close to  
> 8 OS Load (which is amusing, considering there are only 4 CPU cores on the  
> machine), while all the other machines operate normally way below those  
> figures. Naturally, this behavior is accompanied by extremely high write  
> times, and read times, as well.
> 
> Here's what Marvel looks like:
> 
> [https://lh5.googleusercontent.com/-bxUFPhqAnVk/U9gKg4c19nI/AAAAAAAAABE/S\_w68vZ63Uo/s1600/Marvel+-+Node+Statistics.png](https://lh5.googleusercontent.com/-bxUFPhqAnVk/U9gKg4c19nI/AAAAAAAAABE/S_w68vZ63Uo/s1600/Marvel+-+Node+Statistics.png)
> 
> Here's all the information we could gather:
> 
> - Full thread dump from while this occurred:  
> [gist:ff6c3744197f2c748632 · GitHub](https://gist.github.com/danielschonfeld/ff6c3744197f2c748632)
> - GET \_nodes/stats:  
> [GET \_nodes/stats · GitHub](https://gist.github.com/schonfeld/693c8dbf0dd57e4cff7c)
> - GET \_nodes/hot\_threads:  
> [GET \_nodes/hot\_threads · GitHub](https://gist.github.com/schonfeld/766d771d211e452a7100)
> - GET \_cluster/stats:  
> [GET \_cluster/stats · GitHub](https://gist.github.com/schonfeld/d5395f97e3a87745cc1f)
> 
> Thoughts? insights? Any clues would be greatly appreciated.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/6660d09b-98f9-41c4-87d0-9ee56890c7b9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/6660d09b-98f9-41c4-87d0-9ee56890c7b9%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Greg\_Murnane](https://avatars.discourse-cdn.com/v4/letter/g/b9bd4f/32.png) [@Greg\_Murnane](https://discuss.elastic.co/u/Greg_Murnane)\
**Post date:** [August 1, 2014, 2:28pm UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/3 "2014-08-01T14:28:42Z")

</div>

From the Marvel image, it looks like the heap utilization isn't dropping  
periodically as it does on the other nodes. Can you verify that GC is  
behaving nicely while this occurs?

--  
The information transmitted in this email is intended only for the  
person(s) or entity to which it is addressed and may contain confidential  
and/or privileged material. Any review, retransmission, dissemination or  
other use of, or taking of any action in reliance upon, this information by  
persons or entities other than the intended recipient is prohibited. If you  
received this email in error, please contact the sender and permanently  
delete the email from any computer.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/67a3ebba-540b-49f3-ae33-20bc1705c4e6%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/67a3ebba-540b-49f3-ae33-20bc1705c4e6%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Luis\_Garcia\_Acosta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/luis_garcia_acosta/32/1266_2.png) [@Luis\_Garcia\_Acosta](https://discuss.elastic.co/u/Luis_Garcia_Acosta)\
**Post date:** [August 2, 2014, 9:40am UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/4 "2014-08-02T09:40:35Z")

</div>

Sorry for hickhacking the post with more questions, but why is the memory going all the way up and then dropping for the other nodes, is that normal?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/db6237d5-2ec2-4a02-b451-e69934d01691%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/db6237d5-2ec2-4a02-b451-e69934d01691%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![smonasco\_2](https://avatars.discourse-cdn.com/v4/letter/s/848f3c/32.png) [@smonasco\_2](https://discuss.elastic.co/u/smonasco_2)\
**Post date:** [August 3, 2014, 6:53am UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/5 "2014-08-03T06:53:42Z")

</div>

What jumps out at me is your that the CPU work you're doing seems to be very index related, your garbage collections are trying hard on the errant machine and not getting anywhere and you have a lot of deleted docs.

Tell us about your indexing strategy? Tell us things like routing, how bursty it is and maybe why you have so many deleted docs.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/2a79ccc0-d2cd-4d27-b66e-715d709c838b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/2a79ccc0-d2cd-4d27-b66e-715d709c838b%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![smonasco\_2](https://avatars.discourse-cdn.com/v4/letter/s/848f3c/32.png) [@smonasco\_2](https://discuss.elastic.co/u/smonasco_2)\
**Post date:** [August 3, 2014, 6:55am UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/6 "2014-08-03T06:55:02Z")

</div>

Also is it always the master node that goes awry?  
On Aug 3, 2014 12:54 AM, "smonasco" [smonasco@gmail.com](mailto:smonasco@gmail.com) wrote:

> What jumps out at me is your that the CPU work you're doing seems to be  
> very index related, your garbage collections are trying hard on the errant  
> machine and not getting anywhere and you have a lot of deleted docs.
> 
> Tell us about your indexing strategy? Tell us things like routing, how  
> bursty it is and maybe why you have so many deleted docs.
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/e1RBjvFSKGU/unsubscribe](https://groups.google.com/d/topic/elasticsearch/e1RBjvFSKGU/unsubscribe).  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/2a79ccc0-d2cd-4d27-b66e-715d709c838b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/2a79ccc0-d2cd-4d27-b66e-715d709c838b%40googlegroups.com)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAFDU5WKEx-jwUVKqrOkXLzfKiOKgUrphzKec6S7E5Z%2Bsm-YjDA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAFDU5WKEx-jwUVKqrOkXLzfKiOKgUrphzKec6S7E5Z%2Bsm-YjDA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![michael\_4](https://avatars.discourse-cdn.com/v4/letter/m/cc9497/32.png) [@michael\_4](https://discuss.elastic.co/u/michael_4)\
**Post date:** [August 4, 2014, 4:45pm UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/7 "2014-08-04T16:45:48Z")

</div>

Thanks everyone for replying. As it turns out, all our problems stemmed  
from our index schema.

Since our app was heavily modeled after social networks, we had to store  
our users' followers, and their IDs. To do that, each of our users had an  
array called "follower\_ids" -- the IDs of the people who are following our  
user. Now, that's all fine, until a user like the NBA, or Pepsi comes in  
with millions and millions of followers, and that array turns into an  
immensely giant array. We also tried turning the array into a nested object  
of [{id: 1,}, {id: 2}, ...], but because of the indexing strategy, that  
ended up even worse.

We've pinpointed the problem to the part in which we add IDs to the  
follower\_ids array. Ultimately, we swapped the schema around -- instead of  
storing a giant "follower\_ids", we started storing "following\_ids" --  
meaning, each person's document stores which users they follow.

Our current schema works great! CPU never goes above 25%, OS Load stays  
consistent, and our cluster is functioning super fast.

On Sunday, August 3, 2014 2:55:21 AM UTC-4, smonasco wrote:

> Also is it always the master node that goes awry?  
> On Aug 3, 2014 12:54 AM, "smonasco" \<[smon...@gmail.com](mailto:smon...@gmail.com) \<javascript:\>\>  
> wrote:
> 
> > What jumps out at me is your that the CPU work you're doing seems to be  
> > very index related, your garbage collections are trying hard on the errant  
> > machine and not getting anywhere and you have a lot of deleted docs.
> > 
> > Tell us about your indexing strategy? Tell us things like routing, how  
> > bursty it is and maybe why you have so many deleted docs.
> > 
> > --  
> > You received this message because you are subscribed to a topic in the  
> > Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit  
> > [https://groups.google.com/d/topic/elasticsearch/e1RBjvFSKGU/unsubscribe](https://groups.google.com/d/topic/elasticsearch/e1RBjvFSKGU/unsubscribe).  
> > To unsubscribe from this group and all its topics, send an email to  
> > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/2a79ccc0-d2cd-4d27-b66e-715d709c838b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/2a79ccc0-d2cd-4d27-b66e-715d709c838b%40googlegroups.com)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/eecf7019-3ecc-4ad2-805b-a6c16cdd2122%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/eecf7019-3ecc-4ad2-805b-a6c16cdd2122%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:10am UTC](https://discuss.elastic.co/t/constant-high-99-cpu-on-1-of-5-nodes-in-cluster/18961/8 "2017-07-06T01:10:57Z")

</div>


