# High cpu usage on large ec2 nodes

**URL:** <https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528>\
**Category:** Elasticsearch\
**Created:** [January 28, 2013, 5:52pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528 "2013-01-28T17:52:19Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![rohit\_reddy](https://avatars.discourse-cdn.com/v4/letter/r/ed655f/32.png) [@rohit\_reddy](https://discuss.elastic.co/u/rohit_reddy)\
**Post date:** [January 28, 2013, 5:52pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528/1 "2013-01-28T17:52:19Z")

</div>

Hi,

I'm pretty new to elasticsearch, though i have extensively used lucene.  
We are currently migrating from lucene to elasticsearch in our project.

We create a basic elasticsearch setup on AWS cloud and are trying to test  
the performance of the same.

The configuration:  
EC2 Nodes - 2 Large nodes  
Shards - 5  
Replication - 1  
Memory settings - 4GB

We have created a basic index whose size is about 7GB. For the performance  
tests, we have pretty much maintained a constant index, ie., the index is  
not getting updated. There are no index events to the elasticsearch server.

Not we are bombarding \*each \*elasticsearch node with about 100 search  
requets per sec (using a single jmeter client for this). Each search query  
is a boolean query with 5-6 term query criteria.

For this load the CPU utilization is going upto 75%. The performance of  
each query is still good. One query took about\* 90ms\* to return the result.

We then reduced the shards to 3 and ran the same tests.  
The CPU usage remained the same but the performance degraded. Now each  
request took about _180ms_ to return the result.

We expected the results to improve since we reduced the number of shards.  
Not the opposite happened. Is this the expected result.  
And is the high CPU usage also expected?

Thanks  
Rohit

--

---

<div class="post-metadata">

**Author:** ![karmi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karmi/32/44951_2.png) [@karmi](https://discuss.elastic.co/u/karmi)\
**Post date:** [January 29, 2013, 7:42am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528/2 "2013-01-29T07:42:34Z")

</div>

In general, yes, decreasing the number of shards should improve search  
performance (less Lucene indices to search against), but I suspect in your  
benchmarking scenario, there are many variables and it's hard to keep them  
consistent:

- The m1.large instance type is quite small, in a sense it has lot of  
"neighbours" -- you never know who is doing what in the same rack
- The m2.xlarge is better in this sense, and also allows you to use the  
high I/O EBS volumes
- A _lot_ depends on the disk used for ES -- are you using the EBS-backed  
instance disk? The "physical" ephemeral disk for the instance? Extra EBS  
volume, possibly IOPS?
- Regarding the CPU, I'd say it's expected you'll saturate the resources of  
the machine at one point, and ~100 req/sec sounds kinda OK to me for the  
type of machine in question. You can use the `hot_threads` API to check  
where the time is  
spent: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-cluster-nodes-hot-threads.html)

Karel

On Monday, January 28, 2013 6:52:19 PM UTC+1, rohit reddy wrote:

> Hi,
> 
> I'm pretty new to elasticsearch, though i have extensively used lucene.  
> We are currently migrating from lucene to elasticsearch in our project.
> 
> We create a basic elasticsearch setup on AWS cloud and are trying to test  
> the performance of the same.
> 
> The configuration:  
> EC2 Nodes - 2 Large nodes  
> Shards - 5  
> Replication - 1  
> Memory settings - 4GB
> 
> We have created a basic index whose size is about 7GB. For the performance  
> tests, we have pretty much maintained a constant index, ie., the index is  
> not getting updated. There are no index events to the elasticsearch server.
> 
> Not we are bombarding \*each \*elasticsearch node with about 100 search  
> requets per sec (using a single jmeter client for this). Each search query  
> is a boolean query with 5-6 term query criteria.
> 
> For this load the CPU utilization is going upto 75%. The performance of  
> each query is still good. One query took about\* 90ms\* to return the  
> result.
> 
> We then reduced the shards to 3 and ran the same tests.  
> The CPU usage remained the same but the performance degraded. Now each  
> request took about _180ms_ to return the result.
> 
> We expected the results to improve since we reduced the number of shards.  
> Not the opposite happened. Is this the expected result.  
> And is the high CPU usage also expected?
> 
> Thanks  
> Rohit

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![rohit\_reddy](https://avatars.discourse-cdn.com/v4/letter/r/ed655f/32.png) [@rohit\_reddy](https://discuss.elastic.co/u/rohit_reddy)\
**Post date:** [January 31, 2013, 6:07pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528/3 "2013-01-31T18:07:19Z")

</div>

We are using ephemeral disk with s3 backup. Since we expect the performance  
of ephemeral disk to be better than EBS. And since our index does not get  
updated too frequently, the overhead of storing backups in S3 is not huge.

I'll see use the API and try to identify which resource is taking up the  
CPU.

Thanks  
Rohit

On Tuesday, January 29, 2013 1:12:34 PM UTC+5:30, Karel Minařík wrote:

> In general, yes, decreasing the number of shards should improve search  
> performance (less Lucene indices to search against), but I suspect in your  
> benchmarking scenario, there are many variables and it's hard to keep them  
> consistent:
> 
> - The m1.large instance type is quite small, in a sense it has lot of  
> "neighbours" -- you never know who is doing what in the same rack
> - The m2.xlarge is better in this sense, and also allows you to use the  
> high I/O EBS volumes
> - A _lot_ depends on the disk used for ES -- are you using the EBS-backed  
> instance disk? The "physical" ephemeral disk for the instance? Extra EBS  
> volume, possibly IOPS?
> - Regarding the CPU, I'd say it's expected you'll saturate the resources  
> of the machine at one point, and ~100 req/sec sounds kinda OK to me for the  
> type of machine in question. You can use the `hot_threads` API to check  
> where the time is spent:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-cluster-nodes-hot-threads.html)
> 
> Karel
> 
> On Monday, January 28, 2013 6:52:19 PM UTC+1, rohit reddy wrote:
> 
> > Hi,
> > 
> > I'm pretty new to elasticsearch, though i have extensively used lucene.  
> > We are currently migrating from lucene to elasticsearch in our project.
> > 
> > We create a basic elasticsearch setup on AWS cloud and are trying to test  
> > the performance of the same.
> > 
> > The configuration:  
> > EC2 Nodes - 2 Large nodes  
> > Shards - 5  
> > Replication - 1  
> > Memory settings - 4GB
> > 
> > We have created a basic index whose size is about 7GB. For the  
> > performance tests, we have pretty much maintained a constant index, ie.,  
> > the index is not getting updated. There are no index events to the  
> > elasticsearch server.
> > 
> > Not we are bombarding \*each \*elasticsearch node with about 100 search  
> > requets per sec (using a single jmeter client for this). Each search query  
> > is a boolean query with 5-6 term query criteria.
> > 
> > For this load the CPU utilization is going upto 75%. The performance of  
> > each query is still good. One query took about\* 90ms\* to return the  
> > result.
> > 
> > We then reduced the shards to 3 and ran the same tests.  
> > The CPU usage remained the same but the performance degraded. Now each  
> > request took about _180ms_ to return the result.
> > 
> > We expected the results to improve since we reduced the number of shards.  
> > Not the opposite happened. Is this the expected result.  
> > And is the high CPU usage also expected?
> > 
> > Thanks  
> > Rohit

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![rohit\_reddy](https://avatars.discourse-cdn.com/v4/letter/r/ed655f/32.png) [@rohit\_reddy](https://discuss.elastic.co/u/rohit_reddy)\
**Post date:** [February 7, 2013, 9:01am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528/4 "2013-02-07T09:01:47Z")

</div>

Attached the hot-thread snapshot using the elasticsearch api.  
I'm using DFS\_QUERY\_THEN\_FETCH for the search.

> <https://gist.github.com/rohitreddy/4729660>

Seems like most of the threads on waiting on reading from lucene index. Is  
the normal? or should i tweek some configurations to reduce this. Using all  
defaults for now.

On Thursday, January 31, 2013 11:37:19 PM UTC+5:30, rohit reddy wrote:

> We are using ephemeral disk with s3 backup. Since we expect the  
> performance of ephemeral disk to be better than EBS. And since our index  
> does not get updated too frequently, the overhead of storing backups in S3  
> is not huge.
> 
> I'll see use the API and try to identify which resource is taking up the  
> CPU.
> 
> Thanks  
> Rohit
> 
> On Tuesday, January 29, 2013 1:12:34 PM UTC+5:30, Karel Minařík wrote:
> 
> > In general, yes, decreasing the number of shards should improve search  
> > performance (less Lucene indices to search against), but I suspect in your  
> > benchmarking scenario, there are many variables and it's hard to keep them  
> > consistent:
> > 
> > - The m1.large instance type is quite small, in a sense it has lot of  
> > "neighbours" -- you never know who is doing what in the same rack
> > - The m2.xlarge is better in this sense, and also allows you to use the  
> > high I/O EBS volumes
> > - A _lot_ depends on the disk used for ES -- are you using the EBS-backed  
> > instance disk? The "physical" ephemeral disk for the instance? Extra EBS  
> > volume, possibly IOPS?
> > - Regarding the CPU, I'd say it's expected you'll saturate the resources  
> > of the machine at one point, and ~100 req/sec sounds kinda OK to me for the  
> > type of machine in question. You can use the `hot_threads` API to check  
> > where the time is spent:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-cluster-nodes-hot-threads.html)
> > 
> > Karel
> > 
> > On Monday, January 28, 2013 6:52:19 PM UTC+1, rohit reddy wrote:
> > 
> > > Hi,
> > > 
> > > I'm pretty new to elasticsearch, though i have extensively used lucene.  
> > > We are currently migrating from lucene to elasticsearch in our project.
> > > 
> > > We create a basic elasticsearch setup on AWS cloud and are trying to  
> > > test the performance of the same.
> > > 
> > > The configuration:  
> > > EC2 Nodes - 2 Large nodes  
> > > Shards - 5  
> > > Replication - 1  
> > > Memory settings - 4GB
> > > 
> > > We have created a basic index whose size is about 7GB. For the  
> > > performance tests, we have pretty much maintained a constant index, ie.,  
> > > the index is not getting updated. There are no index events to the  
> > > elasticsearch server.
> > > 
> > > Not we are bombarding \*each \*elasticsearch node with about 100 search  
> > > requets per sec (using a single jmeter client for this). Each search query  
> > > is a boolean query with 5-6 term query criteria.
> > > 
> > > For this load the CPU utilization is going upto 75%. The performance of  
> > > each query is still good. One query took about\* 90ms\* to return the  
> > > result.
> > > 
> > > We then reduced the shards to 3 and ran the same tests.  
> > > The CPU usage remained the same but the performance degraded. Now each  
> > > request took about _180ms_ to return the result.
> > > 
> > > We expected the results to improve since we reduced the number of  
> > > shards. Not the opposite happened. Is this the expected result.  
> > > And is the high CPU usage also expected?
> > > 
> > > Thanks  
> > > Rohit

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 12, 2013, 10:35pm UTC](https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528/5 "2013-02-12T22:35:47Z")

</div>

It seems like its waiting most of the time on read. Which instance type of AWS are you using? Make sure to have ~50% of the memory allocated to ES (ES\_HEAP\_SIZE), and the other half to the OS.

Also, which java version are you using? Make sure you are on the latest 1.6 (update 34 and above) or 1.7. This makes a big difference and in older Linux distro, the default java provided is pretty old (4 years old).

I will shy away from the DFS\_ type, typically, you don't really need it with big enough data set.

Last, the reason why more shards performed better is because, even on 2 nodes, each search request was being parallelized across more shards (and less data). Note, if you start running concurrent client tests, make sure to configure the search thread pool with a fixed thread size of about 4 times the CPUs you have, so it won't overflow the concurrent execution.

On Feb 7, 2013, at 10:01 AM, rohit reddy [rohit.kommareddy@gmail.com](mailto:rohit.kommareddy@gmail.com) wrote:

> Attached the hot-thread snapshot using the elasticsearch api.  
> I'm using DFS\_QUERY\_THEN\_FETCH for the search.
> 
> [Thread stack - elasticsearch · GitHub](https://gist.github.com/rohitreddy/4729660)
> 
> Seems like most of the threads on waiting on reading from lucene index. Is the normal? or should i tweek some configurations to reduce this. Using all defaults for now.
> 
> On Thursday, January 31, 2013 11:37:19 PM UTC+5:30, rohit reddy wrote:  
> We are using ephemeral disk with s3 backup. Since we expect the performance of ephemeral disk to be better than EBS. And since our index does not get updated too frequently, the overhead of storing backups in S3 is not huge.
> 
> I'll see use the API and try to identify which resource is taking up the CPU.
> 
> Thanks  
> Rohit
> 
> On Tuesday, January 29, 2013 1:12:34 PM UTC+5:30, Karel Minařík wrote:  
> In general, yes, decreasing the number of shards should improve search performance (less Lucene indices to search against), but I suspect in your benchmarking scenario, there are many variables and it's hard to keep them consistent:
> 
> - The m1.large instance type is quite small, in a sense it has lot of "neighbours" -- you never know who is doing what in the same rack
> - The m2.xlarge is better in this sense, and also allows you to use the high I/O EBS volumes
> - A _lot_ depends on the disk used for ES -- are you using the EBS-backed instance disk? The "physical" ephemeral disk for the instance? Extra EBS volume, possibly IOPS?
> - Regarding the CPU, I'd say it's expected you'll saturate the resources of the machine at one point, and ~100 req/sec sounds kinda OK to me for the type of machine in question. You can use the `hot_threads` API to check where the time is spent: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/admin-cluster-nodes-hot-threads.html)
> 
> Karel
> 
> On Monday, January 28, 2013 6:52:19 PM UTC+1, rohit reddy wrote:  
> Hi,
> 
> I'm pretty new to elasticsearch, though i have extensively used lucene.  
> We are currently migrating from lucene to elasticsearch in our project.
> 
> We create a basic elasticsearch setup on AWS cloud and are trying to test the performance of the same.
> 
> The configuration:  
> EC2 Nodes - 2 Large nodes  
> Shards - 5  
> Replication - 1  
> Memory settings - 4GB
> 
> We have created a basic index whose size is about 7GB. For the performance tests, we have pretty much maintained a constant index, ie., the index is not getting updated. There are no index events to the elasticsearch server.
> 
> Not we are bombarding each elasticsearch node with about 100 search requets per sec (using a single jmeter client for this). Each search query is a boolean query with 5-6 term query criteria.
> 
> For this load the CPU utilization is going upto 75%. The performance of each query is still good. One query took about 90ms to return the result.
> 
> We then reduced the shards to 3 and ran the same tests.  
> The CPU usage remained the same but the performance degraded. Now each request took about 180ms to return the result.
> 
> We expected the results to improve since we reduced the number of shards. Not the opposite happened. Is this the expected result.  
> And is the high CPU usage also expected?
> 
> Thanks  
> Rohit
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:51am UTC](https://discuss.elastic.co/t/high-cpu-usage-on-large-ec2-nodes/10528/6 "2017-07-06T02:51:48Z")

</div>


