# High CPU use and CouchDB river outage at times

**URL:** <https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927>\
**Category:** Elasticsearch\
**Created:** [February 27, 2013, 2:59pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927 "2013-02-27T14:59:21Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Milan\_Gornik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/milan_gornik/32/2605_2.png) [@Milan\_Gornik](https://discuss.elastic.co/u/Milan_Gornik)\
**Post date:** [February 27, 2013, 2:59pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/1 "2013-02-27T14:59:21Z")

</div>

Hi guys,

We are running ElasticSearch version 0.20.5 on two nodes. Right now we  
store around 95 million documents in it. We are getting a lot of new  
documents each day, so we are trimming the dataset by deleting outdated  
data. Each day it deletes around 2M of old documents. I noticed  
deleted\_docs property (in docs object of index status) is growing while we  
are deleting. We had couple of incidents lately during which our river that  
copies data from CouchDB to ElasticSearch would stop getting new data in.  
When we SSH into servers and check, it shows really high CPU use (up to  
900% for ES service). I restarted service on both nodes and after it system  
got back into normal (from yellow state of the cluster, back to green, CPU  
use dropped and got back to normal). I also noticed that deleted\_docs was  
cut in half after I restarted service.

In a separate topic I asked about running \_optimize and others explained  
that running it manually is not really needed. We let the ES to do merges  
based on internal logic. Is it possible that when ES decides it is time to  
optimize/compact it causes a big load which would cause effects as I  
described? If it is the case, is there a way for us to control how this  
process goes, to avoid the big impact this has on the system (if it really  
is a reason)? It is necessary for us to remain online during all times, so  
we are trying to determine if the way we are removing outdated data is  
killing our ES nodes and if there is a better way to do it.

Thanks a lot for your time!  
Milan Gornik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 27, 2013, 3:47pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/2 "2013-02-27T15:47:14Z")

</div>

As you receive data on a daily basis, how about using index aliases for  
the indices containing daily data? So, you will not need to delete  
documents, which is every expensive. With index alias, you can just  
re-assign the alias to the date period you actually want, and you can  
safely and quickly drop obsolete indices.

Jörg

Am 27.02.13 15:59, schrieb Milan Gornik:

> In a separate topic I asked about running \_optimize and others  
> explained that running it manually is not really needed. We let the ES  
> to do merges based on internal logic. Is it possible that when ES  
> decides it is time to optimize/compact it causes a big load which  
> would cause effects as I described? If it is the case, is there a way  
> for us to control how this process goes, to avoid the big impact this  
> has on the system (if it really is a reason)? It is necessary for us  
> to remain online during all times, so we are trying to determine if  
> the way we are removing outdated data is killing our ES nodes and if  
> there is a better way to do it.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Milan\_Gornik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/milan_gornik/32/2605_2.png) [@Milan\_Gornik](https://discuss.elastic.co/u/Milan_Gornik)\
**Post date:** [February 27, 2013, 4:31pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/3 "2013-02-27T16:31:51Z")

</div>

Hi Jörg,

I'm not sure if that would work for our use-case (I didn't provide enough  
details in the first post). We have a working set of last 40 days. That's  
the current data that we are using. And we are getting new records all the  
time (in the first post I probably oversimplified that). System is used  
24/7 and the new records are coming in constantly. We have a background  
script that runs each day with the same schedule and it selects everything  
older than 40 days and removes it. So, I'm not sure how we would achieve  
chopping off these expired documents with using indexes and aliases. I  
guess we would need to have separate index for each day and then one alias  
that joins all these indexes which are current (so every days except those  
older than 40 days). Then in the removal script, we would need to drop  
older indexes and recreate alias to include only current indexes. Is this  
possible?

Thanks!  
Milan

On Wednesday, February 27, 2013 4:47:14 PM UTC+1, Jörg Prante wrote:

> As you receive data on a daily basis, how about using index aliases for  
> the indices containing daily data? So, you will not need to delete  
> documents, which is every expensive. With index alias, you can just  
> re-assign the alias to the date period you actually want, and you can  
> safely and quickly drop obsolete indices.
> 
> Jörg
> 
> Am 27.02.13 15:59, schrieb Milan Gornik:
> 
> > In a separate topic I asked about running \_optimize and others  
> > explained that running it manually is not really needed. We let the ES  
> > to do merges based on internal logic. Is it possible that when ES  
> > decides it is time to optimize/compact it causes a big load which  
> > would cause effects as I described? If it is the case, is there a way  
> > for us to control how this process goes, to avoid the big impact this  
> > has on the system (if it really is a reason)? It is necessary for us  
> > to remain online during all times, so we are trying to determine if  
> > the way we are removing outdated data is killing our ES nodes and if  
> > there is a better way to do it.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Randall\_McRee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/randall_mcree/32/47177_2.png) [@Randall\_McRee](https://discuss.elastic.co/u/Randall_McRee)\
**Post date:** [February 27, 2013, 6:04pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/4 "2013-02-27T18:04:02Z")

</div>

In our production servers (live 24\*7 like yours) we call optimize  
max\_num\_segments setting it to 3 after updating/deleting out-of-date  
documents once an hour. See Elastic guide to  
optimize[http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html)

We have noticed that this api actually needs to be called twice in  
succession. You can tell if it actually succeeded (it always says it did)  
by monitoring the index in the data directory--the number of files will  
decrease proportionately.

Also we set index.merge.policy.segments\_per\_tier  
and index.merge.policy.max\_merge\_at\_once to 3 (default is 10).

We do this because we see performance advantages for querying during the  
remainder of the hour. But it should also provide the benefits that you  
seek, namely decreasing any large merges later in the day.

You can use the curl command forms to experiment.

Randy

On Wed, Feb 27, 2013 at 6:59 AM, Milan Gornik [milan@wildbit.com](mailto:milan@wildbit.com) wrote:

> Hi guys,
> 
> We are running Elasticsearch version 0.20.5 on two nodes. Right now we  
> store around 95 million documents in it. We are getting a lot of new  
> documents each day, so we are trimming the dataset by deleting outdated  
> data. Each day it deletes around 2M of old documents. I noticed  
> deleted\_docs property (in docs object of index status) is growing while we  
> are deleting. We had couple of incidents lately during which our river that  
> copies data from CouchDB to Elasticsearch would stop getting new data in.  
> When we SSH into servers and check, it shows really high CPU use (up to  
> 900% for ES service). I restarted service on both nodes and after it system  
> got back into normal (from yellow state of the cluster, back to green, CPU  
> use dropped and got back to normal). I also noticed that deleted\_docs was  
> cut in half after I restarted service.
> 
> In a separate topic I asked about running \_optimize and others explained  
> that running it manually is not really needed. We let the ES to do merges  
> based on internal logic. Is it possible that when ES decides it is time to  
> optimize/compact it causes a big load which would cause effects as I  
> described? If it is the case, is there a way for us to control how this  
> process goes, to avoid the big impact this has on the system (if it really  
> is a reason)? It is necessary for us to remain online during all times, so  
> we are trying to determine if the way we are removing outdated data is  
> killing our ES nodes and if there is a better way to do it.
> 
> Thanks a lot for your time!  
> Milan Gornik
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Milan\_Gornik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/milan_gornik/32/2605_2.png) [@Milan\_Gornik](https://discuss.elastic.co/u/Milan_Gornik)\
**Post date:** [February 28, 2013, 4:13pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/5 "2013-02-28T16:13:49Z")

</div>

Hi Randy,

Thanks a lot for sharing your experience on this topic here. One question:  
have you tried modifying these parameters on live system? I am interested  
to know if modifying them can cause issues for cluster which is currently  
running. I also read about Store Level Throttling  
here: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html),  
so these two combined should be able to fully control merge process in ES.

Regards,  
Milan

On Wednesday, February 27, 2013 7:04:02 PM UTC+1, RKM wrote:

> In our production servers (live 24\*7 like yours) we call optimize  
> max\_num\_segments setting it to 3 after updating/deleting out-of-date  
> documents once an hour. See Elastic guide to optimize[http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html)
> 
> We have noticed that this api actually needs to be called twice in  
> succession. You can tell if it actually succeeded (it always says it did)  
> by monitoring the index in the data directory--the number of files will  
> decrease proportionately.
> 
> Also we set index.merge.policy.segments\_per\_tier  
> and index.merge.policy.max\_merge\_at\_once to 3 (default is 10).
> 
> We do this because we see performance advantages for querying during the  
> remainder of the hour. But it should also provide the benefits that you  
> seek, namely decreasing any large merges later in the day.
> 
> You can use the curl command forms to experiment.
> 
> Randy
> 
> On Wed, Feb 27, 2013 at 6:59 AM, Milan Gornik \<[mi...@wildbit.com](mailto:mi...@wildbit.com)\<javascript:\>
> 
> > wrote:
> 
> > Hi guys,
> > 
> > We are running Elasticsearch version 0.20.5 on two nodes. Right now we  
> > store around 95 million documents in it. We are getting a lot of new  
> > documents each day, so we are trimming the dataset by deleting outdated  
> > data. Each day it deletes around 2M of old documents. I noticed  
> > deleted\_docs property (in docs object of index status) is growing while we  
> > are deleting. We had couple of incidents lately during which our river that  
> > copies data from CouchDB to Elasticsearch would stop getting new data in.  
> > When we SSH into servers and check, it shows really high CPU use (up to  
> > 900% for ES service). I restarted service on both nodes and after it system  
> > got back into normal (from yellow state of the cluster, back to green, CPU  
> > use dropped and got back to normal). I also noticed that deleted\_docs was  
> > cut in half after I restarted service.
> > 
> > In a separate topic I asked about running \_optimize and others explained  
> > that running it manually is not really needed. We let the ES to do merges  
> > based on internal logic. Is it possible that when ES decides it is time to  
> > optimize/compact it causes a big load which would cause effects as I  
> > described? If it is the case, is there a way for us to control how this  
> > process goes, to avoid the big impact this has on the system (if it really  
> > is a reason)? It is necessary for us to remain online during all times, so  
> > we are trying to determine if the way we are removing outdated data is  
> > killing our ES nodes and if there is a better way to do it.
> > 
> > Thanks a lot for your time!  
> > Milan Gornik
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Dan\_Fairs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_fairs/32/1205_2.png) [@Dan\_Fairs](https://discuss.elastic.co/u/Dan_Fairs)\
**Post date:** [February 28, 2013, 6:14pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/6 "2013-02-28T18:14:26Z")

</div>

Hi Milan,

> I'm not sure if that would work for our use-case (I didn't provide enough details in the first post). We have a working set of last 40 days. That's the current data that we are using. And we are getting new records all the time (in the first post I probably oversimplified that). System is used 24/7 and the new records are coming in constantly. We have a background script that runs each day with the same schedule and it selects everything older than 40 days and removes it.

I think I'd probably still support Jörg's multiple-index approach.

I'd have an index per day, with a name like docs-20130228. I'd then have an alias 'docs' that pointed to the 'current' day.

I'd modify my application so that it always wrote to an index called 'docs' (the alias will route this to the correct underlying index for the day) - I guess this is actually your CouchDB river, in this case.

I'd also have a script which ran every hour, which looked at the current date/time. If we're about to move to a new day, it would create a new index (with appropriate mappings etc.) then update the 'docs' alias to point to it. The application would then transparently start writing documents to that new index.

At that point, I can figure out what the name of the index for 41-days-ago was, and delete it.

I guess something like that is what Jörg was getting at.

Remember that searches can easily span multiple indices, so you don't limit your flexibility to search across your whole 40-day history with this approach.

## Cheers, Dan

Dan Fairs | [dan.fairs@gmail.com](mailto:dan.fairs@gmail.com) | @danfairs | [secondsync.com](http://secondsync.com)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Randall\_McRee](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/randall_mcree/32/47177_2.png) [@Randall\_McRee](https://discuss.elastic.co/u/Randall_McRee)\
**Post date:** [February 28, 2013, 8:29pm UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/7 "2013-02-28T20:29:57Z")

</div>

Yes we issue these commands on our live system, once an hour after a batch  
index. Other than optimize being somewhat flaky (i.e. saying success when  
in actuality the switch to the merged segment did not happen) all has been  
well.

We have not tried the store level throttling. We do have mmapfs set to  
true, however.

So, yes, there are a lot of knobs and switches you can throw.

On Thu, Feb 28, 2013 at 8:13 AM, Milan Gornik [milan@wildbit.com](mailto:milan@wildbit.com) wrote:

> Hi Randy,
> 
> Thanks a lot for sharing your experience on this topic here. One question:  
> have you tried modifying these parameters on live system? I am interested  
> to know if modifying them can cause issues for cluster which is currently  
> running. I also read about Store Level Throttling here:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html), so  
> these two combined should be able to fully control merge process in ES.
> 
> Regards,  
> Milan
> 
> On Wednesday, February 27, 2013 7:04:02 PM UTC+1, RKM wrote:
> 
> > In our production servers (live 24\*7 like yours) we call optimize  
> > max\_num\_segments setting it to 3 after updating/deleting out-of-date  
> > documents once an hour. See Elastic guide to optimize[http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html)
> > 
> > We have noticed that this api actually needs to be called twice in  
> > succession. You can tell if it actually succeeded (it always says it did)  
> > by monitoring the index in the data directory--the number of files will  
> > decrease proportionately.
> > 
> > Also we set index.merge.policy.\*\*segments\_per\_tier  
> > and index.merge.policy.max\_\*\*merge\_at\_once to 3 (default is 10).
> > 
> > We do this because we see performance advantages for querying during the  
> > remainder of the hour. But it should also provide the benefits that you  
> > seek, namely decreasing any large merges later in the day.
> > 
> > You can use the curl command forms to experiment.
> > 
> > Randy
> > 
> > On Wed, Feb 27, 2013 at 6:59 AM, Milan Gornik [mi...@wildbit.com](mailto:mi...@wildbit.com) wrote:
> > 
> > > Hi guys,
> > > 
> > > We are running Elasticsearch version 0.20.5 on two nodes. Right now we  
> > > store around 95 million documents in it. We are getting a lot of new  
> > > documents each day, so we are trimming the dataset by deleting outdated  
> > > data. Each day it deletes around 2M of old documents. I noticed  
> > > deleted\_docs property (in docs object of index status) is growing while we  
> > > are deleting. We had couple of incidents lately during which our river that  
> > > copies data from CouchDB to Elasticsearch would stop getting new data in.  
> > > When we SSH into servers and check, it shows really high CPU use (up to  
> > > 900% for ES service). I restarted service on both nodes and after it system  
> > > got back into normal (from yellow state of the cluster, back to green, CPU  
> > > use dropped and got back to normal). I also noticed that deleted\_docs was  
> > > cut in half after I restarted service.
> > > 
> > > In a separate topic I asked about running \_optimize and others explained  
> > > that running it manually is not really needed. We let the ES to do merges  
> > > based on internal logic. Is it possible that when ES decides it is time to  
> > > optimize/compact it causes a big load which would cause effects as I  
> > > described? If it is the case, is there a way for us to control how this  
> > > process goes, to avoid the big impact this has on the system (if it really  
> > > is a reason)? It is necessary for us to remain online during all times, so  
> > > we are trying to determine if the way we are removing outdated data is  
> > > killing our ES nodes and if there is a better way to do it.
> > > 
> > > Thanks a lot for your time!  
> > > Milan Gornik
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > 
> > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > .
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:48am UTC](https://discuss.elastic.co/t/high-cpu-use-and-couchdb-river-outage-at-times/10927/8 "2017-07-06T02:48:57Z")

</div>


