# Indexing slows down dramatically as index size grows

**URL:** https://discuss.elastic.co/t/indexing-slows-down-dramatically-as-index-size-grows/7390
**Category:** Elasticsearch
**Created:** [April 18, 2012, 6:13pm UTC](https://discuss.elastic.co/t/indexing-slows-down-dramatically-as-index-size-grows/7390 "2012-04-18T18:13:41Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![Gregory\_Rice](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gregory_rice/32/2884_2.png) [@Gregory\_Rice](https://discuss.elastic.co/u/Gregory_Rice)
#### Post date: [April 18, 2012, 6:13pm UTC](https://discuss.elastic.co/t/indexing-slows-down-dramatically-as-index-size-grows/7390/1 "2012-04-18T18:13:41Z")

</div>

Hey guys,

I've got a two-node ES cluster running on two four-core HPs and  
feeding data into a pretty fast datastore.

When I start my logstash -\> ElasticSearch pipeline, I'm able to get  
about 1200 documents indexed per second. This is far in excess of the  
speed at which they're coming in over the AMQP pipe, so it's  
sufficient for my needs.

After a few hours of running, however, performance drops to a mere  
60-80 documents indexed per second. Why does this happen, and what can  
I do to fix it? Emptying out my data.path location speeds things back  
up immediately, but wipes everything out in the process, of course.

Thanks,  
Greg Rice

---

<div class="post-metadata">

### Author: ![Radu\_Gheorghe1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe1/32/2688_2.png) [@Radu\_Gheorghe1](https://discuss.elastic.co/u/Radu_Gheorghe1)
#### Post date: [April 19, 2012, 6:51am UTC](https://discuss.elastic.co/t/indexing-slows-down-dramatically-as-index-size-grows/7390/2 "2012-04-19T06:51:29Z")

</div>

Hi Greg,

How much memory do you have allocated for ES? The rule of thumb would  
be around half the total amount of RAM.

Also, can you check the logs to see if you run into a max-open-files  
limit?

Also, it might be useful to take a look at this thread:  
[http://elasticsearch-users.115913.n3.nabble.com/Problems-with-GrayLog2-ES-setup-long-td3859362.html](http://elasticsearch-users.115913.n3.nabble.com/Problems-with-GrayLog2-ES-setup-long-td3859362.html)

I suppose that optimizations for Graylog would also apply to Logstash.

On Apr 18, 9:13 pm, Greg Rice [gregr...@gmail.com](mailto:gregr...@gmail.com) wrote:

> Hey guys,
> 
> I've got a two-node ES cluster running on two four-core HPs and  
> feeding data into a pretty fast datastore.
> 
> When I start my logstash -\> Elasticsearch pipeline, I'm able to get  
> about 1200 documents indexed per second. This is far in excess of the  
> speed at which they're coming in over the AMQP pipe, so it's  
> sufficient for my needs.
> 
> After a few hours of running, however, performance drops to a mere  
> 60-80 documents indexed per second. Why does this happen, and what can  
> I do to fix it? Emptying out my data.path location speeds things back  
> up immediately, but wipes everything out in the process, of course.
> 
> Thanks,  
> Greg Rice

---

<div class="post-metadata">

### Author: ![Gregory\_Rice](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gregory_rice/32/2884_2.png) [@Gregory\_Rice](https://discuss.elastic.co/u/Gregory_Rice)
#### Post date: [April 19, 2012, 7:07pm UTC](https://discuss.elastic.co/t/indexing-slows-down-dramatically-as-index-size-grows/7390/3 "2012-04-19T19:07:33Z")

</div>

Radu,

I'm seeing lots of I/O wait (40-75%) on the nodes. Increasing to  
ES\_MAX\_MEM="8g" ES\_HEAP\_SIZE="8g" has made the bursts of activity on  
disc more productive. This is all happening on a shared storage system  
with all of our VM data on it, mounted to an ESXi as a VM disk over  
NFS.

I'm able to just barely keep up with input now, but is I/O wait like  
that normal for ES? Would I be a whole ton better off with local  
storage?

Thanks,  
Greg Rice

On Apr 18, 11:51 pm, Radu Gheorghe [radu0gheor...@gmail.com](mailto:radu0gheor...@gmail.com) wrote:

> Hi Greg,
> 
> How much memory do you have allocated for ES? The rule of thumb would  
> be around half the total amount of RAM.
> 
> Also, can you check the logs to see if you run into a max-open-files  
> limit?
> 
> Also, it might be useful to take a look at this thread:[http://elasticsearch-users.115913.n3.nabble.com/Problems-with-GrayLog](http://elasticsearch-users.115913.n3.nabble.com/Problems-with-GrayLog)...
> 
> I suppose that optimizations for Graylog would also apply to Logstash.
> 
> On Apr 18, 9:13 pm, Greg Rice [gregr...@gmail.com](mailto:gregr...@gmail.com) wrote:
> 
> > Hey guys,
> 
> > I've got a two-node ES cluster running on two four-core HPs and  
> > feeding data into a pretty fast datastore.
> 
> > When I start my logstash -\> Elasticsearch pipeline, I'm able to get  
> > about 1200 documents indexed per second. This is far in excess of the  
> > speed at which they're coming in over the AMQP pipe, so it's  
> > sufficient for my needs.
> 
> > After a few hours of running, however, performance drops to a mere  
> > 60-80 documents indexed per second. Why does this happen, and what can  
> > I do to fix it? Emptying out my data.path location speeds things back  
> > up immediately, but wipes everything out in the process, of course.
> 
> > Thanks,  
> > Greg Rice

---

<div class="post-metadata">

### Author: ![Radu\_Gheorghe1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe1/32/2688_2.png) [@Radu\_Gheorghe1](https://discuss.elastic.co/u/Radu_Gheorghe1)
#### Post date: [April 20, 2012, 6:55am UTC](https://discuss.elastic.co/t/indexing-slows-down-dramatically-as-index-size-grows/7390/4 "2012-04-20T06:55:31Z")

</div>

Hi Gregory,

I think you would be better off with local storage, yes. I'm having  
the same problem, with slow inserts from VMs with shared storage.

I'm not sure what can be done from the ES side, especially since other  
databases seem to suffer as well in the same conditions. Which kind of  
makes sense: you have to dump your changes to disk every once in a  
while.

There might be some settings to optimize this, but I'm not aware of  
one that helps. All I found was here:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

If you find any solution, please write it here as well. I would really  
appreciate it 🙂

Best regards,  
Radu

On Apr 19, 10:07 pm, Gregory Rice [gregr...@gmail.com](mailto:gregr...@gmail.com) wrote:

> Radu,
> 
> I'm seeing lots of I/O wait (40-75%) on the nodes. Increasing to  
> ES\_MAX\_MEM="8g" ES\_HEAP\_SIZE="8g" has made the bursts of activity on  
> disc more productive. This is all happening on a shared storage system  
> with all of our VM data on it, mounted to an ESXi as a VM disk over  
> NFS.
> 
> I'm able to just barely keep up with input now, but is I/O wait like  
> that normal for ES? Would I be a whole ton better off with local  
> storage?
> 
> Thanks,  
> Greg Rice
> 
> On Apr 18, 11:51 pm, Radu Gheorghe [radu0gheor...@gmail.com](mailto:radu0gheor...@gmail.com) wrote:
> 
> > Hi Greg,
> 
> > How much memory do you have allocated for ES? The rule of thumb would  
> > be around half the total amount of RAM.
> 
> > Also, can you check the logs to see if you run into a max-open-files  
> > limit?
> 
> > Also, it might be useful to take a look at this thread:[http://elasticsearch-users.115913.n3.nabble.com/Problems-with-GrayLog](http://elasticsearch-users.115913.n3.nabble.com/Problems-with-GrayLog)...
> 
> > I suppose that optimizations for Graylog would also apply to Logstash.
> 
> > On Apr 18, 9:13 pm, Greg Rice [gregr...@gmail.com](mailto:gregr...@gmail.com) wrote:
> 
> > > Hey guys,
> 
> > > I've got a two-node ES cluster running on two four-core HPs and  
> > > feeding data into a pretty fast datastore.
> 
> > > When I start my logstash -\> Elasticsearch pipeline, I'm able to get  
> > > about 1200 documents indexed per second. This is far in excess of the  
> > > speed at which they're coming in over the AMQP pipe, so it's  
> > > sufficient for my needs.
> 
> > > After a few hours of running, however, performance drops to a mere  
> > > 60-80 documents indexed per second. Why does this happen, and what can  
> > > I do to fix it? Emptying out my data.path location speeds things back  
> > > up immediately, but wipes everything out in the process, of course.
> 
> > > Thanks,  
> > > Greg Rice

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 3:31am UTC](https://discuss.elastic.co/t/indexing-slows-down-dramatically-as-index-size-grows/7390/5 "2017-07-06T03:31:43Z")

</div>


