# Very high disk IO while indexing

**URL:** <https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199>\
**Category:** Elasticsearch\
**Created:** [March 19, 2013, 4:32am UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199 "2013-03-19T04:32:02Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bruno\_Miranda](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_miranda/32/746_2.png) [@Bruno\_Miranda](https://discuss.elastic.co/u/Bruno_Miranda)\
**Post date:** [March 19, 2013, 4:32am UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/1 "2013-03-19T04:32:02Z")

</div>

While indexing to our QA environment, 2 nodes (EC2 m1.large) each with 2  
CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
[http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)

I have about 20 processes indexing simultaneously. The index is 3.3 million  
documents and about 8GB on disk. The refresh internal is set at 5 minutes  
while the reindex is going, the index is 5 shards and 1 replica. I haven't  
changed the merge policy.

Cutting the concurrent process to 20 to 10 definitely lowers the IO from  
98% to around 50%.

Any tips on lowering disk IO? Is this normal? Reindex seems to finish  
properly, all records are properly indexed but the 100% disk IO scares me.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Bruno\_Miranda](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_miranda/32/746_2.png) [@Bruno\_Miranda](https://discuss.elastic.co/u/Bruno_Miranda)\
**Post date:** [March 19, 2013, 4:45am UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/2 "2013-03-19T04:45:05Z")

</div>

Here are the index  
\_settings: [https://gist.github.com/brupm/d7fe657a9501e617d46c](https://gist.github.com/brupm/d7fe657a9501e617d46c)

On Monday, March 18, 2013 9:32:02 PM UTC-7, Bruno Miranda wrote:

> While indexing to our QA environment, 2 nodes (EC2 m1.large) each with 2  
> CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
> [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)
> 
> I have about 20 processes indexing simultaneously. The index is 3.3  
> million documents and about 8GB on disk. The refresh internal is set at 5  
> minutes while the reindex is going, the index is 5 shards and 1 replica. I  
> haven't changed the merge policy.
> 
> Cutting the concurrent process to 20 to 10 definitely lowers the IO from  
> 98% to around 50%.
> 
> Any tips on lowering disk IO? Is this normal? Reindex seems to finish  
> properly, all records are properly indexed but the 100% disk IO scares me.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [March 19, 2013, 8:22am UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/3 "2013-03-19T08:22:42Z")

</div>

Beside tackling I/O with ES, you should also use disk monitoring, maybe  
there is a faulty disk drive... but, don't ask me how this works in EC2  
environment, I assume it's now difference to local disk management.

Jörg

Am 19.03.13 05:45, schrieb Bruno Miranda:

> Here are the index  
> \_settings: [https://gist.github.com/brupm/d7fe657a9501e617d46c](https://gist.github.com/brupm/d7fe657a9501e617d46c)
> 
> On Monday, March 18, 2013 9:32:02 PM UTC-7, Bruno Miranda wrote:
> 
> ```
> While indexing to our QA environment, 2 nodes (EC2 m1.large) each
> with 2 CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.
> http://cl.ly/image/371H01382h2O <http://cl.ly/image/371H01382h2O>
> 
> I have about 20 processes indexing simultaneously. The index is
> 3.3 million documents and about 8GB on disk. The refresh internal
> is set at 5 minutes while the reindex is going, the index is 5
> shards and 1 replica. I haven't changed the merge policy.
> Cutting the concurrent process to 20 to 10 definitely lowers the
> IO from 98% to around 50%.
> 
> Any tips on lowering disk IO? Is this normal? Reindex seems to
> finish properly, all records are properly indexed but the 100%
> disk IO scares me.
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google  
> Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send  
> an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [March 19, 2013, 8:32am UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/4 "2013-03-19T08:32:01Z")

</div>

Hey,

you could play around with throttling the merge throughput in order to  
lower disk utilization, see

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

In addition, try to check if you can decrease your index size as well  
(maybe by deactivating the all field, storing less fields, etc)

If I check your first graph correctly, the writing rate per second is not  
that high (but I do not know what is defined fast or slow on AWS either)

--Alex

On Tue, Mar 19, 2013 at 9:22 AM, Jörg Prante [joergprante@gmail.com](mailto:joergprante@gmail.com) wrote:

> Beside tackling I/O with ES, you should also use disk monitoring, maybe  
> there is a faulty disk drive... but, don't ask me how this works in EC2  
> environment, I assume it's now difference to local disk management.
> 
> Jörg
> 
> Am 19.03.13 05:45, schrieb Bruno Miranda:
> 
> > Here are the index \_settings: [brupm’s gists · GitHub](https://gist.github.com/brupm/)\*\*  
> > d7fe657a9501e617d46c [https://gist.github.com/brupm/d7fe657a9501e617d46c](https://gist.github.com/brupm/d7fe657a9501e617d46c)
> > 
> > On Monday, March 18, 2013 9:32:02 PM UTC-7, Bruno Miranda wrote:
> > 
> > ```
> > While indexing to our QA environment, 2 nodes (EC2 m1.large) each
> > with 2 CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.
> > http://cl.ly/image/**371H01382h2O <http://cl.ly/image/371H01382h2O> <
> > 
> > ```
> > 
> > [http://cl.ly/image/\*\*371H01382h2O](http://cl.ly/image/**371H01382h2O) [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)\>
> > 
> > ```
> > I have about 20 processes indexing simultaneously. The index is
> > 3.3 million documents and about 8GB on disk. The refresh internal
> > is set at 5 minutes while the reindex is going, the index is 5
> > shards and 1 replica. I haven't changed the merge policy.
> > Cutting the concurrent process to 20 to 10 definitely lowers the
> > IO from 98% to around 50%.
> > 
> > Any tips on lowering disk IO? Is this normal? Reindex seems to
> > finish properly, all records are properly indexed but the 100%
> > disk IO scares me.
> > 
> > ```
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to elasticsearch+unsubscribe@\*\*[googlegroups.com](http://googlegroups.com)[elasticsearch%2Bunsubscribe@googlegroups.com](mailto:elasticsearch%2Bunsubscribe@googlegroups.com)  
> > .  
> > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > .
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to elasticsearch+unsubscribe@\*\*[googlegroups.com](http://googlegroups.com)[elasticsearch%2Bunsubscribe@googlegroups.com](mailto:elasticsearch%2Bunsubscribe@googlegroups.com)  
> .  
> For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> .

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 19, 2013, 11:53am UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/5 "2013-03-19T11:53:31Z")

</div>

On Mon, 2013-03-18 at 21:32 -0700, Bruno Miranda wrote:

> While indexing to our QA environment, 2 nodes (EC2 m1.large) each with  
> 2 CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
> [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)
> 
> I have about 20 processes indexing simultaneously. The index is 3.3  
> million documents and about 8GB on disk. The refresh internal is set  
> at 5 minutes while the reindex is going, the index is 5 shards and 1  
> replica. I haven't changed the merge policy.

Look at using merge throttling. Merges can be heavy and use a lot of IO.  
With throttling, merges will still happen, but won't swamp your system

clint

> 

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [March 19, 2013, 7:01pm UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/6 "2013-03-19T19:01:14Z")

</div>

Hi,

I see you have "SPM" bookmarked in your browser. You should look at graphs  
under the "Index Stats" tab -- these:  
[https://apps.sematext.com/spm-reports/mainPage.do#report\_anchor\_esRefreshFlushMerge](https://apps.sematext.com/spm-reports/mainPage.do#report_anchor_esRefreshFlushMerge)  
to see what's going on with ES/Lucene refreshing, flushing, and merging as  
you make changes to throttle merges that others have suggested.

## Otis

ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Tuesday, March 19, 2013 12:32:02 AM UTC-4, Bruno Miranda wrote:

> While indexing to our QA environment, 2 nodes (EC2 m1.large) each with 2  
> CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
> [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)
> 
> I have about 20 processes indexing simultaneously. The index is 3.3  
> million documents and about 8GB on disk. The refresh internal is set at 5  
> minutes while the reindex is going, the index is 5 shards and 1 replica. I  
> haven't changed the merge policy.
> 
> Cutting the concurrent process to 20 to 10 definitely lowers the IO from  
> 98% to around 50%.
> 
> Any tips on lowering disk IO? Is this normal? Reindex seems to finish  
> properly, all records are properly indexed but the 100% disk IO scares me.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Bruno\_Miranda](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_miranda/32/746_2.png) [@Bruno\_Miranda](https://discuss.elastic.co/u/Bruno_Miranda)\
**Post date:** [March 19, 2013, 7:03pm UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/7 "2013-03-19T19:03:56Z")

</div>

Need to convince the CEO to pay for the plan. I can only see 30 minutes.  
(Which should help me diagnose)

Thank you for the reminder.

On Tuesday, March 19, 2013 12:01:14 PM UTC-7, Otis Gospodnetic wrote:

> Hi,
> 
> I see you have "SPM" bookmarked in your browser. You should look at  
> graphs under the "Index Stats" tab -- these:  
> [https://apps.sematext.com/spm-reports/mainPage.do#report\_anchor\_esRefreshFlushMergeto](https://apps.sematext.com/spm-reports/mainPage.do#report_anchor_esRefreshFlushMergeto) see what's going on with ES/Lucene refreshing, flushing, and merging as  
> you make changes to throttle merges that others have suggested.
> 
> ## Otis
> 
> ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Tuesday, March 19, 2013 12:32:02 AM UTC-4, Bruno Miranda wrote:
> 
> > While indexing to our QA environment, 2 nodes (EC2 m1.large) each with 2  
> > CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
> > [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)
> > 
> > I have about 20 processes indexing simultaneously. The index is 3.3  
> > million documents and about 8GB on disk. The refresh internal is set at 5  
> > minutes while the reindex is going, the index is 5 shards and 1 replica. I  
> > haven't changed the merge policy.
> > 
> > Cutting the concurrent process to 20 to 10 definitely lowers the IO from  
> > 98% to around 50%.
> > 
> > Any tips on lowering disk IO? Is this normal? Reindex seems to finish  
> > properly, all records are properly indexed but the 100% disk IO scares me.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [March 20, 2013, 8:35pm UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/8 "2013-03-20T20:35:16Z")

</div>

which version of ES are you using?

simon

On Tuesday, March 19, 2013 8:03:56 PM UTC+1, Bruno Miranda wrote:

> Need to convince the CEO to pay for the plan. I can only see 30 minutes.  
> (Which should help me diagnose)
> 
> Thank you for the reminder.
> 
> On Tuesday, March 19, 2013 12:01:14 PM UTC-7, Otis Gospodnetic wrote:
> 
> > Hi,
> > 
> > I see you have "SPM" bookmarked in your browser. You should look at  
> > graphs under the "Index Stats" tab -- these:  
> > [https://apps.sematext.com/spm-reports/mainPage.do#report\_anchor\_esRefreshFlushMergeto](https://apps.sematext.com/spm-reports/mainPage.do#report_anchor_esRefreshFlushMergeto) see what's going on with ES/Lucene refreshing, flushing, and merging as  
> > you make changes to throttle merges that others have suggested.
> > 
> > ## Otis
> > 
> > ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > 
> > On Tuesday, March 19, 2013 12:32:02 AM UTC-4, Bruno Miranda wrote:
> > 
> > > While indexing to our QA environment, 2 nodes (EC2 m1.large) each with 2  
> > > CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
> > > [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)
> > > 
> > > I have about 20 processes indexing simultaneously. The index is 3.3  
> > > million documents and about 8GB on disk. The refresh internal is set at 5  
> > > minutes while the reindex is going, the index is 5 shards and 1 replica. I  
> > > haven't changed the merge policy.
> > > 
> > > Cutting the concurrent process to 20 to 10 definitely lowers the IO from  
> > > 98% to around 50%.
> > > 
> > > Any tips on lowering disk IO? Is this normal? Reindex seems to finish  
> > > properly, all records are properly indexed but the 100% disk IO scares me.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Bruno\_Miranda](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_miranda/32/746_2.png) [@Bruno\_Miranda](https://discuss.elastic.co/u/Bruno_Miranda)\
**Post date:** [March 20, 2013, 9:43pm UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/9 "2013-03-20T21:43:33Z")

</div>

0.20.5

On Wednesday, March 20, 2013 1:35:16 PM UTC-7, simonw wrote:

> which version of ES are you using?
> 
> simon
> 
> On Tuesday, March 19, 2013 8:03:56 PM UTC+1, Bruno Miranda wrote:
> 
> > Need to convince the CEO to pay for the plan. I can only see 30 minutes.  
> > (Which should help me diagnose)
> > 
> > Thank you for the reminder.
> > 
> > On Tuesday, March 19, 2013 12:01:14 PM UTC-7, Otis Gospodnetic wrote:
> > 
> > > Hi,
> > > 
> > > I see you have "SPM" bookmarked in your browser. You should look at  
> > > graphs under the "Index Stats" tab -- these:  
> > > [https://apps.sematext.com/spm-reports/mainPage.do#report\_anchor\_esRefreshFlushMergeto](https://apps.sematext.com/spm-reports/mainPage.do#report_anchor_esRefreshFlushMergeto) see what's going on with ES/Lucene refreshing, flushing, and merging as  
> > > you make changes to throttle merges that others have suggested.
> > > 
> > > ## Otis
> > > 
> > > ELASTICSEARCH Performance Monitoring -  
> > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > 
> > > On Tuesday, March 19, 2013 12:32:02 AM UTC-4, Bruno Miranda wrote:
> > > 
> > > > While indexing to our QA environment, 2 nodes (EC2 m1.large) each with  
> > > > 2 CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
> > > > [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)
> > > > 
> > > > I have about 20 processes indexing simultaneously. The index is 3.3  
> > > > million documents and about 8GB on disk. The refresh internal is set at 5  
> > > > minutes while the reindex is going, the index is 5 shards and 1 replica. I  
> > > > haven't changed the merge policy.
> > > > 
> > > > Cutting the concurrent process to 20 to 10 definitely lowers the IO  
> > > > from 98% to around 50%.
> > > > 
> > > > Any tips on lowering disk IO? Is this normal? Reindex seems to finish  
> > > > properly, all records are properly indexed but the 100% disk IO scares me.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Bruno\_Miranda](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruno_miranda/32/746_2.png) [@Bruno\_Miranda](https://discuss.elastic.co/u/Bruno_Miranda)\
**Post date:** [March 20, 2013, 9:44pm UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/10 "2013-03-20T21:44:09Z")

</div>

I seem to have lowered the IO by limiting the max\_bytes\_per\_sec to 10mb, I  
guess there was no limit before by default.

On Wednesday, March 20, 2013 2:43:33 PM UTC-7, Bruno Miranda wrote:

> 0.20.5
> 
> On Wednesday, March 20, 2013 1:35:16 PM UTC-7, simonw wrote:
> 
> > which version of ES are you using?
> > 
> > simon
> > 
> > On Tuesday, March 19, 2013 8:03:56 PM UTC+1, Bruno Miranda wrote:
> > 
> > > Need to convince the CEO to pay for the plan. I can only see 30 minutes.  
> > > (Which should help me diagnose)
> > > 
> > > Thank you for the reminder.
> > > 
> > > On Tuesday, March 19, 2013 12:01:14 PM UTC-7, Otis Gospodnetic wrote:
> > > 
> > > > Hi,
> > > > 
> > > > I see you have "SPM" bookmarked in your browser. You should look at  
> > > > graphs under the "Index Stats" tab -- these:  
> > > > [https://apps.sematext.com/spm-reports/mainPage.do#report\_anchor\_esRefreshFlushMergeto](https://apps.sematext.com/spm-reports/mainPage.do#report_anchor_esRefreshFlushMergeto) see what's going on with ES/Lucene refreshing, flushing, and merging as  
> > > > you make changes to throttle merges that others have suggested.
> > > > 
> > > > ## Otis
> > > > 
> > > > ELASTICSEARCH Performance Monitoring -  
> > > > [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > > > 
> > > > On Tuesday, March 19, 2013 12:32:02 AM UTC-4, Bruno Miranda wrote:
> > > > 
> > > > > While indexing to our QA environment, 2 nodes (EC2 m1.large) each with  
> > > > > 2 CPUs 7.3 GBs RAM I am seeing exceptionally high Disk IO.  
> > > > > [http://cl.ly/image/371H01382h2O](http://cl.ly/image/371H01382h2O)
> > > > > 
> > > > > I have about 20 processes indexing simultaneously. The index is 3.3  
> > > > > million documents and about 8GB on disk. The refresh internal is set at 5  
> > > > > minutes while the reindex is going, the index is 5 shards and 1 replica. I  
> > > > > haven't changed the merge policy.
> > > > > 
> > > > > Cutting the concurrent process to 20 to 10 definitely lowers the IO  
> > > > > from 98% to around 50%.
> > > > > 
> > > > > Any tips on lowering disk IO? Is this normal? Reindex seems to finish  
> > > > > properly, all records are properly indexed but the 100% disk IO scares me.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:45am UTC](https://discuss.elastic.co/t/very-high-disk-io-while-indexing/11199/11 "2017-07-06T02:45:31Z")

</div>


