# Slow Indexing Speed

**URL:** <https://discuss.elastic.co/t/slow-indexing-speed/13484>\
**Category:** Elasticsearch\
**Created:** [September 5, 2013, 11:39pm UTC](https://discuss.elastic.co/t/slow-indexing-speed/13484 "2013-09-05T23:39:13Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Robert\_Navarro](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robert_navarro/32/2125_2.png) [@Robert\_Navarro](https://discuss.elastic.co/u/Robert_Navarro)\
**Post date:** [September 5, 2013, 11:39pm UTC](https://discuss.elastic.co/t/slow-indexing-speed/13484/1 "2013-09-05T23:39:13Z")

</div>

Hello,

I have a single server ES "cluster" setup right now and it's struggling to  
keep up with our indexing load.

The server has 30GB of ram, 15GB locked for elasticsearch.

Here are the ES details:

{  
"ok" : true,  
"status" : 200,  
"name" : "esls1",  
"version" : {  
"number" : "0.90.3",  
"build\_hash" : "5c38d6076448b899d758f29443329571e2522410",  
"build\_timestamp" : "2013-08-06T13:18:31Z",  
"build\_snapshot" : false,  
"lucene\_version" : "4.4"  
},  
"tagline" : "You Know, for Search"  
}

Here is the java version:

root@esls1:~# java -version  
java version "1.7.0\_25"  
Java(TM) SE Runtime Environment (build 1.7.0\_25-b15)  
Java HotSpot(TM) 64-Bit Server VM (build 23.25-b01, mixed mode)

Here are some of the knobs I've tried to tweak for our indexes...this is  
just a snapshot of one index:

> <https://gist.github.com/rnavarro/490196cf73ff46e33a8b>

Operating System:

Ubuntu 12.04.2 LTS

There are 15 indexes on this node, rotated daily and removed after 14 days.

The incoming index requests are coming out of logstash and it's all logging data.

I suspect the server is IO bound as there is are bursts of 5-10s sustained 10%+ iowait.

What other knobs can I tweak to help speed things along?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [September 6, 2013, 12:23am UTC](https://discuss.elastic.co/t/slow-indexing-speed/13484/2 "2013-09-06T00:23:20Z")

</div>

What's your indexing load (docs/sec) ? Are you querying at the same time?  
Often, if you are bound by Disk IO, there isn't much you can do except get  
faster disks or more nodes. Do you have SSDs? They are a great investment  
if you can afford them. And adding more nodes is almost a linear increase  
in indexing speed.

Some more things you can do:

- If you don't need it, disable the \_all field. This bloats the doc  
size (so more bytes to write) and eats up a bit of CPU.
- I'd put the index.merge.policy.segments\_per\_tier back to it's default  
(10). By having it set so high, Lucene is going to perform big bursts of  
merging which can easily eat up all your IO and a considerable amount of  
CPU. In general, I've spent a lot of time fiddling with the merge policy  
settings and never found a configuration better than the defaults. Mike  
McCandless[http://blog.mikemccandless.com/2011/02/visualizing-lucenes-segment-merges.html](http://blog.mikemccandless.com/2011/02/visualizing-lucenes-segment-merges.html)knows best 🙂
- I noticed you have Term Vectors compressed. Are you actually using  
term vectors? They double your index size and eat up more IO.
- You could check the indexing thread count and see if you are routinely  
queuing indexing threads. May help to increase that some (although, it may  
not)

-Zach

On Thursday, September 5, 2013 7:39:13 PM UTC-4, Robert Navarro wrote:

> Hello,
> 
> I have a single server ES "cluster" setup right now and it's struggling to  
> keep up with our indexing load.
> 
> The server has 30GB of ram, 15GB locked for elasticsearch.
> 
> Here are the ES details:
> 
> {  
> "ok" : true,  
> "status" : 200,  
> "name" : "esls1",  
> "version" : {  
> "number" : "0.90.3",  
> "build\_hash" : "5c38d6076448b899d758f29443329571e2522410",  
> "build\_timestamp" : "2013-08-06T13:18:31Z",  
> "build\_snapshot" : false,  
> "lucene\_version" : "4.4"  
> },  
> "tagline" : "You Know, for Search"  
> }
> 
> Here is the java version:
> 
> root@esls1:~# java -version  
> java version "1.7.0\_25"  
> Java(TM) SE Runtime Environment (build 1.7.0\_25-b15)  
> Java HotSpot(TM) 64-Bit Server VM (build 23.25-b01, mixed mode)
> 
> Here are some of the knobs I've tried to tweak for our indexes...this is  
> just a snapshot of one index:
> 
> [gist:490196cf73ff46e33a8b · GitHub](https://gist.github.com/rnavarro/490196cf73ff46e33a8b)
> 
> Operating System:
> 
> Ubuntu 12.04.2 LTS
> 
> There are 15 indexes on this node, rotated daily and removed after 14 days.
> 
> The incoming index requests are coming out of logstash and it's all logging data.
> 
> I suspect the server is IO bound as there is are bursts of 5-10s sustained 10%+ iowait.
> 
> What other knobs can I tweak to help speed things along?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Robert\_Navarro](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robert_navarro/32/2125_2.png) [@Robert\_Navarro](https://discuss.elastic.co/u/Robert_Navarro)\
**Post date:** [September 6, 2013, 12:41am UTC](https://discuss.elastic.co/t/slow-indexing-speed/13484/3 "2013-09-06T00:41:17Z")

</div>

Hey Zach,

Thanks for the response!

The indexing load isn't particularly high, but the documents being indexed  
are pretty large....many of them \>1MB. Looking at my logstash indexing  
machines I'd say that the docs/sec is in the realm of 200 ish? Is there a  
way to see this from the ES side of things?

We do query at the same time, but we generally query indexes that are a  
day+ old and have been optimized by a nightly cron....less querying happens  
on the "active" daily index.

I don't have SSDs right now, this is running on a rackspace cloud server  
with 5 non-ssd block storage volumes attached in a raid5 config. We index  
~300GB of data a day...so moving to SSDs is likely very very cost  
prohibitive for us.

We have \_all disabled in our index template, so that should help things.

I'll drop the index.merge.policy.segments\_per\_tier back down to 10 to see  
how that effects things.

We're using kibana to do the searching on our indexes and from what I  
understand (and looking at some of the requests) there are no term vectors  
being used.

I'd be more than happy to add nodes to the cluster, but I wasn't certain if  
that would help indexing speed as much as it would query speed. However,  
now that you mention it...if the shards were all split up one per node that  
would make sense that I would have gains there.

Also, I forgot to add a copy of our indexing template...so here it is  
(updated to reflect the default merge policy settings):

> <https://gist.github.com/rnavarro/5abf0f74983eabb14071>

Thanks for your time and all the food for thought! Much appreciated! 🙂

On Thursday, September 5, 2013 5:23:20 PM UTC-7, Zachary Tong wrote:

> What's your indexing load (docs/sec) ? Are you querying at the same time?  
> Often, if you are bound by Disk IO, there isn't much you can do except get  
> faster disks or more nodes. Do you have SSDs? They are a great investment  
> if you can afford them. And adding more nodes is almost a linear increase  
> in indexing speed.
> 
> Some more things you can do:
> 
> - If you don't need it, disable the \_all field. This bloats the doc  
> size (so more bytes to write) and eats up a bit of CPU.
> - I'd put the index.merge.policy.segments\_per\_tier back to it's  
> default (10). By having it set so high, Lucene is going to perform big  
> bursts of merging which can easily eat up all your IO and a considerable  
> amount of CPU. In general, I've spent a lot of time fiddling with the  
> merge policy settings and never found a configuration better than the  
> defaults. Mike McCandless[http://blog.mikemccandless.com/2011/02/visualizing-lucenes-segment-merges.html](http://blog.mikemccandless.com/2011/02/visualizing-lucenes-segment-merges.html)knows best 🙂
> - I noticed you have Term Vectors compressed. Are you actually using  
> term vectors? They double your index size and eat up more IO.
> - You could check the indexing thread count and see if you are  
> routinely queuing indexing threads. May help to increase that some  
> (although, it may not)
> 
> -Zach
> 
> On Thursday, September 5, 2013 7:39:13 PM UTC-4, Robert Navarro wrote:
> 
> > Hello,
> > 
> > I have a single server ES "cluster" setup right now and it's struggling  
> > to keep up with our indexing load.
> > 
> > The server has 30GB of ram, 15GB locked for elasticsearch.
> > 
> > Here are the ES details:
> > 
> > {  
> > "ok" : true,  
> > "status" : 200,  
> > "name" : "esls1",  
> > "version" : {  
> > "number" : "0.90.3",  
> > "build\_hash" : "5c38d6076448b899d758f29443329571e2522410",  
> > "build\_timestamp" : "2013-08-06T13:18:31Z",  
> > "build\_snapshot" : false,  
> > "lucene\_version" : "4.4"  
> > },  
> > "tagline" : "You Know, for Search"  
> > }
> > 
> > Here is the java version:
> > 
> > root@esls1:~# java -version  
> > java version "1.7.0\_25"  
> > Java(TM) SE Runtime Environment (build 1.7.0\_25-b15)  
> > Java HotSpot(TM) 64-Bit Server VM (build 23.25-b01, mixed mode)
> > 
> > Here are some of the knobs I've tried to tweak for our indexes...this is  
> > just a snapshot of one index:
> > 
> > [gist:490196cf73ff46e33a8b · GitHub](https://gist.github.com/rnavarro/490196cf73ff46e33a8b)
> > 
> > Operating System:
> > 
> > Ubuntu 12.04.2 LTS
> > 
> > There are 15 indexes on this node, rotated daily and removed after 14 days.
> > 
> > The incoming index requests are coming out of logstash and it's all logging data.
> > 
> > I suspect the server is IO bound as there is are bursts of 5-10s sustained 10%+ iowait.
> > 
> > What other knobs can I tweak to help speed things along?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Israel\_Ekpo](https://avatars.discourse-cdn.com/v4/letter/i/49beb7/32.png) [@Israel\_Ekpo](https://discuss.elastic.co/u/Israel_Ekpo)\
**Post date:** [September 6, 2013, 2:47am UTC](https://discuss.elastic.co/t/slow-indexing-speed/13484/4 "2013-09-06T02:47:44Z")

</div>

Since you are not actively searching on the indices that are currently  
being indexed I would recommend for you to increase the refresh interval  
for the index.

Checkout some benchmarks :

> **[Elasticsearch Refresh Interval vs Indexing Performance - Sematext](https://sematext.com/blog/elasticsearch-refresh-interval-vs-indexing-performance/)**
>
> Elasticsearch is near-realtime, in the sense that when you index a document, you need to wait for the next refresh for that document to appear in a search. Refreshing is an expensive operation and that is why by default it’s made at a regular...

_Author and Instructor for the Upcoming Book and Lecture Series_  
_Massive Log Data Aggregation, Processing, Searching and Visualization with  
Open Source Software_  
_[http://massivelogdata.com](http://massivelogdata.com)_

On Thu, Sep 5, 2013 at 7:39 PM, Robert Navarro [crshman@gmail.com](mailto:crshman@gmail.com) wrote:

> Hello,
> 
> I have a single server ES "cluster" setup right now and it's struggling to  
> keep up with our indexing load.
> 
> The server has 30GB of ram, 15GB locked for elasticsearch.
> 
> Here are the ES details:
> 
> {  
> "ok" : true,  
> "status" : 200,  
> "name" : "esls1",  
> "version" : {  
> "number" : "0.90.3",  
> "build\_hash" : "5c38d6076448b899d758f29443329571e2522410",  
> "build\_timestamp" : "2013-08-06T13:18:31Z",  
> "build\_snapshot" : false,  
> "lucene\_version" : "4.4"  
> },  
> "tagline" : "You Know, for Search"  
> }
> 
> Here is the java version:
> 
> root@esls1:~# java -version  
> java version "1.7.0\_25"  
> Java(TM) SE Runtime Environment (build 1.7.0\_25-b15)  
> Java HotSpot(TM) 64-Bit Server VM (build 23.25-b01, mixed mode)
> 
> Here are some of the knobs I've tried to tweak for our indexes...this is  
> just a snapshot of one index:
> 
> [gist:490196cf73ff46e33a8b · GitHub](https://gist.github.com/rnavarro/490196cf73ff46e33a8b)
> 
> Operating System:
> 
> Ubuntu 12.04.2 LTS
> 
> There are 15 indexes on this node, rotated daily and removed after 14 days.
> 
> The incoming index requests are coming out of logstash and it's all logging data.
> 
> I suspect the server is IO bound as there is are bursts of 5-10s sustained 10%+ iowait.
> 
> What other knobs can I tweak to help speed things along?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![polyfractal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/polyfractal/32/48162_2.png) [@polyfractal](https://discuss.elastic.co/u/polyfractal)\
**Post date:** [September 6, 2013, 7:58pm UTC](https://discuss.elastic.co/t/slow-indexing-speed/13484/5 "2013-09-06T19:58:25Z")

</div>

> The indexing load isn't particularly high, but the documents being indexed  
> are pretty large....many of them \>1MB. Looking at my logstash indexing  
> machines I'd say that the docs/sec is in the realm of 200 ish? Is there a  
> way to see this from the ES side of things?

You can use the Indices stats API[http://www.elasticsearch.org/guide/reference/api/admin-indices-stats/](http://www.elasticsearch.org/guide/reference/api/admin-indices-stats/)to see the indexing rate. You'll get an output that includes indexing  
stats (from the viewpoint of the index, regardless of which shard/machine  
it goes towards):

curl -XGET 'localhost:9200/\_stats'

[...]  
"indexing": {  
"index\_total": 3,  
"index\_time\_in\_millis": 49,  
"index\_current": 0,  
"delete\_total": 0,  
"delete\_time\_in\_millis": 0,  
"delete\_current": 0  
},  
[...]

Another place to look is the Indexing threadpool via the Cluster stats API[http://www.elasticsearch.org/guide/reference/api/admin-cluster-nodes-stats/](http://www.elasticsearch.org/guide/reference/api/admin-cluster-nodes-stats/)- check to see if you are queuing a lot of threads:

curl -XGET '[http://localhost:9200/\_nodes/stats?clear=true&thread\_pool=true](http://localhost:9200/_nodes/stats?clear=true&thread_pool=true)'

[...]  
"index": {  
"threads": 3,  
"queue": 0,  
"active": 0,  
"rejected": 0,  
"largest": 3,  
"completed": 3  
},  
[...]

Plugins like Bigdesk [https://github.com/lukas-vlcek/bigdesk/](https://github.com/lukas-vlcek/bigdesk/) and Paramedic[https://github.com/karmi/elasticsearch-paramedic](https://github.com/karmi/elasticsearch-paramedic)are basically graphical wrappers for these APIs.

I'd be more than happy to add nodes to the cluster, but I wasn't certain if

> that would help indexing speed as much as it would query speed. However,  
> now that you mention it...if the shards were all split up one per node that  
> would make sense that I would have gains there.

Yep, you'll definitely see an increase as each node adds more indexing  
throughput. It isn't exactly linear, but it is fairly close (especially if  
your query load is low, as in most logging environments). Realistically,  
this is the easiest and fastest way to increase your indexing speed if you  
can afford the cost of another node

Hope this helps! Keep us updated if you have any more questions  
-Zach

> On Thursday, September 5, 2013 5:23:20 PM UTC-7, Zachary Tong wrote:
> 
> > What's your indexing load (docs/sec) ? Are you querying at the same  
> > time? Often, if you are bound by Disk IO, there isn't much you can do  
> > except get faster disks or more nodes. Do you have SSDs? They are a great  
> > investment if you can afford them. And adding more nodes is almost a  
> > linear increase in indexing speed.
> > 
> > Some more things you can do:
> > 
> > - If you don't need it, disable the \_all field. This bloats the doc  
> > size (so more bytes to write) and eats up a bit of CPU.
> > - I'd put the index.merge.policy.segments\_per\_tier back to it's  
> > default (10). By having it set so high, Lucene is going to perform big  
> > bursts of merging which can easily eat up all your IO and a considerable  
> > amount of CPU. In general, I've spent a lot of time fiddling with the  
> > merge policy settings and never found a configuration better than the  
> > defaults. Mike McCandless[http://blog.mikemccandless.com/2011/02/visualizing-lucenes-segment-merges.html](http://blog.mikemccandless.com/2011/02/visualizing-lucenes-segment-merges.html)knows best 🙂
> > - I noticed you have Term Vectors compressed. Are you actually using  
> > term vectors? They double your index size and eat up more IO.
> > - You could check the indexing thread count and see if you are  
> > routinely queuing indexing threads. May help to increase that some  
> > (although, it may not)
> > 
> > -Zach
> > 
> > On Thursday, September 5, 2013 7:39:13 PM UTC-4, Robert Navarro wrote:
> > 
> > > Hello,
> > > 
> > > I have a single server ES "cluster" setup right now and it's struggling  
> > > to keep up with our indexing load.
> > > 
> > > The server has 30GB of ram, 15GB locked for elasticsearch.
> > > 
> > > Here are the ES details:
> > > 
> > > {  
> > > "ok" : true,  
> > > "status" : 200,  
> > > "name" : "esls1",  
> > > "version" : {  
> > > "number" : "0.90.3",  
> > > "build\_hash" : "5c38d6076448b899d758f29443329571e2522410",  
> > > "build\_timestamp" : "2013-08-06T13:18:31Z",  
> > > "build\_snapshot" : false,  
> > > "lucene\_version" : "4.4"  
> > > },  
> > > "tagline" : "You Know, for Search"  
> > > }
> > > 
> > > Here is the java version:
> > > 
> > > root@esls1:~# java -version  
> > > java version "1.7.0\_25"  
> > > Java(TM) SE Runtime Environment (build 1.7.0\_25-b15)  
> > > Java HotSpot(TM) 64-Bit Server VM (build 23.25-b01, mixed mode)
> > > 
> > > Here are some of the knobs I've tried to tweak for our indexes...this is  
> > > just a snapshot of one index:
> > > 
> > > [gist:490196cf73ff46e33a8b · GitHub](https://gist.github.com/rnavarro/490196cf73ff46e33a8b)
> > > 
> > > Operating System:
> > > 
> > > Ubuntu 12.04.2 LTS
> > > 
> > > There are 15 indexes on this node, rotated daily and removed after 14 days.
> > > 
> > > The incoming index requests are coming out of logstash and it's all logging data.
> > > 
> > > I suspect the server is IO bound as there is are bursts of 5-10s sustained 10%+ iowait.
> > > 
> > > What other knobs can I tweak to help speed things along?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:17am UTC](https://discuss.elastic.co/t/slow-indexing-speed/13484/6 "2017-07-06T02:17:43Z")

</div>


