# Improve indexing throughput

**URL:** <https://discuss.elastic.co/t/improve-indexing-throughput/6877>\
**Category:** Elasticsearch\
**Created:** [March 4, 2012, 9:05am UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877 "2012-03-04T09:05:58Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lior\_Cohen\_barnea](https://avatars.discourse-cdn.com/v4/letter/l/dec6dc/32.png) [@Lior\_Cohen\_barnea](https://discuss.elastic.co/u/Lior_Cohen_barnea)\
**Post date:** [March 4, 2012, 9:05am UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/1 "2012-03-04T09:05:58Z")

</div>

Hi all,

I am doing a POC on ES and Solr cloud.

I want to improve indexing throughput which is quite less than Solr's.  
What approach should i take:  
Increase number of shards and/or nodes?  
Config my client with all nodes IP's?  
Run 1 thread per node and configure 1 node IP per client?  
Changes the segments settings ? if so, how?

Thanks in advance,  
Lior.

---

<div class="post-metadata">

**Author:** ![Mark\_Waddle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_waddle/32/2608_2.png) [@Mark\_Waddle](https://discuss.elastic.co/u/Mark_Waddle)\
**Post date:** [March 5, 2012, 6:57am UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/2 "2012-03-05T06:57:48Z")

</div>

Interesting. My experience so far has been that ES has equivalent or faster  
indexing than Solr. Are you doing bulk indexing or incremental indexing? If  
you are doing bulk indexing you should read

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

.

On Sunday, March 4, 2012 1:05:58 AM UTC-8, Lior Cohen barnea wrote:

> Hi all,
> 
> I am doing a POC on ES and Solr cloud.
> 
> I want to improve indexing throughput which is quite less than Solr's.  
> What approach should i take:  
> Increase number of shards and/or nodes?  
> Config my client with all nodes IP's?  
> Run 1 thread per node and configure 1 node IP per client?  
> Changes the segments settings ? if so, how?
> 
> Thanks in advance,  
> Lior.

---

<div class="post-metadata">

**Author:** ![Radu\_Gheorghe1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe1/32/2688_2.png) [@Radu\_Gheorghe1](https://discuss.elastic.co/u/Radu_Gheorghe1)\
**Post date:** [March 5, 2012, 8:36am UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/3 "2012-03-05T08:36:24Z")

</div>

Hi,

I just had a "facepalm moment" while doing some performance tests on  
cloud machines. Maybe it also applies to you:

the default memory settings for Elasticsearch are:

# grep \_MEM elasticsearch/bin/elasticsearch.in.sh

if ["x$ES\_MIN\_MEM" = "x"]; then  
ES\_MIN\_MEM=256m  
if ["x$ES\_MAX\_MEM" = "x"]; then  
ES\_MAX\_MEM=1g  
JAVA\_OPTS="$JAVA\_OPTS -Xms${ES\_MIN\_MEM}"  
JAVA\_OPTS="$JAVA\_OPTS -Xmx${ES\_MAX\_MEM}"

This might turn out to be too little when inserting a huge amount of  
logs for performance tests.

When I raised this limit (eg: 4g min, 20g max), the inserting  
performance suddenly multiplied by 3.

Take a look here for some more info:

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

For performance, it's best to set the minimum and maximum memory to  
the same value, in order to avoid using the Garbage Collector as much  
as possible.

On Mar 4, 11:05 am, Lior Cohen barnea [liorcohenbar...@gmail.com](mailto:liorcohenbar...@gmail.com)  
wrote:

> Hi all,
> 
> I am doing a POC on ES and Solr cloud.
> 
> I want to improve indexing throughput which is quite less than Solr's.  
> What approach should i take:  
> Increase number of shards and/or nodes?  
> Config my client with all nodes IP's?  
> Run 1 thread per node and configure 1 node IP per client?  
> Changes the segments settings ? if so, how?
> 
> Thanks in advance,  
> Lior.

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [March 5, 2012, 11:58am UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/4 "2012-03-05T11:58:37Z")

</div>

Hello Lior,

Well, Elasticsearch has real-time replication, so ES has to do more work.  
Also, ES has index refreshing happening every 1 sec by default, while Solr  
doesn't (or are you comparing ES and SolrCloud?).  
...  
I think more shards would speed things up, as would bulk API and setting  
the number of replicas to 0.  
You could also increase the number of concurrent indexers in order to  
maximize CPU and/or disk IO.

## Otis

Sematext is hiring Elasticsearch engineers World-Wide --

> **[Jobs](https://sematext.com/jobs/)**
>
> We’re Hiring We are always looking for smart, passionate, motivated, and independent people regardless of where on the planet they may be. Learn more about the company Agent & Backend Engineer Full Stack Developer Backend Engineer Frontend...

On Sunday, March 4, 2012 5:05:58 PM UTC+8, Lior Cohen barnea wrote:

> Hi all,
> 
> I am doing a POC on ES and Solr cloud.
> 
> I want to improve indexing throughput which is quite less than Solr's.  
> What approach should i take:  
> Increase number of shards and/or nodes?  
> Config my client with all nodes IP's?  
> Run 1 thread per node and configure 1 node IP per client?  
> Changes the segments settings ? if so, how?
> 
> Thanks in advance,  
> Lior.

On Sunday, March 4, 2012 5:05:58 PM UTC+8, Lior Cohen barnea wrote:

> Hi all,
> 
> I am doing a POC on ES and Solr cloud.
> 
> I want to improve indexing throughput which is quite less than Solr's.  
> What approach should i take:  
> Increase number of shards and/or nodes?  
> Config my client with all nodes IP's?  
> Run 1 thread per node and configure 1 node IP per client?  
> Changes the segments settings ? if so, how?
> 
> Thanks in advance,  
> Lior.

---

<div class="post-metadata">

**Author:** ![Thomas\_Peuss](https://avatars.discourse-cdn.com/v4/letter/t/22d042/32.png) [@Thomas\_Peuss](https://discuss.elastic.co/u/Thomas_Peuss)\
**Post date:** [March 5, 2012, 2:00pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/5 "2012-03-05T14:00:42Z")

</div>

Hi Lior!

Am Sonntag, 4. März 2012 10:05:58 UTC+1 schrieb Lior Cohen barnea:

> I want to improve indexing throughput which is quite less than Solr's.  
> What approach should i take:  
> Increase number of shards and/or nodes?  
> Config my client with all nodes IP's?  
> Run 1 thread per node and configure 1 node IP per client?  
> Changes the segments settings ? if so, how?

First of all make sure that you don't compare apples with oranges. When you  
calculate the indexing speed of Solr you need to take the "commit"-time  
into account you don't need with ES. And you will see that Solr gets  
"slower".

Increasing the shard count is only useful if you have enough CPU cores and  
disk spindles available. You should increase the refresh\_interval or  
disable it alltogether and trigger it manually if you get to really big  
index sizes (we have \>1 billion docs in some indices while having \>500  
indices). If you run on many nodes your network can be the next

Bulk indexing gives you a big performance gain over doing single updates so  
you should consider that.

But in the end you need to benchmark with your amount of data (docs per  
second, index size etc.) and different settings. If need to guarantee some  
sort of SLA there is no other way around it.

CU  
Thomas

---

<div class="post-metadata">

**Author:** ![haarts](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/haarts/32/2972_2.png) [@haarts](https://discuss.elastic.co/u/haarts)\
**Post date:** [March 5, 2012, 2:21pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/6 "2012-03-05T14:21:11Z")

</div>

Hi Thomas,

Those are some impressive numbers. Would you mind sharing on what kind of  
machines you are running? We are struggling indexing 500M documents,  
reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1  
simple spinner). Performance indexing is acceptable. But first time query  
performance isn't great (seconds...).

On Monday, 5 March 2012 15:00:42 UTC+1, Thomas Peuss wrote:

> Hi Lior!
> 
> Am Sonntag, 4. März 2012 10:05:58 UTC+1 schrieb Lior Cohen barnea:
> 
> > I want to improve indexing throughput which is quite less than Solr's.  
> > What approach should i take:  
> > Increase number of shards and/or nodes?  
> > Config my client with all nodes IP's?  
> > Run 1 thread per node and configure 1 node IP per client?  
> > Changes the segments settings ? if so, how?
> 
> First of all make sure that you don't compare apples with oranges. When  
> you calculate the indexing speed of Solr you need to take the "commit"-time  
> into account you don't need with ES. And you will see that Solr gets  
> "slower".
> 
> Increasing the shard count is only useful if you have enough CPU cores and  
> disk spindles available. You should increase the refresh\_interval or  
> disable it alltogether and trigger it manually if you get to really big  
> index sizes (we have \>1 billion docs in some indices while having \>500  
> indices). If you run on many nodes your network can be the next
> 
> Bulk indexing gives you a big performance gain over doing single updates  
> so you should consider that.
> 
> But in the end you need to benchmark with your amount of data (docs per  
> second, index size etc.) and different settings. If need to guarantee some  
> sort of SLA there is no other way around it.
> 
> CU  
> Thomas

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [March 5, 2012, 2:54pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/7 "2012-03-05T14:54:56Z")

</div>

Hi,

here are some techniques I used for Elasticsearch indexing:

- parallel multithreading custom remote index client (12 instances in  
parallel, depends on your input data streams)
- remote index client for distributing feed load and ES node load. ES  
TransportClient is threadsafe, a single instance can be re-used for  
parallel indexing.
- sharding (12 shards / primaries plus 1x replica = 24 Lucene indexes  
over 3 nodes in total)
- bulk indexing, balanced by waiting for call backs, limited to max.  
30 jobs with 100 docs each, each doc is ~2-200 KB, 4KB average
- compressed input (reduces I/O bandwidth with large text documents)
- RHEL, ES data volume mounted with noatime option, SAS-2 (6Gbit/s)  
drives, RAID 5
- default ES config changes, no change of refresh\_interval, merge  
policies etc. = no ES tuning at all

This config can index constantly ~3000 docs per second for ~2 hours to  
Elasticsearch (for a total of ~20 mio. documents). Each ~20 min on  
each node, load spikes happen, because Lucene merges take heavy I/O.  
This could be optimized but for my purposes it suffices right now.

I am afraid comparing an ES cluster setup to a single Solr index seems  
quite unfair.

Jörg

On Mar 5, 12:58 pm, Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)  
wrote:

> Hello Lior,
> 
> Well, Elasticsearch has real-time replication, so ES has to do more work.  
> Also, ES has index refreshing happening every 1 sec by default, while Solr  
> doesn't (or are you comparing ES and SolrCloud?).  
> ...  
> I think more shards would speed things up, as would bulk API and setting  
> the number of replicas to 0.  
> You could also increase the number of concurrent indexers in order to  
> maximize CPU and/or disk IO.
> 
> ## Otis
> 
> Sematext is hiring Elasticsearch engineers World-Wide --[Jobs - Sematext](http://sematext.com/about/jobs.html)
> 
> On Sunday, March 4, 2012 5:05:58 PM UTC+8, Lior Cohen barnea wrote:
> 
> > Hi all,
> 
> > I am doing a POC on ES and Solr cloud.
> 
> > I want to improve indexing throughput which is quite less than Solr's.  
> > What approach should i take:  
> > Increase number of shards and/or nodes?  
> > Config my client with all nodes IP's?  
> > Run 1 thread per node and configure 1 node IP per client?  
> > Changes the segments settings ? if so, how?
> 
> > Thanks in advance,  
> > Lior.  
> > On Sunday, March 4, 2012 5:05:58 PM UTC+8, Lior Cohen barnea wrote:
> 
> > Hi all,
> 
> > I am doing a POC on ES and Solr cloud.
> 
> > I want to improve indexing throughput which is quite less than Solr's.  
> > What approach should i take:  
> > Increase number of shards and/or nodes?  
> > Config my client with all nodes IP's?  
> > Run 1 thread per node and configure 1 node IP per client?  
> > Changes the segments settings ? if so, how?
> 
> > Thanks in advance,  
> > Lior.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 5, 2012, 3:02pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/8 "2012-03-05T15:02:08Z")

</div>

A few more things to add to this great thread:

1. By default, elasticsearch also indexed all the json data into an \_all field. This is very convenient, but does add to the indexing time (and index size). You can easily disable it with the all mapping, but its different compared to Solr, which doesn't do it.
2. It was already mentioned, but indexing process is different compared to Solr (I am not sure which SolrCloud you are using, but it has changed considerable over time). It does active replication, and there is no need to commit (ES commits periodically by itself, and has a transaction log). See this video for a bit more info: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/videos/2011/08/09/road-to-a-distributed-searchengine-berlinbuzzwords.html).
3. By default, ES stores the whole json so you can easily get it back if you issue a GET or a search. If you don't store anything in your solr mappings, then it will incur that additional overhead comparing the two.
4. When indexing data, if your client ends up doing batch indexing with Solr, make sure to do batch indexing with ES as well.

On Monday, March 5, 2012 at 4:21 PM, haarts wrote:

> Hi Thomas,
> 
> Those are some impressive numbers. Would you mind sharing on what kind of machines you are running? We are struggling indexing 500M documents, reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1 simple spinner). Performance indexing is acceptable. But first time query performance isn't great (seconds...).
> 
> On Monday, 5 March 2012 15:00:42 UTC+1, Thomas Peuss wrote:
> 
> > Hi Lior!
> > 
> > Am Sonntag, 4. März 2012 10:05:58 UTC+1 schrieb Lior Cohen barnea:
> > 
> > > I want to improve indexing throughput which is quite less than Solr's.  
> > > What approach should i take:  
> > > Increase number of shards and/or nodes?  
> > > Config my client with all nodes IP's?  
> > > Run 1 thread per node and configure 1 node IP per client?  
> > > Changes the segments settings ? if so, how?
> > 
> > First of all make sure that you don't compare apples with oranges. When you calculate the indexing speed of Solr you need to take the "commit"-time into account you don't need with ES. And you will see that Solr gets "slower".
> > 
> > Increasing the shard count is only useful if you have enough CPU cores and disk spindles available. You should increase the refresh\_interval or disable it alltogether and trigger it manually if you get to really big index sizes (we have \>1 billion docs in some indices while having \>500 indices). If you run on many nodes your network can be the next
> > 
> > Bulk indexing gives you a big performance gain over doing single updates so you should consider that.
> > 
> > But in the end you need to benchmark with your amount of data (docs per second, index size etc.) and different settings. If need to guarantee some sort of SLA there is no other way around it.
> > 
> > CU  
> > Thomas

---

<div class="post-metadata">

**Author:** ![Thomas\_Peuss](https://avatars.discourse-cdn.com/v4/letter/t/22d042/32.png) [@Thomas\_Peuss](https://discuss.elastic.co/u/Thomas_Peuss)\
**Post date:** [March 5, 2012, 4:03pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/9 "2012-03-05T16:03:38Z")

</div>

Hi!

Am Montag, 5. März 2012 15:21:11 UTC+1 schrieb haarts:

> Those are some impressive numbers. Would you mind sharing on what kind of  
> machines you are running? We are struggling indexing 500M documents,  
> reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1  
> simple spinner). Performance indexing is acceptable. But first time query  
> performance isn't great (seconds...).

We are running a 8-node cluster in two datacenters (4 nodes per DC). Each  
machine has 24 cores, 32GB RAM and 8 disks (extendable to 16 disks) running  
RHEL 6.1. The machines are not dedicated to ES alone (we use 50% of the  
cores for number crunching without I/O involved). Currently we are running  
with 16 shards and 1 replica.

We are currently peaking at 400 docs/s but the numbers are rising... 😉

You should try to insert with many threads in parallel (we use 16).  
Important here is that you wait for the response from ES because otherwise  
you will overload ES.

CU  
Thomas

---

<div class="post-metadata">

**Author:** ![haarts](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/haarts/32/2972_2.png) [@haarts](https://discuss.elastic.co/u/haarts)\
**Post date:** [March 5, 2012, 4:11pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/10 "2012-03-05T16:11:25Z")

</div>

Thanks a lot for the insight! I'd better convince by boss to buy 16 disk  
machines. 🙂

On Monday, 5 March 2012 17:03:38 UTC+1, Thomas Peuss wrote:

> Hi!
> 
> Am Montag, 5. März 2012 15:21:11 UTC+1 schrieb haarts:
> 
> > Those are some impressive numbers. Would you mind sharing on what kind of  
> > machines you are running? We are struggling indexing 500M documents,  
> > reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1  
> > simple spinner). Performance indexing is acceptable. But first time query  
> > performance isn't great (seconds...).
> 
> We are running a 8-node cluster in two datacenters (4 nodes per DC). Each  
> machine has 24 cores, 32GB RAM and 8 disks (extendable to 16 disks) running  
> RHEL 6.1. The machines are not dedicated to ES alone (we use 50% of the  
> cores for number crunching without I/O involved). Currently we are running  
> with 16 shards and 1 replica.
> 
> We are currently peaking at 400 docs/s but the numbers are rising... 😉
> 
> You should try to insert with many threads in parallel (we use 16).  
> Important here is that you wait for the response from ES because otherwise  
> you will overload ES.
> 
> CU  
> Thomas

---

<div class="post-metadata">

**Author:** ![Craig\_Brown](https://avatars.discourse-cdn.com/v4/letter/c/ce7236/32.png) [@Craig\_Brown](https://discuss.elastic.co/u/Craig_Brown)\
**Post date:** [March 5, 2012, 4:45pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/11 "2012-03-05T16:45:16Z")

</div>

We're running on AWS, 4 C1-XL nodes - 7GB ram, 20 compute units (8 virtual  
cores). We allocate 4GB ram to ES. Each node as 1-500GB EBS instance for  
storage. We run 26 shards with 0 replicas when indexing. It's MUCH faster  
to index with 0 replicas if you can, then up the replica number after  
indexing, than it is to index with 1 or more replicas. We set  
refresh\_interval to 30s and merge.policy.merge\_factor to 30. After  
indexing, we set them back to 1s and 1 and run optimize. This really helps.  
Our documents are about 2k-5k in size and we index about 10k-12k docs/sec  
initially. After 240m docs, we're in the 5k-6k docs/sec range. We wrote our  
own multi-threaded indexing tool to do the work. We enable \_source and  
compression on \_source. We still have \_all enabled though we are not using  
it. We'll disable that in the next round.

- Craig

On Mon, Mar 5, 2012 at 9:11 AM, haarts [harmaarts@gmail.com](mailto:harmaarts@gmail.com) wrote:

> Thanks a lot for the insight! I'd better convince by boss to buy 16 disk  
> machines. 🙂
> 
> On Monday, 5 March 2012 17:03:38 UTC+1, Thomas Peuss wrote:
> 
> > Hi!
> > 
> > Am Montag, 5. März 2012 15:21:11 UTC+1 schrieb haarts:
> > 
> > > Those are some impressive numbers. Would you mind sharing on what kind  
> > > of machines you are running? We are struggling indexing 500M documents,  
> > > reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1  
> > > simple spinner). Performance indexing is acceptable. But first time query  
> > > performance isn't great (seconds...).
> > 
> > We are running a 8-node cluster in two datacenters (4 nodes per DC). Each  
> > machine has 24 cores, 32GB RAM and 8 disks (extendable to 16 disks) running  
> > RHEL 6.1. The machines are not dedicated to ES alone (we use 50% of the  
> > cores for number crunching without I/O involved). Currently we are running  
> > with 16 shards and 1 replica.
> > 
> > We are currently peaking at 400 docs/s but the numbers are rising... 😉
> > 
> > You should try to insert with many threads in parallel (we use 16).  
> > Important here is that you wait for the response from ES because otherwise  
> > you will overload ES.
> > 
> > CU  
> > Thomas

--  
…  
CRAIG BROWN  
chief architect  
youwho, Inc.

_[www.youwho.com](http://www.youwho.com)_ [http://www.youwho.com/](http://www.youwho.com/)

T: 801.855. 0921  
M: 801.913. 0939

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 5, 2012, 8:31pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/12 "2012-03-05T20:31:24Z")

</div>

Note that the merge factor parameter does not apply to the default tiered merge policy. In any case, setting it to 1 is not recommended, since you can always control the number of shards it will optimize down to in the optimize call API.

On Monday, March 5, 2012 at 6:45 PM, Craig Brown wrote:

> We're running on AWS, 4 C1-XL nodes - 7GB ram, 20 compute units (8 virtual cores). We allocate 4GB ram to ES. Each node as 1-500GB EBS instance for storage. We run 26 shards with 0 replicas when indexing. It's MUCH faster to index with 0 replicas if you can, then up the replica number after indexing, than it is to index with 1 or more replicas. We set refresh\_interval to 30s and merge.policy.merge\_factor to 30. After indexing, we set them back to 1s and 1 and run optimize. This really helps.  
> Our documents are about 2k-5k in size and we index about 10k-12k docs/sec initially. After 240m docs, we're in the 5k-6k docs/sec range. We wrote our own multi-threaded indexing tool to do the work. We enable \_source and compression on \_source. We still have \_all enabled though we are not using it. We'll disable that in the next round.
> 
> - Craig
> 
> On Mon, Mar 5, 2012 at 9:11 AM, haarts \<[harmaarts@gmail.com](mailto:harmaarts@gmail.com) ([mailto:harmaarts@gmail.com](mailto:harmaarts@gmail.com))\> wrote:
> 
> > Thanks a lot for the insight! I'd better convince by boss to buy 16 disk machines. 🙂
> > 
> > On Monday, 5 March 2012 17:03:38 UTC+1, Thomas Peuss wrote:
> > 
> > > Hi!
> > > 
> > > Am Montag, 5. März 2012 15:21:11 UTC+1 schrieb haarts:
> > > 
> > > > Those are some impressive numbers. Would you mind sharing on what kind of machines you are running? We are struggling indexing 500M documents, reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1 simple spinner). Performance indexing is acceptable. But first time query performance isn't great (seconds...).
> > > 
> > > We are running a 8-node cluster in two datacenters (4 nodes per DC). Each machine has 24 cores, 32GB RAM and 8 disks (extendable to 16 disks) running RHEL 6.1. The machines are not dedicated to ES alone (we use 50% of the cores for number crunching without I/O involved). Currently we are running with 16 shards and 1 replica.
> > > 
> > > We are currently peaking at 400 docs/s but the numbers are rising... 😉
> > > 
> > > You should try to insert with many threads in parallel (we use 16). Important here is that you wait for the response from ES because otherwise you will overload ES.
> > > 
> > > CU  
> > > Thomas
> 
> --  
> …  
> CRAIG BROWN  
> chief architect  
> youwho, Inc.
> 
> [www.youwho.com](http://www.youwho.com) ([http://www.youwho.com/](http://www.youwho.com/))
> 
> T: 801.855. 0921  
> M: 801.913. 0939

---

<div class="post-metadata">

**Author:** ![Lior\_Cohen\_barnea](https://avatars.discourse-cdn.com/v4/letter/l/dec6dc/32.png) [@Lior\_Cohen\_barnea](https://discuss.elastic.co/u/Lior_Cohen_barnea)\
**Post date:** [March 6, 2012, 7:52am UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/13 "2012-03-06T07:52:53Z")

</div>

Thanks for all the answers.

It seem that when i used a client node

`
Node node = nodeBuilder().client(true).node();
Client client = node.client()
`

my indexing time was much faster than when i used transport client to  
a local node ...

should it be like that ?

On Mar 5, 10:31 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Note that the merge factor parameter does not apply to the default tiered merge policy. In any case, setting it to 1 is not recommended, since you can always control the number of shards it will optimize down to in the optimize call API.
> 
> On Monday, March 5, 2012 at 6:45 PM, Craig Brown wrote:
> 
> > We're running on AWS, 4 C1-XL nodes - 7GB ram, 20 compute units (8 virtual cores). We allocate 4GB ram to ES. Each node as 1-500GB EBS instance for storage. We run 26 shards with 0 replicas when indexing. It's MUCH faster to index with 0 replicas if you can, then up the replica number after indexing, than it is to index with 1 or more replicas. We set refresh\_interval to 30s and merge.policy.merge\_factor to 30. After indexing, we set them back to 1s and 1 and run optimize. This really helps.  
> > Our documents are about 2k-5k in size and we index about 10k-12k docs/sec initially. After 240m docs, we're in the 5k-6k docs/sec range. We wrote our own multi-threaded indexing tool to do the work. We enable \_source and compression on \_source. We still have \_all enabled though we are not using it. We'll disable that in the next round.
> 
> > - Craig
> 
> > On Mon, Mar 5, 2012 at 9:11 AM, haarts \<[harmaa...@gmail.com](mailto:harmaa...@gmail.com) ([mailto:harmaa...@gmail.com](mailto:harmaa...@gmail.com))\> wrote:
> > 
> > > Thanks a lot for the insight! I'd better convince by boss to buy 16 disk machines. 🙂
> 
> > > On Monday, 5 March 2012 17:03:38 UTC+1, Thomas Peuss wrote:
> > > 
> > > > Hi!
> 
> > > > Am Montag, 5. März 2012 15:21:11 UTC+1 schrieb haarts:
> > > > 
> > > > > Those are some impressive numbers. Would you mind sharing on what kind of machines you are running? We are struggling indexing 500M documents, reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1 simple spinner). Performance indexing is acceptable. But first time query performance isn't great (seconds...).
> 
> > > > We are running a 8-node cluster in two datacenters (4 nodes per DC). Each machine has 24 cores, 32GB RAM and 8 disks (extendable to 16 disks) running RHEL 6.1. The machines are not dedicated to ES alone (we use 50% of the cores for number crunching without I/O involved). Currently we are running with 16 shards and 1 replica.
> 
> > > > We are currently peaking at 400 docs/s but the numbers are rising... 😉
> 
> > > > You should try to insert with many threads in parallel (we use 16). Important here is that you wait for the response from ES because otherwise you will overload ES.
> 
> > > > CU  
> > > > Thomas
> 
> > --  
> > …  
> > CRAIG BROWN  
> > chief architect  
> > youwho, Inc.
> 
> > [www.youwho.com](http://www.youwho.com)([http://www.youwho.com/](http://www.youwho.com/))
> 
> > T: 801.855. 0921  
> > M: 801.913. 0939

---

<div class="post-metadata">

**Author:** ![Craig\_Brown](https://avatars.discourse-cdn.com/v4/letter/c/ce7236/32.png) [@Craig\_Brown](https://discuss.elastic.co/u/Craig_Brown)\
**Post date:** [March 6, 2012, 4:59pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/14 "2012-03-06T16:59:57Z")

</div>

OK, thanks for the advice.

- Craig

On Mon, Mar 5, 2012 at 1:31 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Note that the merge factor parameter does not apply to the default tiered  
> merge policy. In any case, setting it to 1 is not recommended, since you  
> can always control the number of shards it will optimize down to in the  
> optimize call API.
> 
> On Monday, March 5, 2012 at 6:45 PM, Craig Brown wrote:
> 
> We're running on AWS, 4 C1-XL nodes - 7GB ram, 20 compute units (8 virtual  
> cores). We allocate 4GB ram to ES. Each node as 1-500GB EBS instance for  
> storage. We run 26 shards with 0 replicas when indexing. It's MUCH faster  
> to index with 0 replicas if you can, then up the replica number after  
> indexing, than it is to index with 1 or more replicas. We set  
> refresh\_interval to 30s and merge.policy.merge\_factor to 30. After  
> indexing, we set them back to 1s and 1 and run optimize. This really helps.  
> Our documents are about 2k-5k in size and we index about 10k-12k docs/sec  
> initially. After 240m docs, we're in the 5k-6k docs/sec range. We wrote our  
> own multi-threaded indexing tool to do the work. We enable \_source and  
> compression on \_source. We still have \_all enabled though we are not using  
> it. We'll disable that in the next round.
> 
> - Craig
> 
> On Mon, Mar 5, 2012 at 9:11 AM, haarts [harmaarts@gmail.com](mailto:harmaarts@gmail.com) wrote:
> 
> Thanks a lot for the insight! I'd better convince by boss to buy 16 disk  
> machines. 🙂
> 
> On Monday, 5 March 2012 17:03:38 UTC+1, Thomas Peuss wrote:
> 
> Hi!
> 
> Am Montag, 5. März 2012 15:21:11 UTC+1 schrieb haarts:
> 
> Those are some impressive numbers. Would you mind sharing on what kind of  
> machines you are running? We are struggling indexing 500M documents,  
> reaching 1000+ inserts per second on a 3 node cluster (8 core i7 24GB, 1  
> simple spinner). Performance indexing is acceptable. But first time query  
> performance isn't great (seconds...).
> 
> We are running a 8-node cluster in two datacenters (4 nodes per DC). Each  
> machine has 24 cores, 32GB RAM and 8 disks (extendable to 16 disks) running  
> RHEL 6.1. The machines are not dedicated to ES alone (we use 50% of the  
> cores for number crunching without I/O involved). Currently we are running  
> with 16 shards and 1 replica.
> 
> We are currently peaking at 400 docs/s but the numbers are rising... 😉
> 
> You should try to insert with many threads in parallel (we use 16).  
> Important here is that you wait for the response from ES because otherwise  
> you will overload ES.
> 
> CU  
> Thomas
> 
> --  
> …  
> CRAIG BROWN  
> chief architect  
> youwho, Inc.
> 
> _[www.youwho.com](http://www.youwho.com)_ [http://www.youwho.com/](http://www.youwho.com/)
> 
> T: 801.855. 0921  
> M: 801.913. 0939

--  
…  
CRAIG BROWN  
chief architect  
youwho, Inc.

_[www.youwho.com](http://www.youwho.com)_ [http://www.youwho.com/](http://www.youwho.com/)

T: 801.855. 0921  
M: 801.913. 0939

---

<div class="post-metadata">

**Author:** ![Craig\_Brown](https://avatars.discourse-cdn.com/v4/letter/c/ce7236/32.png) [@Craig\_Brown](https://discuss.elastic.co/u/Craig_Brown)\
**Post date:** [March 6, 2012, 5:02pm UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/15 "2012-03-06T17:02:43Z")

</div>

When you connect to the cluster this way, your client becomes a node in the  
cluster and has access to all of the routing and other information. The  
client is therefore much more efficient because it can directly communicate  
with the node that has to do the work.

- Craig

On Tue, Mar 6, 2012 at 12:52 AM, Lior Cohen barnea \<  
[liorcohenbarnea@gmail.com](mailto:liorcohenbarnea@gmail.com)\> wrote:

> Thanks for all the answers.
> 
> It seem that when i used a client node
> 
> `
> Node node = nodeBuilder().client(true).node();
> Client client = node.client()
> `
> 
> my indexing time was much faster than when i used transport client to  
> a local node ...
> 
> should it be like that ?
> 
> On Mar 5, 10:31 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > Note that the merge factor parameter does not apply to the default  
> > tiered merge policy. In any case, setting it to 1 is not recommended, since  
> > you can always control the number of shards it will optimize down to in the  
> > optimize call API.
> > 
> > On Monday, March 5, 2012 at 6:45 PM, Craig Brown wrote:
> > 
> > > We're running on AWS, 4 C1-XL nodes - 7GB ram, 20 compute units (8  
> > > virtual cores). We allocate 4GB ram to ES. Each node as 1-500GB EBS  
> > > instance for storage. We run 26 shards with 0 replicas when indexing. It's  
> > > MUCH faster to index with 0 replicas if you can, then up the replica number  
> > > after indexing, than it is to index with 1 or more replicas. We set  
> > > refresh\_interval to 30s and merge.policy.merge\_factor to 30. After  
> > > indexing, we set them back to 1s and 1 and run optimize. This really helps.  
> > > Our documents are about 2k-5k in size and we index about 10k-12k  
> > > docs/sec initially. After 240m docs, we're in the 5k-6k docs/sec range. We  
> > > wrote our own multi-threaded indexing tool to do the work. We enable  
> > > \_source and compression on \_source. We still have \_all enabled though we  
> > > are not using it. We'll disable that in the next round.
> > 
> > > - Craig
> > 
> > > On Mon, Mar 5, 2012 at 9:11 AM, haarts \<[harmaa...@gmail.com](mailto:harmaa...@gmail.com) (mailto:  
> > > [harmaa...@gmail.com](mailto:harmaa...@gmail.com))\> wrote:
> > > 
> > > > Thanks a lot for the insight! I'd better convince by boss to buy 16  
> > > > disk machines. 🙂
> > 
> > > > On Monday, 5 March 2012 17:03:38 UTC+1, Thomas Peuss wrote:
> > > > 
> > > > > Hi!
> > 
> > > > > Am Montag, 5. März 2012 15:21:11 UTC+1 schrieb haarts:
> > > > > 
> > > > > > Those are some impressive numbers. Would you mind sharing on  
> > > > > > what kind of machines you are running? We are struggling indexing 500M  
> > > > > > documents, reaching 1000+ inserts per second on a 3 node cluster (8 core i7  
> > > > > > 24GB, 1 simple spinner). Performance indexing is acceptable. But first time  
> > > > > > query performance isn't great (seconds...).
> > 
> > > > > We are running a 8-node cluster in two datacenters (4 nodes per  
> > > > > DC). Each machine has 24 cores, 32GB RAM and 8 disks (extendable to 16  
> > > > > disks) running RHEL 6.1. The machines are not dedicated to ES alone (we use  
> > > > > 50% of the cores for number crunching without I/O involved). Currently we  
> > > > > are running with 16 shards and 1 replica.
> > 
> > > > > We are currently peaking at 400 docs/s but the numbers are  
> > > > > rising... 😉
> > 
> > > > > You should try to insert with many threads in parallel (we use  
> > > > > 16). Important here is that you wait for the response from ES because  
> > > > > otherwise you will overload ES.
> > 
> > > > > CU  
> > > > > Thomas
> > 
> > > --  
> > > …  
> > > CRAIG BROWN  
> > > chief architect  
> > > youwho, Inc.
> > 
> > > [www.youwho.com](http://www.youwho.com)([http://www.youwho.com/](http://www.youwho.com/))
> > 
> > > T: 801.855. 0921  
> > > M: 801.913. 0939

--  
…  
CRAIG BROWN  
chief architect  
youwho, Inc.

_[www.youwho.com](http://www.youwho.com)_ [http://www.youwho.com/](http://www.youwho.com/)

T: 801.855. 0921  
M: 801.913. 0939

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:37am UTC](https://discuss.elastic.co/t/improve-indexing-throughput/6877/16 "2017-07-06T03:37:11Z")

</div>


