# ElasticSearch bulk api performance

**URL:** <https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818>\
**Category:** Elasticsearch\
**Created:** [July 8, 2011, 6:10am UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818 "2011-07-08T06:10:51Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![allwefantasy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/allwefantasy/32/3143_2.png) [@allwefantasy](https://discuss.elastic.co/u/allwefantasy)\
**Post date:** [July 8, 2011, 6:10am UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/1 "2011-07-08T06:10:51Z")

</div>

here is my code using bulk api:  
[https://gist.github.com/1071233](https://gist.github.com/1071233)

single node (node info and performance log)  
[https://gist.github.com/1071237](https://gist.github.com/1071237)

i find it really really slow. 20 documents per second! if i start two nodes in different machine,  
it becomes only 9 documents per second.  
anyone know why?

---

<div class="post-metadata">

**Author:** ![imarcticblue](https://avatars.discourse-cdn.com/v4/letter/i/c67d28/32.png) [@imarcticblue](https://discuss.elastic.co/u/imarcticblue)\
**Post date:** [July 8, 2011, 5:15pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/2 "2011-07-08T17:15:10Z")

</div>

How large are your documents and how many are you indexing at a time? We're not using a DataItem, just raw JSON and we can index 40M records in about 3:30 at 3200 docs/sec. Average record size is .5k. This is on AWS with 2 large nodes, 32 shards, 2 replicas. Your hardware looks beefier than what we have on AWS.

- Craig

---

<div class="post-metadata">

**Author:** ![jrawlings](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jrawlings/32/3183_2.png) [@jrawlings](https://discuss.elastic.co/u/jrawlings)\
**Post date:** [July 8, 2011, 9:25pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/3 "2011-07-08T21:25:32Z")

</div>

Are you indexing to a local ES node? How large are the DataItems?

From my experience doing bulk prepares to a non-local ES node, I  
noticed I maxed out my network connection (10mb) quite quickly..

On Jul 7, 11:10 pm, allwefantasy [allwefant...@gmail.com](mailto:allwefant...@gmail.com) wrote:

> here is my code using bulk api:[https://gist.github.com/1071233https://gist.github.com/1071233](https://gist.github.com/1071233https://gist.github.com/1071233)
> 
> single node (node info and performance log)[https://gist.github.com/1071237https://gist.github.com/1071237](https://gist.github.com/1071237https://gist.github.com/1071237)
> 
> i find it really really slow. 20 documents per second! if i start two nodes  
> in different machine,  
> it becomes only 9 documents per second.  
> anyone know why?
> 
> --  
> View this message in context:[http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap](http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap)...  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![allwefantasy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/allwefantasy/32/3143_2.png) [@allwefantasy](https://discuss.elastic.co/u/allwefantasy)\
**Post date:** [July 9, 2011, 3:02pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/4 "2011-07-09T15:02:17Z")

</div>

DataItem Class contains id and source fields. source is raw json data from blog article. 3200 docs/sec is really awesome! i still have no idea why it so slow in my application

---

<div class="post-metadata">

**Author:** ![allwefantasy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/allwefantasy/32/3143_2.png) [@allwefantasy](https://discuss.elastic.co/u/allwefantasy)\
**Post date:** [July 9, 2011, 3:08pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/5 "2011-07-09T15:08:55Z")

</div>

yes.Index to a local node. DataItem contains one blog article. and each time I bulk index 1000 DataItems.

From: jrawlings [via Elasticsearch Users]  
Sent: Saturday, July 09, 2011 5:25 AM  
To: allwefantasy  
Subject: Re: Elasticsearch bulk api performance

Are you indexing to a local ES node? How large are the DataItems?

From my experience doing bulk prepares to a non-local ES node, I  
noticed I maxed out my network connection (10mb) quite quickly..

On Jul 7, 11:10 pm, allwefantasy \<[hidden email]\> wrote:

> here is my code using bulk api:[https://gist.github.com/1071233https://gist.github.com/1071233](https://gist.github.com/1071233https://gist.github.com/1071233)
> 
> single node (node info and performance log)[https://gist.github.com/1071237https://gist.github.com/1071237](https://gist.github.com/1071237https://gist.github.com/1071237)
> 
> i find it really really slow. 20 documents per second! if i start two nodes  
> in different machine,  
> it becomes only 9 documents per second.  
> anyone know why?
> 
> --  
> View this message in context:[http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap](http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap)...  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

* * *

If you reply to this email, your message will be added to the discussion below:  
[http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-api-performance-tp3150866p3153370.html](http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-api-performance-tp3150866p3153370.html)  
To unsubscribe from Elasticsearch bulk api performance, click here.

---

<div class="post-metadata">

**Author:** ![allwefantasy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/allwefantasy/32/3143_2.png) [@allwefantasy](https://discuss.elastic.co/u/allwefantasy)\
**Post date:** [July 9, 2011, 3:11pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/6 "2011-07-09T15:11:06Z")

</div>

6 milliam blog articles will be indexed and whole index files are 62G .  
1000 DataItems indexed at a time.

From: imarcticblue [via ElasticSearch Users]  
Sent: Saturday, July 09, 2011 1:15 AM  
To: allwefantasy  
Subject: Re: ElasticSearch bulk api performance

How large are your documents and how many are you indexing at a time? We're not using a DataItem, just raw JSON and we can index 40M records in about 3:30 at 3200 docs/sec. Average record size is .5k. This is on AWS with 2 large nodes, 32 shards, 2 replicas. Your hardware looks beefier than what we have on AWS.

- Craig

* * *

If you reply to this email, your message will be added to the discussion below:  
[http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-api-performance-tp3150866p3152481.html](http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-api-performance-tp3150866p3152481.html)  
To unsubscribe from ElasticSearch bulk api performance, click here.

---

<div class="post-metadata">

**Author:** ![Craig\_Brown](https://avatars.discourse-cdn.com/v4/letter/c/ce7236/32.png) [@Craig\_Brown](https://discuss.elastic.co/u/Craig_Brown)\
**Post date:** [July 11, 2011, 11:38pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/7 "2011-07-11T23:38:34Z")

</div>

So your docs are about 10K each? Are you doing any kind of other  
transformation on your data? The code you showed is virtually the same  
as mine, but I'm not using a DataItem. My data files are JSON, one per  
row. I simple read each row, set the id and source, then use  
client.prepareIndex(). We index 10,000 docs at a time. THe files  
contain 20M docs and the files are compressed using gz. It's basically  
just as fast to read from compressed files plus you get much smaller  
files to push around 🙂

- Craig

On Jul 9, 8:58 pm, allwefantasy [allwefant...@gmail.com](mailto:allwefant...@gmail.com) wrote:

> 6 milliam blog articles will be indexed and whole index files are 62G .  
> 1000 DataItems indexed at a time.
> 
> From: imarcticblue [via Elasticsearch Users]  
> Sent: Saturday, July 09, 2011 1:15 AM  
> To: allwefantasy  
> Subject: Re: Elasticsearch bulk api performance
> 
> How large are your documents and how many are you indexing at a time? We're not using a DataItem, just raw JSON and we can index 40M records in about 3:30 at 3200 docs/sec. Average record size is .5k. This is on AWS with 2 large nodes, 32 shards, 2 replicas. Your hardware looks beefier than what we have on AWS.
> 
> - Craig
> 
> * * *
> 
> If you reply to this email, your message will be added to the discussion below:[http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap](http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap)...  
> To unsubscribe from Elasticsearch bulk api performance, click here.
> 
> --  
> View this message in context:[http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap](http://elasticsearch-users.115913.n3.nabble.com/ElasticSearch-bulk-ap)...  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:00am UTC](https://discuss.elastic.co/t/elasticsearch-bulk-api-performance/4818/8 "2017-07-06T04:00:58Z")

</div>


