# Bulk load performance

**URL:** <https://discuss.elastic.co/t/bulk-load-performance/20829>\
**Category:** Elasticsearch\
**Created:** [November 19, 2014, 7:53am UTC](https://discuss.elastic.co/t/bulk-load-performance/20829 "2014-11-19T07:53:08Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![xaviertrujillo111](https://avatars.discourse-cdn.com/v4/letter/x/f1d935/32.png) [@xaviertrujillo111](https://discuss.elastic.co/u/xaviertrujillo111)\
**Post date:** [November 19, 2014, 7:53am UTC](https://discuss.elastic.co/t/bulk-load-performance/20829/1 "2014-11-19T07:53:08Z")

</div>

Hello,

I'm trying to do a bulk load of ~10M JSON docs (12.8Gb) with some  
geographical information into an elasticsearch index. With our current  
params, the loading is taking around 20-25 minutes to run, but we think it  
should be faster. Are these numbers similar to what other users are  
getting? Do you have any hints on how to get better performance? Any help  
will be appreciated. Please find the details below.

Our ES cluster is version 1.1.1 with 11 nodes, and we are using  
Elasticsearch-MapReduce libraries 2.0.2 to do the bulk-load, setting the  
numbers of reducers to 11. Other params we use are:

es.input.json=true  
es.mapping.id=id  
es.batch.size.bytes=10M  
es.batch.size.entries=10000

The average doc size is 1.3Kb, and each doc contains a "bbox" field with  
the shape definition like this:

"bbox": {  
"type": "envelope",  
"coordinates": [  
[  
-77.08488844489459,  
38.9502995339637  
],  
[  
-77.0844224567727,  
38.9502305534064  
]  
]  
}

We are using the following mapping for this index, because these are the 3  
fields of our docs we are more interested in:

{  
"properties": {  
"bbox": {  
"precision": "10m",  
"tree": "quadtree",  
"type": "geo\_shape"  
},  
"id": {  
"type": "string",  
"index": "not\_analyzed"  
},  
"streets": {  
"type": "string"  
}  
}  
}

This is a typical output of the MapReduce job:

14/11/17 09:05:44 INFO mapred.JobClient: Elasticsearch Hadoop Counters  
14/11/17 09:05:44 INFO mapred.JobClient: Bulk Retries=0  
14/11/17 09:05:44 INFO mapred.JobClient: Bulk Retries Total Time(ms)=0  
14/11/17 09:05:44 INFO mapred.JobClient: Bulk Total=1375  
14/11/17 09:05:44 INFO mapred.JobClient: Bulk Total Time(ms)=11714959  
14/11/17 09:05:44 INFO mapred.JobClient: Bytes Accepted=14351811146  
14/11/17 09:05:44 INFO mapred.JobClient: Bytes Received=5498829  
14/11/17 09:05:44 INFO mapred.JobClient: Bytes Retried=0  
14/11/17 09:05:44 INFO mapred.JobClient: Bytes Sent=14351811146  
14/11/17 09:05:44 INFO mapred.JobClient: Documents Accepted=10129699  
14/11/17 09:05:44 INFO mapred.JobClient: Documents Received=0  
14/11/17 09:05:44 INFO mapred.JobClient: Documents Retried=0  
14/11/17 09:05:44 INFO mapred.JobClient: Documents Sent=10129699  
14/11/17 09:05:44 INFO mapred.JobClient: Network Retries=0  
14/11/17 09:05:44 INFO mapred.JobClient: Network Total Time(ms)=11732552  
14/11/17 09:05:44 INFO mapred.JobClient: Node Retries=0  
14/11/17 09:05:44 INFO mapred.JobClient: Scroll Total=0  
14/11/17 09:05:44 INFO mapred.JobClient: Scroll Total Time(ms)=0

Thanks,  
Xavier.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nickcanz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nickcanz/32/44873_2.png) [@nickcanz](https://discuss.elastic.co/u/nickcanz)\
**Post date:** [November 19, 2014, 3:46pm UTC](https://discuss.elastic.co/t/bulk-load-performance/20829/2 "2014-11-19T15:46:42Z")

</div>

On the index settings side, you can dynamically turn off the index  
refresh\_interval and also reduce the number of shard replicas for the  
duration of the bulk import.

Described here:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

On Wed, Nov 19, 2014 at 2:53 AM, [xaviertrujillo111@gmail.com](mailto:xaviertrujillo111@gmail.com) wrote:

> Hello,
> 
> I'm trying to do a bulk load of ~10M JSON docs (12.8Gb) with some  
> geographical information into an elasticsearch index. With our current  
> params, the loading is taking around 20-25 minutes to run, but we think it  
> should be faster. Are these numbers similar to what other users are  
> getting? Do you have any hints on how to get better performance? Any help  
> will be appreciated. Please find the details below.
> 
> Our ES cluster is version 1.1.1 with 11 nodes, and we are using  
> Elasticsearch-MapReduce libraries 2.0.2 to do the bulk-load, setting the  
> numbers of reducers to 11. Other params we use are:
> 
> es.input.json=true  
> es.mapping.id=id  
> es.batch.size.bytes=10M  
> es.batch.size.entries=10000
> 
> The average doc size is 1.3Kb, and each doc contains a "bbox" field with  
> the shape definition like this:
> 
> "bbox": {  
> "type": "envelope",  
> "coordinates": [  
> [  
> -77.08488844489459,  
> 38.9502995339637  
> ],  
> [  
> -77.0844224567727,  
> 38.9502305534064  
> ]  
> ]  
> }
> 
> We are using the following mapping for this index, because these are the 3  
> fields of our docs we are more interested in:
> 
> {  
> "properties": {  
> "bbox": {  
> "precision": "10m",  
> "tree": "quadtree",  
> "type": "geo\_shape"  
> },  
> "id": {  
> "type": "string",  
> "index": "not\_analyzed"  
> },  
> "streets": {  
> "type": "string"  
> }  
> }  
> }
> 
> This is a typical output of the MapReduce job:
> 
> 14/11/17 09:05:44 INFO mapred.JobClient: Elasticsearch Hadoop Counters  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Retries=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Retries Total Time(ms)=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Total=1375  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Total Time(ms)=11714959  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Accepted=14351811146  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Received=5498829  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Retried=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Sent=14351811146  
> 14/11/17 09:05:44 INFO mapred.JobClient: Documents Accepted=10129699  
> 14/11/17 09:05:44 INFO mapred.JobClient: Documents Received=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Documents Retried=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Documents Sent=10129699  
> 14/11/17 09:05:44 INFO mapred.JobClient: Network Retries=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Network Total  
> Time(ms)=11732552  
> 14/11/17 09:05:44 INFO mapred.JobClient: Node Retries=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Scroll Total=0  
> 14/11/17 09:05:44 INFO mapred.JobClient: Scroll Total Time(ms)=0
> 
> Thanks,  
> Xavier.
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Nick Canzoneri  
Developer, Wildbit [http://wildbit.com/](http://wildbit.com/)  
Beanstalk [http://beanstalkapp.com/](http://beanstalkapp.com/), Postmark [http://postmarkapp.com/](http://postmarkapp.com/),

> **[DeployBot | Code Deployment Tools | Deploy Code Anywhere](https://deploybot.com)**
>
> Push. Build. Deploy! Instantly build and ship code anywhere in one consistent process for your entire team. DeployBot's code deployment tools work with your existing git repository to deploy new code fast, and with zero downtime. These are the...

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKWm5yPDSs\_PABPi7Ydnr0h8utGAwOTOJuyDvEBm4fNMLG-Sqg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKWm5yPDSs_PABPi7Ydnr0h8utGAwOTOJuyDvEBm4fNMLG-Sqg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![xaviertrujillo111](https://avatars.discourse-cdn.com/v4/letter/x/f1d935/32.png) [@xaviertrujillo111](https://discuss.elastic.co/u/xaviertrujillo111)\
**Post date:** [November 19, 2014, 6:30pm UTC](https://discuss.elastic.co/t/bulk-load-performance/20829/3 "2014-11-19T18:30:26Z")

</div>

Thank you Nick, I tried that but I didn't see a noticeable performance  
improvement.

Also, I tried setting the number of replicas to "0", load the data, then  
put it back to "5", but this is causing some problems with our health check  
scripts, because the index is very large, and the shards seems to be in  
"INITIALIZING" status forever.

Regards.

On Wednesday, November 19, 2014 7:47:10 AM UTC-8, Nick Canzoneri wrote:

> On the index settings side, you can dynamically turn off the index  
> refresh\_interval and also reduce the number of shard replicas for the  
> duration of the bulk import.
> 
> Described here:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/indices-update-settings.html#bulk)
> 
> On Wed, Nov 19, 2014 at 2:53 AM, \<[xaviertr...@gmail.com](mailto:xaviertr...@gmail.com) \<javascript:\>\>  
> wrote:
> 
> > Hello,
> > 
> > I'm trying to do a bulk load of ~10M JSON docs (12.8Gb) with some  
> > geographical information into an elasticsearch index. With our current  
> > params, the loading is taking around 20-25 minutes to run, but we think it  
> > should be faster. Are these numbers similar to what other users are  
> > getting? Do you have any hints on how to get better performance? Any help  
> > will be appreciated. Please find the details below.
> > 
> > Our ES cluster is version 1.1.1 with 11 nodes, and we are using  
> > Elasticsearch-MapReduce libraries 2.0.2 to do the bulk-load, setting the  
> > numbers of reducers to 11. Other params we use are:
> > 
> > es.input.json=true  
> > es.mapping.id=id  
> > es.batch.size.bytes=10M  
> > es.batch.size.entries=10000
> > 
> > The average doc size is 1.3Kb, and each doc contains a "bbox" field with  
> > the shape definition like this:
> > 
> > "bbox": {  
> > "type": "envelope",  
> > "coordinates": [  
> > [  
> > -77.08488844489459,  
> > 38.9502995339637  
> > ],  
> > [  
> > -77.0844224567727,  
> > 38.9502305534064  
> > ]  
> > ]  
> > }
> > 
> > We are using the following mapping for this index, because these are the  
> > 3 fields of our docs we are more interested in:
> > 
> > {  
> > "properties": {  
> > "bbox": {  
> > "precision": "10m",  
> > "tree": "quadtree",  
> > "type": "geo\_shape"  
> > },  
> > "id": {  
> > "type": "string",  
> > "index": "not\_analyzed"  
> > },  
> > "streets": {  
> > "type": "string"  
> > }  
> > }  
> > }
> > 
> > This is a typical output of the MapReduce job:
> > 
> > 14/11/17 09:05:44 INFO mapred.JobClient: Elasticsearch Hadoop Counters  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Retries=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Retries Total Time(ms)=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Total=1375  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bulk Total Time(ms)=11714959  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Accepted=14351811146  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Received=5498829  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Retried=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Bytes Sent=14351811146  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Documents Accepted=10129699  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Documents Received=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Documents Retried=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Documents Sent=10129699  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Network Retries=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Network Total  
> > Time(ms)=11732552  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Node Retries=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Scroll Total=0  
> > 14/11/17 09:05:44 INFO mapred.JobClient: Scroll Total Time(ms)=0
> > 
> > Thanks,  
> > Xavier.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/70956234-78d0-4ee2-9536-398ac529b76a%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> Nick Canzoneri  
> Developer, Wildbit [http://wildbit.com/](http://wildbit.com/)  
> Beanstalk [http://beanstalkapp.com/](http://beanstalkapp.com/), Postmark [http://postmarkapp.com/](http://postmarkapp.com/),  
> [dploy.io](http://dploy.io)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4d5bfe04-50a6-497a-8370-642fa0ed56ff%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4d5bfe04-50a6-497a-8370-642fa0ed56ff%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:48am UTC](https://discuss.elastic.co/t/bulk-load-performance/20829/4 "2017-07-06T00:48:57Z")

</div>


