# Why the bulk api (python) has low efficiency in this way?

**URL:** <https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970>\
**Category:** Elasticsearch\
**Created:** [December 26, 2016, 3:28am UTC](https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970 "2016-12-26T03:28:38Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 26, 2016, 3:28am UTC](https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970/1 "2016-12-26T03:28:38Z")

</div>

Hi:  
I am testing the BULK insert(index) efficiency.  
the python code as follow:

```
from elasticsearch import Elasticsearch
from elasticsearch import helpers
import time
es = Elasticsearch("127.0.0.1")
i = 0 
data_list=[]
for i in range(50000000):
    data_list.append({"_index":"stress","_type":"test","_source":{
            "collectTime": 1414709176,  
            "deltatime": 300,  
            "deviceId": "48572",  
            "getway": 0,  
            "ifindiscards": 0,  
            "ifindiscardspps": 0,  
             ...
             ...
             ...
            "ifinunknownprotos": 0,  
            "ifinunknownprotospps": 0
             }})
    if len(data_list) == 5000:
        helpers.bulk(es,data_list)
        data_list[:]=[]
if len(data_list) != 0:
    helpers.bulk(es,data_list) 

```

but the efficiency is very low, about 2000 docs/s，and the cpu usage is only 300% (12 core),it is expected to 1200%.

when i running esrally to benchmark my elasticsearch the speed can reach 9000 docs/s , and the cpu usage can reach 1200%.

Do i miss something?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 26, 2016, 4:36am UTC](https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970/2 "2016-12-26T04:36:45Z")

</div>

The size of documents and the number and types of fields will affect indexing rate as it will define the amount of work Elasticsearch need to do for each document. The reason for the low CPU usage is however probably that Rally has the ability to partition work and use multiple worker processes bulk indexing in parallell, whereas your script appear to be single threaded. Do you see better resource utilization if you increase the number of your scripts that you run concurrently?

---

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 26, 2016, 5:23am UTC](https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970/3 "2016-12-26T05:23:09Z")

</div>

The problem is that when I run eight scripts ,the cpu usage still at 300% - 400%, and the exception [EsRejectedExecutionException[rejected execution (queue capacity 50)] catched.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 26, 2016, 5:35am UTC](https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970/4 "2016-12-26T05:35:03Z")

</div>

Don't jump directly to 8 scripts. Instead start with 2 and slowly increase. You may also want to try with a smaller bulk size. Have you run Rally with this type of events into an index with the same number of shards?

---

<div class="post-metadata">

**Author:** ![a943013827](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@a943013827](https://discuss.elastic.co/u/a943013827)\
**Post date:** [December 28, 2016, 12:21pm UTC](https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970/5 "2016-12-28T12:21:21Z")

</div>

Hi,I found the solution，  
use : es.bulk()  
do not use : helpers.bulk()

the es.bulk has high efficiency.

OK,  
helpers.bulk(chunk\_size=5000) also has high efficiency (a little lower than es.bulk ) .

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 25, 2017, 12:21pm UTC](https://discuss.elastic.co/t/why-the-bulk-api-python-has-low-efficiency-in-this-way/69970/6 "2017-01-25T12:21:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
