# Helpers.parallel\_bulk in Python not working?

**URL:** <https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498>\
**Category:** Elasticsearch\
**Created:** [January 19, 2016, 3:52am UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498 "2016-01-19T03:52:27Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Patrick\_Lam](https://avatars.discourse-cdn.com/v4/letter/p/ed8c4c/32.png) [@Patrick\_Lam](https://discuss.elastic.co/u/Patrick_Lam)\
**Post date:** [January 19, 2016, 3:52am UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498/1 "2016-01-19T03:52:27Z")

</div>

Hi,

I'm trying to test out the parallel\_bulk functionality in the python client for elasticsearch and I can't seem to get helpers.parallel\_bulk to work.

For example, using the regular helpers.bulk works:

```
bulk_data = []
header = data.columns
for i in range(len(data)):
    source_dict = {}
    row = data.iloc[i]
    for k in header:
        source_dict[k] = str(row[k])
    data_dict = {
        '_op_type': 'index',
        '_index': index_name,
        '_type': doc_type,
        '_source': source_dict
    }
    bulk_data.append(data_dict)

es.indices.create(index=index_name, body=settings, ignore=404)
helpers.bulk(client=es, actions=bulk_data)
es.indices.refresh()
es.count(index=index_name)

{'_shards': {'failed': 0, 'successful': 5, 'total': 5}, 'count': 13979}

```

But replacing it with helpers.parallel\_bulk doesn't seem to index anything:

```
bulk_data = []
header = data.columns
for i in range(len(data)):
    source_dict = {}
    row = data.iloc[i]
    for k in header:
        source_dict[k] = str(row[k])
    data_dict = {
        '_op_type': 'index',
        '_index': index_name,
        '_type': doc_type,
        '_source': source_dict
    }
    bulk_data.append(data_dict)
es.indices.create(index=index_name, body=settings, ignore=404)

helpers.parallel_bulk(client=es, actions=bulk_data, thread_count=4)
es.indices.refresh()
es.count(index=index_name)

{'_shards': {'failed': 0, 'successful': 5, 'total': 5}, 'count': 0}

```

Am I missing something? I'm on elasticsearch 2.1.1 with elasticsearch-py 2.1.0.

---

<div class="post-metadata">

**Author:** ![honzakral](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/honzakral/32/44958_2.png) [@honzakral](https://discuss.elastic.co/u/honzakral)\
**Post date:** [January 19, 2016, 4:55pm UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498/2 "2016-01-19T16:55:34Z")

</div>

Hi,

parallel bulk is a `generator`, meaning it is lazy and won't produce any results until you start consuming them. The proper way to use it is:

```
for success, info in parallel_bulk(...):
    if not success:
        print('A document failed:', info)

```

If you don't care about the results (which by default you don't have to since any error will cause an exception) you can use the `consume` function from itertools recipes ([https://docs.python.org/2/library/itertools.html#recipes](https://docs.python.org/2/library/itertools.html#recipes)):

```
from collections import deque
deque(parallel_bulk(...), maxlen=0)

```

Hope this helps.

---

<div class="post-metadata">

**Author:** ![honzakral](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/honzakral/32/44958_2.png) [@honzakral](https://discuss.elastic.co/u/honzakral)\
**Post date:** [January 19, 2016, 5:01pm UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498/3 "2016-01-19T17:01:18Z")

</div>

Btw the reason it is lazy is so that you never have to materialize a list with all the records, which can be potentially very expensive (for example when inserting data from a DB, long file or doing a reindex). It also means that you don't have to pass in a list, you can pass in a generator, thus avoiding creating a huge in-memory list yourself. In your example you could have a generator function:

```
def genereate_actions(data):
    for i in range(len(data)):
        source_dict = {}
        row = data.iloc[i]
        for k in header:
            source_dict[k] = str(row[k])
        yield {
            '_op_type': 'index',
            '_index': index_name,
            '_type': doc_type,
            '_source': source_dict
        }

```

and then call:

```
for success, info in parallel_bulk(es, genereate_actions(data), ...):
    if not success: print('Doc failed', info)

```

which will avoid the need to have all the documents present in memory at any given time. When working with larger dataset it can be a significant memory saving!

---

<div class="post-metadata">

**Author:** ![xamox](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xamox/32/5447_2.png) [@xamox](https://discuss.elastic.co/u/xamox)\
**Post date:** [April 15, 2016, 3:41pm UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498/4 "2016-04-15T15:41:47Z")

</div>

Thanks, I was having same issue. Not sure why this isn't called out in the docs.

---

<div class="post-metadata">

**Author:** ![oabio](https://avatars.discourse-cdn.com/v4/letter/o/ecccb3/32.png) [@oabio](https://discuss.elastic.co/u/oabio)\
**Post date:** [May 17, 2016, 12:38am UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498/5 "2016-05-17T00:38:40Z")

</div>

I am following your recommendation to index a very large file which takes hours. However, along the way, the memory foot print increases until it starts swapping. Internally, I am not saving anything and memory profile shows that it has something to do with parallel\_bulk, but I cannot get deeper than that. Am I missing something here? Does this makes any sense to you at all? I am using Python 2.7.5 with elasticsearch 2.3.0

Thank you for your help

---

<div class="post-metadata">

**Author:** ![minafarid](https://avatars.discourse-cdn.com/v4/letter/m/a4c791/32.png) [@minafarid](https://discuss.elastic.co/u/minafarid)\
**Post date:** [June 14, 2016, 8:16pm UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498/6 "2016-06-14T20:16:35Z")

</div>

This DEFINITELY needs to be in the documentation!!!!!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:43pm UTC](https://discuss.elastic.co/t/helpers-parallel-bulk-in-python-not-working/39498/7 "2017-07-05T22:43:51Z")

</div>


