# Index from pandas to Elastic Search Using BULK and Parallel BULK

**URL:** <https://discuss.elastic.co/t/index-from-pandas-to-elastic-search-using-bulk-and-parallel-bulk/176501>\
**Category:** Elasticsearch\
**Created:** [April 11, 2019, 9:16pm UTC](https://discuss.elastic.co/t/index-from-pandas-to-elastic-search-using-bulk-and-parallel-bulk/176501 "2019-04-11T21:16:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![tanujsawant123](https://avatars.discourse-cdn.com/v4/letter/t/85f322/32.png) [@tanujsawant123](https://discuss.elastic.co/u/tanujsawant123)\
**Post date:** [April 11, 2019, 9:16pm UTC](https://discuss.elastic.co/t/index-from-pandas-to-elastic-search-using-bulk-and-parallel-bulk/176501/1 "2019-04-11T21:16:28Z")

</div>

BELOW is my code and mapping to create a new index and load data from pandas to elastic.

I have around 400K records with the potential to increase to 2 Millions.

from elasticsearch.helpers import bulk , streaming\_bulk, parallel\_bulk  
config = {  
'host': 'xxx.xx.liv','port':xxxx  
}  
es = Elasticsearch([config,], )

MAPPING: --  
request\_body = {  
"settings" : {  
"number\_of\_shards": 5,  
"number\_of\_replicas": 1  
},

```
'mappings': {
    'Product': {
        'properties': {
            'ipc': {'type': 'keyword'},
            'O_D_P': {'type': 'float'},
            'O_D_P_P': {'type': 'float'},
            'A_P': {'type': 'float'},
            'O': {'type': 'integer'},
            'D_P': {'type': 'float'},
            'D_P_P': {'type': 'float'},
            'L_I': {'type': 'float'},
            'H_I': {'type': 'float'}
        }}}

```

}  
print("creating 'example\_index' index...")  
res = es.indices.delete(index = 'ipc')  
es.indices.create(index = 'ipc', body = request\_body)

DATA:

bulk\_data =

for index, row in DataFrame.iterrows():  
data\_dict = {}  
for i in range(len(row)):  
data\_dict[DataFrame.columns[i]] = row[i]  
op\_dict = {  
"index": {  
"\_index": 'ipc',  
"\_type": 'Product',  
"\_id": data\_dict['ipc']  
}  
}  
bulk\_data.append(op\_dict)  
bulk\_data.append(data\_dict)

res = es.bulk(index = 'ipc',body = bulk\_data,refresh=True,request\_timeout=360000)  
es.search(body={"query": {"match\_all": {}}}, index = 'ipc')  
es.indices.get\_mapping(index = 'ipc')

I am getting Broken Pipe error. Is there a way in which we can load data by chunk?

I tried using Streaming Bulk :

res = streaming\_bulk(client=es, actions=bulk\_data,index ='ipc', chunk\_size=1, max\_retries=5,  
initial\_backoff=2, max\_backoff=600, request\_timeout=3600,refresh=True, yield\_ok=True)  
for ok, response in res:  
print(ok, response)

es.search(body={"query": {"match\_all": {}}}, index = 'ipc')  
es.indices.get\_mapping(index = 'ipc')

ERROR:  
RequestError: RequestError(400, 'action\_request\_validation\_exception', 'Validation Failed: 1: type is missing;'

How can I insert in chunks?

---

<div class="post-metadata">

**Author:** ![astievet](https://avatars.discourse-cdn.com/v4/letter/a/5e9695/32.png) [@astievet](https://discuss.elastic.co/u/astievet)\
**Post date:** [April 12, 2019, 9:24am UTC](https://discuss.elastic.co/t/index-from-pandas-to-elastic-search-using-bulk-and-parallel-bulk/176501/2 "2019-04-12T09:24:09Z")

</div>

I don't know the python api, but could it be a missing value in your call? Try to add a type to your parameters, e.g.:  
res = es.bulk(index = 'ipc', **doc\_type** = "\_doc", body = bulk\_data,refresh=True,request\_timeout=360000)

---

<div class="post-metadata">

**Author:** ![tanujsawant123](https://avatars.discourse-cdn.com/v4/letter/t/85f322/32.png) [@tanujsawant123](https://discuss.elastic.co/u/tanujsawant123)\
**Post date:** [April 12, 2019, 4:54pm UTC](https://discuss.elastic.co/t/index-from-pandas-to-elastic-search-using-bulk-and-parallel-bulk/176501/3 "2019-04-12T16:54:03Z")

</div>

Same error:

ConnectionError: ConnectionError(('Connection aborted.', BrokenPipeError(32, 'Broken pipe'))) caused by: ProtocolError(('Connection aborted.', BrokenPipeError(32, 'Broken pipe')))

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 10, 2019, 4:54pm UTC](https://discuss.elastic.co/t/index-from-pandas-to-elastic-search-using-bulk-and-parallel-bulk/176501/4 "2019-05-10T16:54:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
