# Csv parse taking too long using python API

**URL:** <https://discuss.elastic.co/t/csv-parse-taking-too-long-using-python-api/106907>\
**Category:** Elasticsearch\
**Created:** [November 8, 2017, 5:13pm UTC](https://discuss.elastic.co/t/csv-parse-taking-too-long-using-python-api/106907 "2017-11-08T17:13:31Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![henriqueluz](https://avatars.discourse-cdn.com/v4/letter/h/85e7bf/32.png) [@henriqueluz](https://discuss.elastic.co/u/henriqueluz)\
**Post date:** [November 8, 2017, 5:13pm UTC](https://discuss.elastic.co/t/csv-parse-taking-too-long-using-python-api/106907/1 "2017-11-08T17:13:32Z")

</div>

Hello! That's my first post here in elastic discussion area, sorry if any mistake were made.

I'm kinda new using ELK Stack, I'm trying to parse a .CSV file to ElasticSearch using the Python API. The thing is, it's taking way too long to parse just a few logs (312seconds to parse 30000 logs). I'm using ElasticSearch 5-6-3 running on Ubuntu 16.04, 6GB RAM  
The main idea is to convert each row into a json and then parse it to Elastic, the code:

```
import time
import json
import pandas as pd
from elasticsearch import Elasticsearch

class Storage(object):
    
    def __init__ (self,user,password):
        
        self.user = user
        self.password = password
        self.es = Elasticsearch(http_auth=(user, password))
        
    def get_info(self, log=False):
        
        info = self.es.info()
        if info:
            print(json.dumps(info, indent=3))
            return info
    
    def index(self, index, doc_type, _id, json):
    
        status = self.es.index(index=index, doc_type=doc_type, id = _id, body=json)
        return status

        
if __name__ == ' __main__':
    
    st = Storage('elastic','changeme')
    print(st.get_info())
    df = pd.read_csv(file_to_parse, low_memory=False)
    aux = df.to_dict('records')
    
    index = 1
    begin = time.time()
    for register in aux:
        
        reg = json.dumps(register)
        res = st.index('someindex','log',index,reg)
        index += 1
        
    print("Finished", time.time() - begin)

What could possibly be improved here?
```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 8, 2017, 5:19pm UTC](https://discuss.elastic.co/t/csv-parse-taking-too-long-using-python-api/106907/2 "2017-11-08T17:19:22Z")

</div>

Use the [bulk API](https://elasticsearch-py.readthedocs.io/en/master/api.html#elasticsearch.Elasticsearch.bulk) to send multiple documents per indexing request ([docs](https://www.elastic.co/guide/en/elasticsearch/reference/5.6/docs-bulk.html)). This is much more efficient than indexing each document individually.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 6, 2017, 5:19pm UTC](https://discuss.elastic.co/t/csv-parse-taking-too-long-using-python-api/106907/3 "2017-12-06T17:19:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
