# Read and index multiple CSV files

**URL:** <https://discuss.elastic.co/t/read-and-index-multiple-csv-files/235615>\
**Category:** Elasticsearch\
**Created:** [June 3, 2020, 6:45pm UTC](https://discuss.elastic.co/t/read-and-index-multiple-csv-files/235615 "2020-06-03T18:45:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![manideep](https://avatars.discourse-cdn.com/v4/letter/m/6a8cbe/32.png) [@manideep](https://discuss.elastic.co/u/manideep)\
**Post date:** [June 3, 2020, 6:45pm UTC](https://discuss.elastic.co/t/read-and-index-multiple-csv-files/235615/1 "2020-06-03T18:45:39Z")

</div>

Hi Community,

I am working on a task to read multiple CSV files from path and index them in Elasticsearch. all csv files are independent of each other.  
I am using python. I appreciate if there are any alternate methods.

example:

csv files:

- test1.csv
- test2.csv
- test3.csv
- .........

indices:

- test1
- test2
- test3
- .......

below is my code. my code seems working for one file and not working for multiple csv files.  
the issue is at the helpers.bulk statement. I commented that statement and verified looping through all csv files.

also, because of the nulls values, i see below error which is fine. all nulls are ignored and indexed for now. But is this really causing issue to next all CSVs?

```auto
'error': {'type': 'mapper_parsing_exception', 'reason': 'failed to parse', 'caused_by': {'type': 'json_parse_exception', 'reason': "Non-standard token 'NaN': enable JsonParser.Feature.ALLOW_NON_NUMERIC_NUMBERS to allow\n at [Source: org.elasticsearch.common.bytes.BytesReference$MarkSupportingStreamInputWrapper@69daf1f8; line: 1, column: 33]"}}

```

```auto
# import Elasticsearch module

from elasticsearch import Elasticsearch

from elasticsearch import helpers

import pandas as pd

import glob

# read csv file

path = "C:/Users/shanuma3/Desktop/CSV/"

#file = "bigmart_data1.csv"

files = glob.glob(path + "*.csv")

rows = 10000

# create connection

es = Elasticsearch([{'host':'localhost','port':9200}])

for file in files:

    df = pd.read_csv(file, nrows=rows)

    index_name = file.split("\\")[1][:-4]   

    documents = df.to_dict(orient='records')

    print("Index created: " + index_name)

    es.indices.create(index = index_name)

    print("Indexing Start: " + index_name)

    #print(documents)

    helpers.bulk(es, documents, index = index_name, doc_type='_doc', raise_on_error=True)

    print("Index finished:" + index_name)

```

---

<div class="post-metadata">

**Author:** ![manideep](https://avatars.discourse-cdn.com/v4/letter/m/6a8cbe/32.png) [@manideep](https://discuss.elastic.co/u/manideep)\
**Post date:** [June 4, 2020, 2:18pm UTC](https://discuss.elastic.co/t/read-and-index-multiple-csv-files/235615/2 "2020-06-04T14:18:50Z")

</div>

Hi All,

I think I figured out the solution. Since I am using pandas, I am converting 'Nan' to nulls using the below piece of code.

```auto
df = df.where(pd.notnull(df), None)

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 2, 2020, 2:19pm UTC](https://discuss.elastic.co/t/read-and-index-multiple-csv-files/235615/3 "2020-07-02T14:19:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
