# Bulk export of several directories of csv files to elasticsearch

**URL:** <https://discuss.elastic.co/t/bulk-export-of-several-directories-of-csv-files-to-elasticsearch/78474>\
**Category:** Logstash\
**Created:** [March 14, 2017, 9:14am UTC](https://discuss.elastic.co/t/bulk-export-of-several-directories-of-csv-files-to-elasticsearch/78474 "2017-03-14T09:14:37Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Natheer\_Alabsi](https://avatars.discourse-cdn.com/v4/letter/n/e0b2c6/32.png) [@Natheer\_Alabsi](https://discuss.elastic.co/u/Natheer_Alabsi)\
**Post date:** [March 14, 2017, 9:14am UTC](https://discuss.elastic.co/t/bulk-export-of-several-directories-of-csv-files-to-elasticsearch/78474/1 "2017-03-14T09:14:37Z")

</div>

I have several directories and each has several subdirectories and each has csv files within them. I would like to export all these csv files inside all directories to elasticsearch using logstash and automatically take column names. Note that different csv files have different column names. There are about 7 types of csv files in these directories and each type has the same column names.

Please suggest me what to do in such case.

input {  
file {  
path =\> "/dir/dir\*/_/_/\*.csv"  
start\_position =\> "beginning"  
}  
}

filter {  
if ([message] =~ " ") {  
drop { }  
} else {  
csv { }  
}  
}

output {  
elasticsearch {  
hosts =\> ["[http://10.0.0.4:9200](http://10.0.0.4:9200)"]  
}  
stdout{}  
}

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [March 17, 2017, 6:31am UTC](https://discuss.elastic.co/t/bulk-export-of-several-directories-of-csv-files-to-elasticsearch/78474/2 "2017-03-17T06:31:38Z")

</div>

The csv filter has an `autodetect_column_names` option that probably does what you're looking for.

---

<div class="post-metadata">

**Author:** ![Natheer\_Alabsi](https://avatars.discourse-cdn.com/v4/letter/n/e0b2c6/32.png) [@Natheer\_Alabsi](https://discuss.elastic.co/u/Natheer_Alabsi)\
**Post date:** [March 17, 2017, 7:47am UTC](https://discuss.elastic.co/t/bulk-export-of-several-directories-of-csv-files-to-elasticsearch/78474/3 "2017-03-17T07:47:09Z")

</div>

> [@magnusbaeck](#):
>
> autodetect\_column\_names

But nothing in the documentation about how to use it.  
Actually the columns are in the second row, not in the first row. What should I do in such case.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [March 17, 2017, 8:19am UTC](https://discuss.elastic.co/t/bulk-export-of-several-directories-of-csv-files-to-elasticsearch/78474/4 "2017-03-17T08:19:36Z")

</div>

> But nothing in the documentation about how to use it.

Well, there's this: [Csv filter plugin | Logstash Reference [8.11] | Elastic](https://www.elastic.co/guide/en/logstash/current/plugins-filters-csv.html#plugins-filters-csv-autogenerate_column_names)

> Actually the columns are in the second row, not in the first row. What should I do in such case.

That's probably not supported.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 14, 2017, 8:19am UTC](https://discuss.elastic.co/t/bulk-export-of-several-directories-of-csv-files-to-elasticsearch/78474/5 "2017-04-14T08:19:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
