# Parsing csv file with dynamic headers/columns

**URL:** <https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354>\
**Category:** Logstash\
**Created:** [May 11, 2017, 9:01am UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354 "2017-05-11T09:01:28Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![sirsyedian](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@sirsyedian](https://discuss.elastic.co/u/sirsyedian)\
**Post date:** [May 11, 2017, 9:01am UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/1 "2017-05-11T09:01:28Z")

</div>

Hi,  
I need to parse a csv file which contains some dynamic headers/Columns. (lets say)  
Header format-1: Column-1, Column-2, Collumn-3, X  
Header format-2: Column-1, Column-2, Collumn-3, Y

These files are generated every 5 minutes and some have header format-1 while others have header format-2. Without opening the files, we cant say whether they would have Header format-1 or header format-2 (ie nothing in their name suggests if they are of either type).

The header is part of the csv file (ie first line), so I have tried some conditional statements to identify which format is the file and assign path to a variable that i can use to ascertain which file has which format (something like below)

filter {  
if [message] =~ /^Column-1, Column-2, Collumn-3, X/  
{ mutate {  
add\_field =\> { "File-Type-1" =\> "%{[path]}" }  
}

However, since each line (within the csv) is parsed at a time, the above condition only picks up the header line only and doesnt work on the actual data.

Is there a way in logstash i can accomplish this task?

Thanks  
sirsyedian

---

<div class="post-metadata">

**Author:** ![sirsyedian](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@sirsyedian](https://discuss.elastic.co/u/sirsyedian)\
**Post date:** [May 14, 2017, 12:56am UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/2 "2017-05-14T00:56:43Z")

</div>

Has anyone got luck in managing csv files with dynamic headers in logstash?  
Just want to know if this is something supported/manageable in logstash or we need to start exploring other options.

Thanks  
sirsyedian

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 14, 2017, 8:12pm UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/3 "2017-05-14T20:12:22Z")

</div>

> [@sirsyedian](#):
>
> Is there a way in logstash i can accomplish this task?

Unfortunately not ☹

---

<div class="post-metadata">

**Author:** ![sirsyedian](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@sirsyedian](https://discuss.elastic.co/u/sirsyedian)\
**Post date:** [May 15, 2017, 8:06am UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/4 "2017-05-15T08:06:50Z")

</div>

Thanks Mark,  
Are you (or anyone else) aware of other tools capable of handling such tasks?

Regards  
sirsyedian

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 15, 2017, 8:29am UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/5 "2017-05-15T08:29:22Z")

</div>

Does the file path/name give you a clue as to the potential CSV layout of the contents?

---

<div class="post-metadata">

**Author:** ![sirsyedian](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@sirsyedian](https://discuss.elastic.co/u/sirsyedian)\
**Post date:** [May 15, 2017, 5:01pm UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/6 "2017-05-15T17:01:35Z")

</div>

Unfortunately not. Only the file header within it would tell us how to interpret its data.

Isnt there a way we can interpret the header and based on its format, interpret rest of the file?

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 16, 2017, 9:23am UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/7 "2017-05-16T09:23:09Z")

</div>

We don't have an out of the box solution for this.

I guess you could write a script that looked at the first line of data in a file in a 'staging' folder then wrote the file to a 'ingest' folder with a new filename that reflected which pattern the CSV contents what dumped in.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 13, 2017, 9:29am UTC](https://discuss.elastic.co/t/parsing-csv-file-with-dynamic-headers-columns/85354/8 "2017-06-13T09:29:01Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
