# Reading the heading (1st line) of CSV file through logstash

**URL:** <https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930>\
**Category:** Logstash\
**Created:** [November 23, 2017, 5:53pm UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930 "2017-11-23T17:53:02Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![kiranilla](https://avatars.discourse-cdn.com/v4/letter/k/8e7dd6/32.png) [@kiranilla](https://discuss.elastic.co/u/kiranilla)\
**Post date:** [November 23, 2017, 5:53pm UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/1 "2017-11-23T17:53:02Z")

</div>

HI,

Eg: sample CSV file

```
Incident ID,Status,Resolved By,Resolution Breached,Resolution Month,Resolution Date,Resolution Date & Time,
IM02370568,Closed,GUNTURS2,FALSE,1,1/4/2016,1/4/2016 10:35
IM02370648,Closed,PRASADR6,FALSE,1,1/1/2016,1/1/2016 22:51

```

1. i am trying to read a the above csv file
2. The heading(1st line) need to be read and pass it to filter section of csv plugin
3. The data also need to be read and pass it further for parsing.

i would like to pass the heading dynamically to csv plugin of logstash, instead of hardcoding the headings.  
Could you please help me out in resolving this issue or at-least an alternative for this problem.

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [November 23, 2017, 6:53pm UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/2 "2017-11-23T18:53:28Z")

</div>

The CSV filter is stateless, meaning that it handles each event without knowing anything about any previous events. Also the config is parsed and loaded at LS start before any events have been processed.

There is no way at the moment to do this very dynamically.

How many different CSV structures do you need to parse?  
Does the filename and/or path contain a clue to the type of CSV structure?

---

<div class="post-metadata">

**Author:** ![Leandro\_Sampaio](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandro_sampaio/32/18409_2.png) [@Leandro\_Sampaio](https://discuss.elastic.co/u/Leandro_Sampaio)\
**Post date:** [November 23, 2017, 7:13pm UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/3 "2017-11-23T19:13:34Z")

</div>

Maybe you could use a grok condition only the first line and keep it in session... It's not a good practice, but it could resolve your problem...

---

<div class="post-metadata">

**Author:** ![BinaryMonkey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/binarymonkey/32/24625_2.png) [@BinaryMonkey](https://discuss.elastic.co/u/BinaryMonkey)\
**Post date:** [November 24, 2017, 12:46am UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/4 "2017-11-24T00:46:55Z")

</div>

filter {  
csv {  
separator =\> ","  
autodetect\_column\_names =\> true  
autogenerate\_column\_names =\> true  
}  
}

---

<div class="post-metadata">

**Author:** ![kiranilla](https://avatars.discourse-cdn.com/v4/letter/k/8e7dd6/32.png) [@kiranilla](https://discuss.elastic.co/u/kiranilla)\
**Post date:** [November 24, 2017, 6:01am UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/5 "2017-11-24T06:01:18Z")

</div>

Thanks @guyboertje for you quick response..

As per your queries below:

How many different CSV structures do you need to parse?

_\> There is no limit for CSV structures, at present we have done for 2 but in future we are expecting more, so would like to generalize for all the structures._

Does the filename and/or path contain a clue to the type of CSV structure?

_\> Yes, filepath can be in one location(we can hard code the path) but filename varies._

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [November 24, 2017, 11:34am UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/6 "2017-11-24T11:34:15Z")

</div>

As @BinaryMonkey has suggested you can **try** the setting the first or both of:

```auto
autodetect_column_names => true
autogenerate_column_names => true

```

See [https://github.com/logstash-plugins/logstash-filter-csv/blob/master/lib/logstash/filters/csv.rb#L126](https://github.com/logstash-plugins/logstash-filter-csv/blob/master/lib/logstash/filters/csv.rb#L126)  
The line of code above will capture the first event seen by the plugin (on all LS restarts) as the column names.

The real pitfall with this is:

1. Only the very first line of the very first file will be the columns for all other files - because the CSV filter does not keep a map of file\_path -\> columns internally.
2. If you have to restart Logstash while it is half way though a file then the first line of that file will not be re-read and the columns will become some arbitrary set of values.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [November 24, 2017, 11:47am UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/7 "2017-11-24T11:47:13Z")

</div>

> [@kiranilla](#):
>
> Yes, filepath can be in one location(we can hard code the path) but filename varies.

So if you have two different CSV structures, say `structure-1` with columns `"a","b"` and `structure-2` with `"c","d"`  
then...  
**Can you put all files with structure-1 into a sub-folder called `structure-1` and so on?**  
If so then you can use regex if conditionals to separate the files so that they flow through their own csv filter. You still have to take into account the second pitfall of my previous post unless you hard code the columns for each structure.

Logstash do not have an automatic way of doing this, because by design we strive for statelessness and the file input (and Filebeat) is designed to resume reading from where reached the previous time.

---

<div class="post-metadata">

**Author:** ![kiranilla](https://avatars.discourse-cdn.com/v4/letter/k/8e7dd6/32.png) [@kiranilla](https://discuss.elastic.co/u/kiranilla)\
**Post date:** [November 26, 2017, 6:00pm UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/8 "2017-11-26T18:00:03Z")

</div>

Thanks @guyboertje for your response.

As suggested, i will check the feasibility and go ahead with implementation.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 24, 2017, 6:00pm UTC](https://discuss.elastic.co/t/reading-the-heading-1st-line-of-csv-file-through-logstash/108930/9 "2017-12-24T18:00:08Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
