# Multiple CSV files and Multiple Header data \[identifiers\]

**URL:** <https://discuss.elastic.co/t/multiple-csv-files-and-multiple-header-data-identifiers/157077>\
**Category:** Logstash\
**Created:** [November 16, 2018, 2:02pm UTC](https://discuss.elastic.co/t/multiple-csv-files-and-multiple-header-data-identifiers/157077 "2018-11-16T14:02:06Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![wheelq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wheelq/32/35893_2.png) [@wheelq](https://discuss.elastic.co/u/wheelq)\
**Post date:** [November 16, 2018, 2:02pm UTC](https://discuss.elastic.co/t/multiple-csv-files-and-multiple-header-data-identifiers/157077/1 "2018-11-16T14:02:06Z")

</div>

I have read @theuntergeek responses in the [Parse 1st Line of Multiple CSV files and set as Columns thread](https://discuss.elastic.co/t/parse-1st-line-of-multiple-csv-files-and-set-as-columns/2246/4)

But I didn't fully understand this comment:

> csv\_type\_1 identifies a stream (by file type, or whatever you identify it with)

Can I have an example please? Where do I define csv\_type\_1?

---

<div class="post-metadata">

**Author:** ![theuntergeek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/theuntergeek/32/44961_2.png) [@theuntergeek](https://discuss.elastic.co/u/theuntergeek)\
**Post date:** [November 19, 2018, 7:59pm UTC](https://discuss.elastic.co/t/multiple-csv-files-and-multiple-header-data-identifiers/157077/2 "2018-11-19T19:59:47Z")

</div>

You're referring to this post:

> [@Parse 1st Line of Multiple CSV files and set as Columns](https://discuss.elastic.co/t/parse-1st-line-of-multiple-csv-files-and-set-as-columns/2246/4):
>
> The content is always split into lines, but without a conditional, the csv filter is totally unaware of which line of a file it is receiving. This is why a conditionals are essential, and knowing what the possible column types are.
> 
> ```auto
> if [csv_type_1] {
> if [message] =~ /headerpattern/ {
> csv { ... }
> }
> }
> 
> ```
> 
> csv\_type\_1 identifies a stream (by file type, or whatever you identify it with), headerpattern will be the way you can tell the first line is a header and not data, and then you apply the _known_ csv column match to the data.
> 
> Logstash 2.x will simplify this somewhat as we are designing it to allow for multiple pipelines. A (hypothetical) CSV pipeline would allow for ingesting a single file, and using the first line to define column/field names. In the meanwhile, unfortunately there is nothing in Logstash which will do what you are asking.

Included here for quick access.

The idea behind `csv_type_1` is arbitrary. It just needs to be a way you can concretely identify the source of the data to differentiate it from other sources of data. If you only have a single source of data, this conditional is unnecessary.

---

<div class="post-metadata">

**Author:** ![wheelq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wheelq/32/35893_2.png) [@wheelq](https://discuss.elastic.co/u/wheelq)\
**Post date:** [November 19, 2018, 8:59pm UTC](https://discuss.elastic.co/t/multiple-csv-files-and-multiple-header-data-identifiers/157077/3 "2018-11-19T20:59:55Z")

</div>

Thanks, I understand it, but I don't know how can I tag specific stream as csv\_type\_1 or anything else. Maybe I'm overthinking it 😉

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [November 19, 2018, 10:04pm UTC](https://discuss.elastic.co/t/multiple-csv-files-and-multiple-header-data-identifiers/157077/4 "2018-11-19T22:04:47Z")

</div>

Can you post your input so we can see how you're ingesting data?

Basically, @theuntergeek is running two logic checks, one inside the other. Basically, it says:

```
IF the field, `csv_type_1` exists in the event {
  IF the field, `message` matches regular expression pattern `headerpattern` {
    perform csv parsing { ... }
  }
}

```

Both `csv_type_1` and `message` are fields inside the event. `Message` will always exist because thats where logstash sticks the raw data it receives. `Csv_type_1` is a field that he just came up with as an example or exists in his own dataset. Unless you are pulling logs from the same type of device, you are going to use a different field to qualify the statement.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 17, 2018, 10:04pm UTC](https://discuss.elastic.co/t/multiple-csv-files-and-multiple-header-data-identifiers/157077/5 "2018-12-17T22:04:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
