# Remove first few lines of csv in logstash

**URL:** <https://discuss.elastic.co/t/remove-first-few-lines-of-csv-in-logstash/273359>\
**Category:** Logstash\
**Created:** [May 19, 2021, 4:46am UTC](https://discuss.elastic.co/t/remove-first-few-lines-of-csv-in-logstash/273359 "2021-05-19T04:46:35Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Osiris](https://avatars.discourse-cdn.com/v4/letter/o/839c29/32.png) [@Osiris](https://discuss.elastic.co/u/Osiris)\
**Post date:** [May 19, 2021, 4:46am UTC](https://discuss.elastic.co/t/remove-first-few-lines-of-csv-in-logstash/273359/1 "2021-05-19T04:46:35Z")

</div>

This is the input I am using for logstash.

```auto
    ItemId AssetId ItemName Comment
    11111 07 ABCDa XYZa
    11112 07 ABCDb XYZb
    11113 07 ABCDc XYZc
    11114 07 ABCDd XYZd
    11115 07 ABCDe XYZe
    11116 07 ABCDf XYZf
    11117 07 ABCDg XYZg
    Date Time rows columns
    19-05-2020 13:03 2 2
    19-05-2020 13:03 2 2
    19-05-2020 13:03 2 2
    19-05-2020 13:03 2 2
    19-05-2020 13:03 2 2

```

I need to remove first 8 lines from the csv and make the next line as column header and parse rest of lines as usual. Is there a way to do that in logstash?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 19, 2021, 4:42pm UTC](https://discuss.elastic.co/t/remove-first-few-lines-of-csv-in-logstash/273359/2 "2021-05-19T16:42:16Z")

</div>

I would use a multiline codec to combine each group of lines. Hopefully something like

```
pattern => "^\d" negate => false what => previous

```

would work, so that you get two events. The first being

```
"ItemId AssetId ItemName Comment\n11111 07 ABCDa XYZa\n11112 \n7 ABCDb XYZb\n11113 07 ABCDc XYZc\n11114 07 ABCDd XYZd\n11115 07 ABCDe XYZe\n11116 07 ABCDf XYZf\n11117 07 ABCDg XYZg"

```

You may need to clean up the field separators

```
mutate { gsub => ["message", "\s+", " "] }

```

Then tag each event

```
if [message] =~ /^Item/ {
    mutate { add_field => { "[@metadata][format]" => "format1" } }
} else {
    mutate { add_field => { "[@metadata][format]" => "format2" } }
}

```

Then use mutate+split to split [message] into an array of lines, then a split filter to convert the array into multiple events. You can then use csv filters for the two formats

```
if [@metadata][format] == "format1" {
    csv { separator => " " columns => ["ItemId", "AssetId", "ItemName", "Comment"] ... }
    if [ItemId] == "ItemId" { drop {} }
} else {
    csv { separator => " " columns => ["Date", "Time", "rows", "columns"] ... }
    if [Date] == "Date" { drop {} }
}

```

There are other ways of handling the headers (autodetect\_column\_names for example) but then you need pipeline.workers to be one and pipeline.ordered to be true.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 16, 2021, 4:42pm UTC](https://discuss.elastic.co/t/remove-first-few-lines-of-csv-in-logstash/273359/3 "2021-06-16T16:42:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
