# Csv plugin cooperating with multiplying pattern

**URL:** <https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865>\
**Category:** Logstash\
**Created:** [July 18, 2022, 9:34am UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865 "2022-07-18T09:34:25Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 18, 2022, 9:34am UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/1 "2022-07-18T09:34:25Z")

</div>

Hi  
@Badger I need to continue below topic:

> [@Conditional processing in txt file](https://discuss.elastic.co/t/conditional-processing-in-txt-file/306654/5):
>
> See [this](https://discuss.elastic.co/t/using-ruby-to-count-the-number-of-fields-in-a-message-then-apply-csv-depending-on-the-count/177880/2) thread.

referring to above I have a question how I can mark in this code pattern, also count of records is not regular (once it's more once it's less)

a place with different kind of pattern/csv?  
the second case that I need to grab a timestamp from this 1'st line of such file  
`# snapshot,65767220,20220601044503`  
As I've tried under one of method and met the issue because when all messages starting with `#` will be tried to have the same format (a least there will be an issue for the multiline input)

an example:

```auto

# snapshot,65767220,20220601044503
# Network Elements

0097,s,n,2719,,s,,,,3,,p,

C313,s,n,4767,,s,,y,,,,,

# DN Blocks

224135896,224135897,,,,,,,,,,,
224135896,224135897,,,,,,,,,,,

# numbers;

00000801163158,0,n,n,y
00000801163158,0,n,n,y

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 18, 2022, 5:52pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/2 "2022-07-18T17:52:30Z")

</div>

I would consider consuming the file with a multiline codec that collects all the lines for each '# Type', then tag them with the type and use a split filter to break each line into its own event

Suppose we have a file

```auto
# Snapshot
# ABC
a,b,c
d,e,f
# Type2
foo,1,2,3
bar,4,5,6
# DN Blocks
224135896,224135897,,,,,,,,,,,
224135896,224135897,,,,,,,,,,,

```

and we use this input/filter configuration

```
input {
    file {
        path => "/home/user/foo.txt"
        sincedb_path => "/dev/null"
        start_position => beginning
        codec => multiline { pattern => "^#" negate => true what => previous auto_flush_interval => 2 multiline_tag => "" }
    }
}
filter {
    mutate { remove_field => ["[event]", "log" ] }
    if "# Snap" in [message] {
        mutate { add_field => { "eventType" => "Header" } }
    } else if "# ABC" in [message] {
        mutate { add_field => { "eventType" => "ABC" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["c1", "c2", "c3"] }
        }
    } else if "# Type2" in [message] {
        mutate { add_field => { "eventType" => "Type2" } }
        split { field => "message" }
    } else {
        mutate { add_field => { "eventType" => "Unrecognized" } }
    }
}

```

We will get

```auto
{
            "c3" => "c",
            "c1" => "a",
       "message" => "a,b,c",
            "c2" => "b",
    "@timestamp" => 2022-07-18T17:47:47.452450Z,
     "eventType" => "ABC"
}
{
            "c3" => "f",
            "c1" => "d",
       "message" => "d,e,f",
            "c2" => "e",
    "@timestamp" => 2022-07-18T17:47:47.452450Z,
     "eventType" => "ABC"
}

```

as well as

```auto
{
       "message" => "# DN Blocks\n224135896,224135897,,,,,,,,,,,\n224135896,224135897,,,,,,,,,,,",
    "@timestamp" => 2022-07-18T17:47:49.939492Z,
     "eventType" => "Unrecognized"
}

```

Obviously you will have to add either csv or dissect for each line type.

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 19, 2022, 8:14am UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/3 "2022-07-19T08:14:05Z")

</div>

Many thanks @Badger in the meantime I'm trying to get out the timestamp from the first line, but something goes wrong

```auto
filter {
    mutate { remove_field => ["[event]", "log" ] }
    if "# snapshot" in [message] {
        dissect {
            mapping => {
                "[message]" => "# %{activity},%{val},%{time}"
            }
            remove_field => ["[message]"]
        }
        date {
                match => ["time", "yyyyMMddHHmmss"]
                timezone => "Europe/Paris"
            }

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 19, 2022, 3:47pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/4 "2022-07-19T15:47:07Z")

</div>

> [@INS](#):
>
> something goes wrong

What goes wrong? What is the value of [time], what is the value of [@timestamp], is there a parse failure tag on the event?

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 19, 2022, 6:46pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/5 "2022-07-19T18:46:50Z")

</div>

Hmm, I don't know how to use this timestamps for all of event's from the first one of row  
as You see it was output in the first messages but I need the same timestamp for the others

```auto
{
      "activity" => "snapshot",
          "path" => "/opt/data/input/test_npdb.txt",
      "@version" => "1",
          "time" => "20220601044503",
          "host" => "0.0.0.0",
    "@timestamp" => 2022-06-01T02:45:03Z,
           "val" => "65767220"
}
{
          "NSDN" => nil,
            "PT" => "1",
      "@version" => "1",
          "CGBL" => nil,
            "SP" => nil,
            "DN" => "C0108",
           "VMS" => nil,
          "CDBL" => "-2",
            "RN" => nil,
          "path" => "/opt/data/input/test_npdb.txt",
           "ASD" => nil,
          "IMSI" => nil,
           "GRN" => nil,
            "ST" => nil,
          "host" => "0.0.0.0",
     "eventType" => "DNs",
    "@timestamp" => 2022-07-19T18:43:33.343119Z,
       "message" => "C0108,,1,,,,,,,,,-2"
}

```

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 19, 2022, 6:55pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/6 "2022-07-19T18:55:27Z")

</div>

I've tried some code as below the got the same results as above

```auto
filter {
    mutate { remove_field => ["[event]", "log" ] }
    if "# snapshot" in [message] {
        dissect {
            mapping => {
                "[message]" => "# %{activity},%{val},%{[@metadata][timestamp]}"
            }
            remove_field => ["[message]"]
        }
        date {
                match => ["[@metadata][timestamp]", "yyyyMMddHHmmss"]
                timezone => "Europe/Paris"
            }

         # mutate { add_field => { "eventType" => "Header" } }
    } else if "# Network Elements" in [message] {
        mutate { add_field => { "eventType" => "Network Elements" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["ID","Type","PCType","PC","GC","RI","SSN","CCGT","NTT","NNAI","NNP","DA","SR"] }
        date {
                match => ["[@metadata][timestamp]", "yyyyMMddHHmmss"]
                timezone => "Europe/Paris"
            }

        }

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 19, 2022, 7:51pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/7 "2022-07-19T19:51:52Z")

</div>

> [@INS](#):
>
> I don't know how to use this timestamps for all of event's from the first one of row

You could save in a ruby filter and add it to non-snapshot lines. An example is [here](https://discuss.elastic.co/t/how-to-handle-metadata-in-file-headers/167544/2). You will need pipeline.workers 1 and pipeline.ordered true. This kind of ruby solution tends to be fragile.

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 19, 2022, 7:57pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/8 "2022-07-19T19:57:04Z")

</div>

> [@Badger](#):
>
> ou could save in a ruby filter and add it to no

great I will try to do it

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 19, 2022, 8:30pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/9 "2022-07-19T20:30:05Z")

</div>

```auto
{
          "host" => "0.0.0.0",
          "path" => "/opt/data/input/test2.txt",
      "activity" => "snapshot",
           "val" => "65767220",
      "@version" => "1",
          "time" => "20220601044503",
    "@timestamp" => 2022-06-01T02:45:03Z
}
{
          "host" => "0.0.0.0",
       "message" => "0097,s,n,2719,,s,,,,3,,p,",
          "CCGT" => nil,
      "metadata" => {},
           "NTT" => nil,
           "NNP" => nil,
      "@version" => "1",
            "DA" => "p",
       "SRFIMSI" => nil,
            "PC" => "2719",
          "tags" => [
        [0] "_dateparsefailure"
    ],
        "PCType" => "n",
          "path" => "/opt/data/input/test2.txt",
           "SSN" => nil,
          "NNAI" => "3",
     "eventType" => "Network Elements",
    "@timestamp" => 2022-07-19T21:10:17.456544Z,
          "Type" => "s",
            "RI" => "s",
            "GC" => nil,
            "ID" => "0097"
}

```

all of the pipeline code

```auto
input {
    file {
        mode => read
        path => "/opt/data/input/test2.txt"
        sincedb_path => "/dev/null"
        start_position => beginning
        file_completed_action => "log"
        file_completed_log_path => "/opt/data/logstash_files/fin_eir.log"
        codec => multiline { pattern => "^#" negate => true what => previous auto_flush_interval => 2 multiline_tag => "" }
    }
}
filter {
    mutate { remove_field => ["[event]", "log" ] }
    if "# snapshot" in [message] {
         dissect {
            mapping => {
                "[message]" => "# %{activity},%{val},%{time}"
            }
            remove_field => ["[message]"]
        }
        date {
                match => ["time", "yyyyMMddHHmmss"]
                timezone => "Europe/Paris"
            }
        ruby {
        init => '
                @@collectingMetadata = false
            '
            code => '
                unless @@collectingMetadata
                    @@metadata = {}
                    @@collectingMetadata = true
                end
                @@metadata[event.get("time")]
            '
        }

         # mutate { add_field => { "eventType" => "Header" } }
    } else if "# Network Elements" in [message] {
        mutate { add_field => { "eventType" => "Network Elements" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["ID","Type","PCType","PC","GC","RI","SSN","CCGT","NTT","NNAI","NNP","DA","SRF"] }
        }

 ruby {
                code => '
                    event.set("metadata", @@metadata)
                '
                }
 date {
                match => ["metadata", "yyyyMMddHHmmss"]
                timezone => "Europe/Paris"
            }

    } else if "# DNs" in [message] {
        mutate { add_field => { "eventType" => "DNs" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["DN","IMSI","PT","SP","RN","VMS","GRN","ASD","ST","NSDN","CGBL","CDBL"] }

        }
    } else if "# DN Blocks" in [message] {
        mutate { add_field => { "eventType" => "DN Block" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["BDN","EDN","PT","SP","RN","VMS","GRN","ASD","ST","NSDN","CGBL","CDBL"] }
        }
    } else {
        mutate { add_field => {"eventType" => "numbers"}}
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["IMEI","SVN","WHITE","GRAY","BLACK"] }
        }
    }
}

output {
    stdout { codec => rubydebug }
}

```

@Badger You can try on this sample data [test2.txt]

```auto
# snapshot,65767220,20220601044503
# Network Elements

0097,s,n,2719,,s,,,,3,,p,

C313,s,n,4767,,s,,y,,,,,

# DN Blocks

224135896,224135897,,,,,,,,,,,
224135896,224135897,,,,,,,,,,,

# numbers;

00000801163158,0,n,n,y
00000801163158,0,n,n,y

```

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 19, 2022, 9:13pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/10 "2022-07-19T21:13:01Z")

</div>

plus pipelines.yml

```auto
- pipeline.id: test
  path.config: "/usr/share/logstash/pipeline/pipeline_test.yml"
  pipeline.workers: 1
  pipeline.batch.size: 2
  pipeline.batch.delay: 50
  pipeline.ordered: true

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 19, 2022, 10:39pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/11 "2022-07-19T22:39:46Z")

</div>

I would use

```
ruby { code => '@@metadata = event.get("@timestamp")' }

```

for the snapshot lines, and

```
ruby { code => 'event.set("@timestamp", @@metadata)' }

```

for the others.

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 20, 2022, 8:58am UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/12 "2022-07-20T08:58:41Z")

</div>

> [@Badger](#):
>
> `ruby { code => 'event.set("@timestamp", @@metadata)' }`

thanks it works as well 😉

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 21, 2022, 8:16am UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/13 "2022-07-21T08:16:42Z")

</div>

When I'm processing file under ~800Mb it takes to long time through 1 worker

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 21, 2022, 12:01pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/14 "2022-07-21T12:01:01Z")

</div>

I think that I will try split such files on the multiply smaller files by pattern (on Python). It should be efficient than logstash performance.  
BTW. It's interesting how it looks like the limit of  
max\_lines =\> ?  
max\_bytes =\> ?  
when I used 16GB of mem per logstash instanace

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 18, 2022, 12:01pm UTC](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/15 "2022-08-18T12:01:23Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
