# How to match lines in unstructured log starting with specific string

**URL:** <https://discuss.elastic.co/t/how-to-match-lines-in-unstructured-log-starting-with-specific-string/284112>\
**Category:** Logstash\
**Created:** [September 13, 2021, 5:50pm UTC](https://discuss.elastic.co/t/how-to-match-lines-in-unstructured-log-starting-with-specific-string/284112 "2021-09-13T17:50:32Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![ansamHox](https://avatars.discourse-cdn.com/v4/letter/a/54ee81/32.png) [@ansamHox](https://discuss.elastic.co/u/ansamHox)\
**Post date:** [September 13, 2021, 5:50pm UTC](https://discuss.elastic.co/t/how-to-match-lines-in-unstructured-log-starting-with-specific-string/284112/1 "2021-09-13T17:50:32Z")

</div>

I have a very big unstructured file. Firstly, I want to parse all lines that starts with **ABC** or **CDE** and store them as one document in Elasticsearch. One file should be one document in index, so @message should look like all **ABC** + **CDE** lines.

Secondly, I also have "ignore list", like lines starting with empty space, "Timing", "----" etc., and basicaly ALL that is left ater parse is complete, I want to store in new key value (data). Final result should look something like:

```auto
 "_source" : {
          "message" : "ABC V4.1.2 MODEL,CDE: 0 uri: xxxx.xxx"
          "data" : "ALL LINES THAT ARE NOT IN IGNORE LIST"
 }

```

This is my logstash conf file for the first part, but for some reason it is not processing anything because index is empty, and I can't see any error in logs, probably because I'm not saving it properly.

```auto
input {
    file {
        path => "/etc/logstash/files/*"
        start_position => "beginning"
        sincedb_path => "/dev/null"
    }
}

filter {

  if ([message] !~ "^ABC"){
   drop{}
  }
  else if ([message] !~ "^CDE"){
    drop{}
  }
}

output{
    elasticsearch {
        hosts => ["XXX"]
        index => "index1"
 }
}

```

When I add 1 ef expression it works, but when I add second one , it doesn't process anything. Any help? Thank you

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [September 13, 2021, 6:38pm UTC](https://discuss.elastic.co/t/how-to-match-lines-in-unstructured-log-starting-with-specific-string/284112/2 "2021-09-13T18:38:02Z")

</div>

If you want to combine all the lines from one file into a single document then you could do it using an aggregate filter, but I would use a multiline codec to read the entire file as a single event as described [here](https://discuss.elastic.co/t/parsing-array-of-json-objects-with-logstash-and-injesting-to-elastic/203197/2).

I would then do the processing with a ruby filter.

```
    ruby {
        code => '
            lines = event.get("message").lines(chomp: true)
            newMessage = ""
            theRest = ""
            lines.each { |x|
                if x =~ /^(ABC|CDE)/
                    newMessage += x + ","
                else
                    unless x =~ /^(\s|Timing)/
                        theRest += x + ","
                    end
                end
            }
            event.set("message", newMessage)
            event.set("data", theRest)
        '
    }

```

Obviously you will want to tune those regular expressions. With the file you showed that will get you

```
      "data" => "Quilting with 1 groups of 0 I/O tasks.,DYNAMICS OPTION: Eulerian Mass Coordinate,",
   "message" => "ABC V4.1.2 MODEL,ABC restart, LBC starts at 1979-12-19_00:00:00 and restart starts at 1979-12-19_00:00:00,CDE: 0 hostname: xxxx.xxx,"
```

---

<div class="post-metadata">

**Author:** ![ansamHox](https://avatars.discourse-cdn.com/v4/letter/a/54ee81/32.png) [@ansamHox](https://discuss.elastic.co/u/ansamHox)\
**Post date:** [September 13, 2021, 7:14pm UTC](https://discuss.elastic.co/t/how-to-match-lines-in-unstructured-log-starting-with-specific-string/284112/3 "2021-09-13T19:14:41Z")

</div>

Thank you Badger, as usual 🙂

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [September 13, 2021, 7:18pm UTC](https://discuss.elastic.co/t/how-to-match-lines-in-unstructured-log-starting-with-specific-string/284112/4 "2021-09-13T19:18:35Z")

</div>

Sounds like you are not using the multiline codec correctly. Note that you may need to configure max\_lines and max\_bytes if the files are tens of thousands of lines long.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 11, 2021, 7:31pm UTC](https://discuss.elastic.co/t/how-to-match-lines-in-unstructured-log-starting-with-specific-string/284112/7 "2021-10-11T19:31:36Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
