# RegEx in Logstash Conf File: loop through/retrieve all matches

**URL:** <https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020>\
**Category:** Logstash\
**Created:** [August 6, 2019, 1:51pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020 "2019-08-06T13:51:17Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![dundundundun](https://avatars.discourse-cdn.com/v4/letter/d/87869e/32.png) [@dundundundun](https://discuss.elastic.co/u/dundundundun)\
**Post date:** [August 6, 2019, 1:51pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/1 "2019-08-06T13:51:17Z")

</div>

Let's suppose my log messages are as following:

GroupName::Cars: mileage=19274 ~ year=2000  
GroupName:🏫 students = 10000 ~ location = USA ~ classes = 50 ~ staff = 75

How can I pattern match both of these lines using the same RegEx pattern in the logstash configuration file? Is there a way to continue pattern matching until some unspecified number of "~"?

I know that it is possible to create two separate grok filters as below, but wondering if there's a cleaner way to collapse the two pattern matches into one

```
if [message] =~ "Cars" {
grok {
    match => {
	        "message" => "%{GREEDYDATA}%{KEYWORD}=%{MILEAGE:mileage} ~ %{KEYWORD}=%{YEAR:year}"
    }
    pattern_definitions => {
        "KEYWORD" => "[\w]{4,7}"
        "MILEAGE" => "[\d]{1,10}"
        "YEAR" => "[\d]{4}"
    }
}
 }

if [message] =~ "School" {
    grok {
        match => {
	        "message" => "%{GREEDYDATA}%{KEYWORD} = %{STUDENTS:students} ~ %{KEYWORD} = %{LOCATION:location} ~ %{KEYWORD} = %{CLASSES:classes} ~ %{KEYWORD} = %{STAFF:staff}"
        }
        pattern_definitions => {
            "KEYWORD" => "[\w]{4,7}"
            "STUDENTS" => "[\d]{1,10}"
            "LOCATION" => "[\w]{1,10}"
            "CLASSES" => "[\d]{1,10}"
            "STAFF" => "[\d]{1,10}"
        }
    }
}
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 6, 2019, 1:58pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/2 "2019-08-06T13:58:09Z")

</div>

If you really want to match a single pattern you could use alternation, but it doesn't make much sense to me.

```
grok { match => { "message" => "(pattern1|pattern2)" } }

```

I would combine the two groks into one and match [message] against an array of patterns...

```
grok {
    match => {
        "message" => [
                "%{KEYWORD}=%{MILEAGE:mileage} ~ %{KEYWORD}=%{YEAR:year}",
                "%{KEYWORD} = %{STUDENTS:students} ~ %{KEYWORD} = %{LOCATION:location} ~ %{KEYWORD} = %{CLASSES:classes} ~ %{KEYWORD} = %{STAFF:staff}"
        ]
        ...

```

Note that the leading %{GREEDYDATA} is not needed. The patterns are not anchored, so they do not have to match from the beginning of the field.

---

<div class="post-metadata">

**Author:** ![dundundundun](https://avatars.discourse-cdn.com/v4/letter/d/87869e/32.png) [@dundundundun](https://discuss.elastic.co/u/dundundundun)\
**Post date:** [August 6, 2019, 2:04pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/3 "2019-08-06T14:04:39Z")

</div>

Thanks! Let's say that the log messages may have new GroupNames in the future with an unspecified number of parameters. All I know is that the messages will be of this format:  
GroupName:: \<parameter\_1 name\> = \<parameter\_1 value\> ~ \<parameter\_2 name\> = \<parameter\_2 value\> ~ \<parameter\_n name\> = \<parameter\_n value\>

What would be a good way to pattern match?

---

<div class="post-metadata">

**Author:** ![jmilot](https://avatars.discourse-cdn.com/v4/letter/j/77aa72/32.png) [@jmilot](https://discuss.elastic.co/u/jmilot)\
**Post date:** [August 6, 2019, 2:05pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/4 "2019-08-06T14:05:26Z")

</div>

Something like this :

```
%{DATA:data}::%{WORD:category}:%{SPACE}%{WORD:keyword}=%{NUMBER:value}?(%{SPACE} ~ %{WORD:keyword}=%{NUMBER:value})*
```

---

<div class="post-metadata">

**Author:** ![dundundundun](https://avatars.discourse-cdn.com/v4/letter/d/87869e/32.png) [@dundundundun](https://discuss.elastic.co/u/dundundundun)\
**Post date:** [August 6, 2019, 2:08pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/5 "2019-08-06T14:08:54Z")

</div>

The ?(...)\* means it's optional I'm assuming? So if I had up to 5 parameters I would do something like this?

`%{DATA:data}::%{WORD:category}:%{SPACE}?(%{SPACE} ~ %{WORD:keyword}=%{NUMBER:value})*?(%{SPACE} ~ %{WORD:keyword}=%{NUMBER:value})*?(%{SPACE} ~ %{WORD:keyword}=%{NUMBER:value})*?(%{SPACE} ~ %{WORD:keyword}=%{NUMBER:value})*?(%{SPACE} ~ %{WORD:keyword}=%{NUMBER:value})*`

---

<div class="post-metadata">

**Author:** ![jmilot](https://avatars.discourse-cdn.com/v4/letter/j/77aa72/32.png) [@jmilot](https://discuss.elastic.co/u/jmilot)\
**Post date:** [August 6, 2019, 2:11pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/6 "2019-08-06T14:11:13Z")

</div>

'\*' is like in bash regexp : zero or more  
? is for optional

For example :

```
GroupName::Cars: mileage=19274 ~ year=2000

{
  "data": [
[
  "GroupName"
]
  ],
  "category": [
[
  "Cars"
]
  ],
  "keyword": [
[
  "mileage",
  "year"
]
  ],
  "value": [
[
  "19274",
  "2000"
]
  ]
}
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 6, 2019, 2:41pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/7 "2019-08-06T14:41:26Z")

</div>

> [@dundundundun](#):
>
> What would be a good way to pattern match?

Use mutate+gsub to remove the prefix then use a kv filter.

---

<div class="post-metadata">

**Author:** ![dundundundun](https://avatars.discourse-cdn.com/v4/letter/d/87869e/32.png) [@dundundundun](https://discuss.elastic.co/u/dundundundun)\
**Post date:** [August 7, 2019, 8:38pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/8 "2019-08-07T20:38:15Z")

</div>

What's the correct syntax for the statement below?

`if [path] =~ "a" OR [path] =~ "b"?`

I want to perform the same filter if path contains two particular keywords. Rather than copy/paste filter, I want to collapse it into one. So instead of

```
if [path] =~ "a" {
    // some filter
} else if [path] =~ "a" {
    // same filter
}

```

I want  
if [path] =~ "a" OR [path] =~ "b" {  
// some filter  
}

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 7, 2019, 8:48pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/9 "2019-08-07T20:48:34Z")

</div>

Lower case or.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 4, 2019, 8:48pm UTC](https://discuss.elastic.co/t/regex-in-logstash-conf-file-loop-through-retrieve-all-matches/194020/10 "2019-09-04T20:48:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
