# Split grok-pattern into multiple lines

**URL:** <https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366>\
**Category:** Logstash\
**Created:** [August 27, 2025, 8:45am UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366 "2025-08-27T08:45:13Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![sthorn](https://avatars.discourse-cdn.com/v4/letter/s/e47c2d/32.png) [@sthorn](https://discuss.elastic.co/u/sthorn)\
**Post date:** [August 27, 2025, 8:45am UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/1 "2025-08-27T08:45:13Z")

</div>

Hi,

Is it possible to split a grok-pattern into multiple lines instead of have one big?

```ruby
grok {
  pattern_definitions => {
    "CM" => "[/.\-\w\s]*"
  }
  match => {"syslog_message" => "%{CM:unknown_id} %{IP:this_ip},%{HOSTNAME:hostname},%{CM:resource_type},%{CM:some_name},%{CM:unknown_01},%{IP:src_ip},%{CM:unknown_02},%{IP:dst_ip},%{NUMBER:src_port:int},%{NUMBER:dst_port:int},%{CM:partition},%{CM:protocol},%{NUMBER:domain},%{CM:unknown_03},%{CM:unknown_04},%{CM:unknown_05},%{CM:unknown_06},%{CM:unknown_09},%{CM:unknown_10},%{CM:unknown_11},%{CM:policy_type},%{CM:policy_name},%{CM:rule_name},%{CM:unknown_15},%{CM:dev_action},%{CM:unknown_17},%{CM:unknown_18},%{CM:unknown_19},%{CM:unknown_20},%{CM:unknown_21},%{CM:unknown_22},%{CM:unknown_23},%{CM:unknown_24},%{CM:unknown_25},%{CM:unknown_26},%{CM:unknown_27},%{CM:unknown_28},%{CM:unknown_29},%{CM:unknown_30}"}
}

```

This line is way to long.

---

<div class="post-metadata">

**Author:** ![Rios](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rios/32/95745_2.png) [@Rios](https://discuss.elastic.co/u/Rios)\
**Post date:** [August 27, 2025, 9:10am UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/2 "2025-08-27T09:10:50Z")

</div>

Most likely yes, we need to se the org. message and what should be result.

---

<div class="post-metadata">

**Author:** ![sthorn](https://avatars.discourse-cdn.com/v4/letter/s/e47c2d/32.png) [@sthorn](https://discuss.elastic.co/u/sthorn)\
**Post date:** [August 27, 2025, 1:48pm UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/3 "2025-08-27T13:48:41Z")

</div>

The parsing works fine, I see no need for a message.

The issue I’m finding is having a line with 700+ chars in git and editors is not optimal.  
Is there a way to build the grok-pattern over multiple lines, with an `<<` or `+=` operator perhaps?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [August 27, 2025, 1:54pm UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/4 "2025-08-27T13:54:50Z")

</div>

Can you share a sample of your message?

From the pattern you are using in the `grok` filter your message seems to be a csv message, you could use the `csv` filter to parse it instead of `grok`.

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [August 27, 2025, 3:25pm UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/5 "2025-08-27T15:25:43Z")

</div>

@leandrojmp has good suggestion to use csv...

> [@sthorn](#):
>
> The issue I’m finding is having a line with 700+ chars in git and editors is not optimal.  
> Is there a way to build the grok-pattern over multiple lines, with an `<<` or `+=` operator perhaps?

But to answer your question ... Yes... but it will not be effecient....

It would be something like

```auto
grok {
  pattern_definitions => {
    "CM" => "[/.\-\w\s]*"
  }
  match => {"syslog_message" => "%{CM:unknown_id} %{IP:this_ip},%{HOSTNAME:hostname},...%{GREEDYDATA:msg_part2}
}

grok {
  pattern_definitions => {
    "CM" => "[/.\-\w\s]*"
  }
  match => {"msg_part2" => "<Grok Patterns>%{GREEDYDATA:msg_part3}
}

grok {
  pattern_definitions => {
    "CM" => "[/.\-\w\s]*"
  }
  match => {"msg_part3" => "<Grok Patterns>}
}

```

Not as efficient

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 27, 2025, 3:50pm UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/6 "2025-08-27T15:50:39Z")

</div>

> [@sthorn](#):
>
> Is there a way to build the grok-pattern over multiple lines, with an `<<` or `+=` operator perhaps?

Yes, you can use custom pattern definitions within a custom pattern definition.

```auto
input { generator { count => 1 lines => ['Foo, Or Bar,Or Baz'] } }

output { stdout { codec => rubydebug { metadata => false } } }
filter {
    grok {
        pattern_definitions => {
            ONE => "%{WORD}"
            TWO => "(?<Foo>[^,]*),%{GREEDYDATA}"
            OVERALL => "%{ONE},%{TWO}"
        }
        match => { "message" => "^%{OVERALL}" }
    }

```

will produce

```auto
       "Foo" => " Or Bar"

```

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [August 27, 2025, 4:49pm UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/7 "2025-08-27T16:49:14Z")

</div>

> [@Badger](#):
>
> Yes, you can use custom pattern definitions within a custom pattern definition.

TIL!

---

<div class="post-metadata">

**Author:** ![sthorn](https://avatars.discourse-cdn.com/v4/letter/s/e47c2d/32.png) [@sthorn](https://discuss.elastic.co/u/sthorn)\
**Post date:** [August 28, 2025, 10:04am UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/8 "2025-08-28T10:04:43Z")

</div>

Great I will try that.

---

<div class="post-metadata">

**Author:** ![sthorn](https://avatars.discourse-cdn.com/v4/letter/s/e47c2d/32.png) [@sthorn](https://discuss.elastic.co/u/sthorn)\
**Post date:** [August 28, 2025, 10:07am UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/9 "2025-08-28T10:07:51Z")

</div>

Yes, the message looks very much as a csv-row in this example, but in our real config there is multiple “match”-lines and the input varies in number of columns.  
The first and second column determines what column has what values.

The idea with CSV is good and I did not think of it, will try and see if we can use it.

Thank!

---

<div class="post-metadata">

**Author:** ![sthorn](https://avatars.discourse-cdn.com/v4/letter/s/e47c2d/32.png) [@sthorn](https://discuss.elastic.co/u/sthorn)\
**Post date:** [August 28, 2025, 10:09am UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/10 "2025-08-28T10:09:14Z")

</div>

This is one of the possible alternatives that we thought of.  
We also was thinking of the performance of it.

Thanks!

---

<div class="post-metadata">

**Author:** ![Rios](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rios/32/95745_2.png) [@Rios](https://discuss.elastic.co/u/Rios)\
**Post date:** [August 28, 2025, 10:38am UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/11 "2025-08-28T10:38:36Z")

</div>

Not sure which are better performances, CSV or dissect, you can try both in your case. The dissect filter is ~10x faster than grok. However csv is much more useful if you have pure csv format.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [August 28, 2025, 12:47pm UTC](https://discuss.elastic.co/t/split-grok-pattern-into-multiple-lines/381366/12 "2025-08-28T12:47:18Z")

</div>

> [@sthorn](#):
>
> but in our real config there is multiple “match”-lines and the input varies in number of columns.

This is not an issue, if you have different types of message you can still combine othe filters or use conditional to correctly parse it.

> [@sthorn](#):
>
> The first and second column determines what column has what values.

Not clear what you mean with this, without you sharing sample of messages is pretty complicated to provide any insight.

The main thing is that while `grok` can parse almost anything, sometimes you can use other parse filters or combination of other parse filters to make things easier.

Personally I only use `grok` as the last option, when a message cannot be parsed using other filters or combination of filters.
