# Grok matching multi-lines saving first and last value to separate fields

**URL:** <https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958>\
**Category:** Logstash\
**Created:** [February 20, 2021, 8:41pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958 "2021-02-20T20:41:07Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Daniel\_Jankech](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_jankech/32/77691_2.png) [@Daniel\_Jankech](https://discuss.elastic.co/u/Daniel_Jankech)\
**Post date:** [February 20, 2021, 8:41pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/1 "2021-02-20T20:41:07Z")

</div>

Hello everyone, as the caption indicates I am trying to match multiple-lines of logs of which I would like to save first occurence of timestamp to separate field than last occurence of the timestamp. I am using codec multiline to separate tasks from big log file and (?m) regex before timestamp. I need to do this to measure time between process completion. Any ideas how this could be done ?

```
2020-12-16 15:43:31.605 INFO 18020 --- [http-nio-8080-exec-3] c.n.w.workflow.service.DataService : Getting groups of task 5fda1d109ceec746643760f8 in case 11.11.2020 13:20 level: 0
2020-12-16 15:43:34.346 INFO 18020 --- [http-nio-8080-exec-1] c.n.w.workflow.service.TaskService : [5fda1d109ceec746643760f5]: Task [GENERATE] in case [11.11.2020 13:20] assigned to [super@netgrif.com] was finished

```

I would like the output to be  
`Start: 15:43:31.605`  
`Severity: INFO`  
`INT: 18020`  
`Thread: http-nio-8080-exec-3`  
`Class: c.n.w.workflow.service.TaskService`  
`GREEDYDATA: [...,...,...]`  
`End: 15:43:34.346`

I was thinking of removing all items from array of timestamps generated by (?m) between first and last and then separate it into two different fields but Im not really sure how to do so.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [February 20, 2021, 9:19pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/2 "2021-02-20T21:19:31Z")

</div>

```
grok { match => { "message" => "\A%{TIMESTAMP_ISO8601:start}.*^%{TIMESTAMP_ISO8601:end}[^\n]*\Z" } }

```

\A anchors the first timestamp to the beginning of the message field. Then for the end time you anchor it to start of line using ^ and the `[^\n]*\Z` means there cannot be another newline from there to the very end of the text.

---

<div class="post-metadata">

**Author:** ![Daniel\_Jankech](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_jankech/32/77691_2.png) [@Daniel\_Jankech](https://discuss.elastic.co/u/Daniel_Jankech)\
**Post date:** [February 20, 2021, 9:46pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/3 "2021-02-20T21:46:39Z")

</div>

Yes! Thank you , that is exactly what I was looking for. There should be special awards for people like you helping out others at Saturday nights:)

---

<div class="post-metadata">

**Author:** ![Daniel\_Jankech](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_jankech/32/77691_2.png) [@Daniel\_Jankech](https://discuss.elastic.co/u/Daniel_Jankech)\
**Post date:** [February 21, 2021, 10:07am UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/4 "2021-02-21T10:07:32Z")

</div>

I just tried to use it incorporated into my pattern , but im not quite getting the output that I would like to be getting from `GREEDYDATA`. For some reason im getting single value rather than array of values from all the GREEDYDATA values.  
My entire pattern :  
`(?m)\A%{TIMESTAMP_ISO8601:start}.* %{SPACE} %{LOGLEVEL:LEVEL} %{INT:NUMBER} --{2} \[%{DATA:THREAD}] %{DATA:CLASS}\s(?m)%{GREEDYDATA:message}^%{TIMESTAMP_ISO8601:end}[^\n]*\Z`

Is my usage correct ?  
I thought (?m) before greedy data indicates multiline input and that would result into an array of values.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [February 21, 2021, 6:03pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/5 "2021-02-21T18:03:41Z")

</div>

grok will not return an array of matches. If you need multiple matches for a single pattern then use a ruby filter and the String .scan function. There is an example of doing that [here](https://discuss.elastic.co/t/match-multiple-4-digit-numbers-in-field/264494/2).

---

<div class="post-metadata">

**Author:** ![Daniel\_Jankech](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_jankech/32/77691_2.png) [@Daniel\_Jankech](https://discuss.elastic.co/u/Daniel_Jankech)\
**Post date:** [February 21, 2021, 8:48pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/6 "2021-02-21T20:48:44Z")

</div>

Sorry I didnt express myself correctly. I would be totally fine with the output that is in top answer [here](https://stackoverflow.com/questions/50502347/grok-parse-multiple-lines-for-example-exception-stack-trace) in the field "extralines". I thought this could be done with simple (?m) usage.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [February 21, 2021, 9:24pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/7 "2021-02-21T21:24:55Z")

</div>

It works for me. If I start with

```
   "message" => "2020-12-16 15:43:31.605 INFO 18020 --- [http-nio-8080-exec-3] c.n.w.workflow.service.DataService : Getting groups of task 5fda1d109ceec746643760f8 in case 11.11.2020 13:20 level: 0\nFoo\n Bar\n2020-12-16 15:43:34.346 INFO 18020 --- [http-nio-8080-exec-1] c.n.w.workflow.service.TaskService : [5fda1d109ceec746643760f5]: Task [GENERATE] in case [11.11.2020 13:20] assigned to [super@netgrif.com] was finished"

```

then

```
     grok { match => { "message" => "(?m)\A%{TIMESTAMP_ISO8601:start}.* %{SPACE} %{LOGLEVEL:LEVEL} %{INT:NUMBER} --{2} \[%{DATA:THREAD}] %{DATA:CLASS}\s(?m)%{GREEDYDATA:message}^%{TIMESTAMP_ISO8601:end}[^\n]*\Z" } }

```

results in

```
       "end" => "2020-12-16 15:43:34.346",
     "CLASS" => "c.n.w.workflow.service.DataService",
     "LEVEL" => "INFO",
   "message" => [
    [0] "2020-12-16 15:43:31.605 INFO 18020 --- [http-nio-8080-exec-3] c.n.w.workflow.service.DataService : Getting groups of task 5fda1d109ceec746643760f8 in case 11.11.2020 13:20 level: 0\nFoo\n Bar\n2020-12-16 15:43:34.346 INFO 18020 --- [http-nio-8080-exec-1] c.n.w.workflow.service.TaskService : [5fda1d109ceec746643760f5]: Task [GENERATE] in case [11.11.2020 13:20] assigned to [super@netgrif.com] was finished",
    [1] " : Getting groups of task 5fda1d109ceec746643760f8 in case 11.11.2020 13:20 level: 0\nFoo\n Bar\n"
],

```

etc. Note that the array of arrays is just a presentation thing at [https://grokdebug.herokuapp.com/](https://grokdebug.herokuapp.com/). Even in that SO answer you linked to, the GREEDYDATA is matching a single string.

If you want to get each line in a separate array entry then use mutate+split.

---

<div class="post-metadata">

**Author:** ![Daniel\_Jankech](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_jankech/32/77691_2.png) [@Daniel\_Jankech](https://discuss.elastic.co/u/Daniel_Jankech)\
**Post date:** [February 22, 2021, 11:39am UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/8 "2021-02-22T11:39:28Z")

</div>

Im not sure whats causing this but I cant seem to get the same output as you

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/f/cf05c89b43749f70579110b5b9f583af7cda32b0.png)  
This still results in  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/b/6b19f60733575e3cd40bd56a32f86f94b0b4842e.png)  
sorry for the screenshot I hope its good enough quality to see the content.

I copied your input message and pattern just to be 100% sure im not screwing up anywhere myself. Thanks for the tip for mutate + split I will certainly do that after I figure this out.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [February 22, 2021, 2:29pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/9 "2021-02-22T14:29:31Z")

</div>

That is exactly what I would expect you to get. What do you think should be different?

---

<div class="post-metadata">

**Author:** ![Daniel\_Jankech](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_jankech/32/77691_2.png) [@Daniel\_Jankech](https://discuss.elastic.co/u/Daniel_Jankech)\
**Post date:** [February 22, 2021, 2:42pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/10 "2021-02-22T14:42:53Z")

</div>

Sorry i re-read your answer and I thought I saw something different in output , I get why Im getting a single value now. Is it possible for me to get ALL the GREEDYDATA values in a single field ? Thanks:)

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [February 22, 2021, 2:57pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/11 "2021-02-22T14:57:28Z")

</div>

That is what you get by default. I do not understand what you want that is different from what you are getting.

---

<div class="post-metadata">

**Author:** ![Daniel\_Jankech](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/daniel_jankech/32/77691_2.png) [@Daniel\_Jankech](https://discuss.elastic.co/u/Daniel_Jankech)\
**Post date:** [February 22, 2021, 3:08pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/12 "2021-02-22T15:08:30Z")

</div>

I mean greedydata from ALL lines accumulated in a single field.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 22, 2021, 3:08pm UTC](https://discuss.elastic.co/t/grok-matching-multi-lines-saving-first-and-last-value-to-separate-fields/264958/13 "2021-03-22T15:08:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
