# How to capture all entries that matches the same search pattern

**URL:** <https://discuss.elastic.co/t/how-to-capture-all-entries-that-matches-the-same-search-pattern/360473>\
**Category:** Logstash\
**Created:** [May 29, 2024, 1:22pm UTC](https://discuss.elastic.co/t/how-to-capture-all-entries-that-matches-the-same-search-pattern/360473 "2024-05-29T13:22:47Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![thosch](https://avatars.discourse-cdn.com/v4/letter/t/4da419/32.png) [@thosch](https://discuss.elastic.co/u/thosch)\
**Post date:** [May 29, 2024, 1:22pm UTC](https://discuss.elastic.co/t/how-to-capture-all-entries-that-matches-the-same-search-pattern/360473/1 "2024-05-29T13:22:47Z")

</div>

Hello,

i am new at using logstash and its filters. I got log files looking like this:

\</\>  
Started 'DB to PostgreSQL transfer' workflow at 2024.05.19 09:31:07

database v87\_orig\_rec\_8904\_03  
########################################################

timestamp | module name | result

* * *

2024.05.19 13:37:06 | DbToPostgres | success  
2024.05.19 13:00:49 | PgDump | success  
2024.05.19 09:43:30 | Converter | success  
2024.05.19 10:00:40 | Import | success  
2024.05.19 10:29:59 | Validation | failure  
2024.05.19 12:31:40 | DbUpdate | success  
2024.05.19 12:45:46 | Final | success  
2024.05.19 11:21:25 | CreateArchive | failure  
2024.05.19 13:37:02 | DbInsert | success  
2024.05.19 11:35:37 | GenerateBucket | success  
2024.05.19 09:36:10 | PrepareDB | success

Finishing workflow at 2024.05.19 13:37:06

\</\>  
I create a multiline in filebeat. This is the original event:

\</\>  
"original" =\> "Started 'DB to PostgreSQL transfer' workflow at 2024.05.19 09:31:07\n\ndatabase v87\_orig\_rec\_r8904\_03\n########################################################\n\ntimestamp | module name | result\n\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\n2024.05.19 13:37:06 | DbToPostgres | success\n2024.05.19 13:00:49 | PgDump | success\n2024.05.19 09:43:30 | Converter | success\n2024.05.19 10:00:40 | Import | success\n2024.05.19 10:29:59 | Validation | failure\n2024.05.19 12:31:40 | DbUpdate | success\n2024.05.19 12:45:46 | Final | success\n2024.05.19 11:21:25 | CreateArchive | failure\n2024.05.19 13:37:02 | DbInsert | success\n2024.05.19 11:35:37 | GenerateBucket | success\n2024.05.19 09:36:10 | PrepareDB | success\n\nFinishing workflow at 2024.05.19 13:37:06"  
\</\>

I would like capture all matches from the search pattern(s) in one message so that one log appears in elastic/opensearch as one hit. That works so far but only for search patterns witch find only a single match. The search pattern which match multiple lines (like 2024.05.19 13:37:06 | DbToPostgres | success and so on) only captures the first match when i use this pattern:

\</\>  
"Started 'DB to PostgreSQL transfer' workflow at (?\<workflow\_start\_timestamp\>%{YEAR:workflow\_start\_year}.%{MONTHNUM:workflow\_start\_month}.%{MONTHDAY:workflow\_start\_day} %{HOUR:workflow\_start\_hour}:%{MINUTE:workflow\_start\_minute}:%{SECOND:workflow\_start\_second})\n\ndatabase(?%{DATA:DB\_VERSION}_%{DATA:DB\_TYPE}_%{DATA:DB\_TYPE\_NEW}_%{DATA:DB\_ID}_%{NUMBER:REV\_NO})\n#+\n\ntimestamp\s\*|\s_module name\s_|\s_result\n\_+\n(?m)(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s_(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}%{GREEDYDATA:message}Finishing workflow at (?\<workflow\_finish\_timestamp\>%{YEAR:workflow\_finish\_year}.%{WORD:workflow\_finish\_month}.%{WORD:workflow\_finish\_day} %{TIME:workflow\_finish\_time})".  
\</\>

When i use this pattern:

\</\>  
"Started 'DB to PostgreSQL transfer' workflow at (?\<workflow\_start\_timestamp\>%{YEAR:workflow\_start\_year}.%{MONTHNUM:workflow\_start\_month}.%{MONTHDAY:workflow\_start\_day} %{HOUR:workflow\_start\_hour}:%{MINUTE:workflow\_start\_minute}:%{SECOND:workflow\_start\_second})\n\ndatabase(?%{DATA:DB\_VERSION}_%{DATA:DB\_TYPE}_%{DATA:DB\_TYPE\_NEW}_%{DATA:DB\_ID}_%{NUMBER:REV\_NO})\n#+\n\ntimestamp\s\*|\s_module name\s_|\s_result\n\_+\n((?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s_(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n)+%{GREEDYDATA:message}Finishing workflow at %{YEAR:workflow\_finish\_year}.%{WORD:workflow\_finish\_month}.%{WORD:workflow\_finish\_day} %{TIME:workflow\_finish\_time}"

\</\>

it only caputres the last match oft he multiple matches (2024.05.19 09:36:10 | PrepareDB ) . I would like to capture all logtimestamps, module names and results in arrays like this:

\</\>  
"workflow\_start\_year" =\> "2024",  
"log\_timestamp" =\> [  
[0] "2024.05.19 13:37:06",  
[1] "2024.05.19 13:00:49",  
[2] "2024.05.19 09:43:30",  
[3] "2024.05.19 10:00:40"  
],  
"workflow\_start\_day" =\> "19",  
"log.date" =\> [  
[0] "2024.05.19",  
[1] "2024.05.19",  
[2] "2024.05.19",  
[3] "2024.05.19"  
],  
"module\_name" =\> [  
[0] "DbToPostgres",  
[1] "PgDump",  
[2] "Converter",  
[3] "Import"  
],  
"log\_day" =\> [  
[0] "19",  
[1] "19",  
[2] "19",  
[3] "19"  
],

\</\>

That only works if i repeat the pattern multiple times like this:

\</\>  
"Started 'DB2 to PostgreSQL transfer' workflow at (?\<workflow\_start\_timestamp\>%{YEAR:workflow\_start\_year}.%{MONTHNUM:workflow\_start\_month}.%{MONTHDAY:workflow\_start\_day} %{HOUR:workflow\_start\_hour}:%{MINUTE:workflow\_start\_minute}:%{SECOND:workflow\_start\_second})\n\ndatabase(?%{DATA:DB\_VERSION}_%{DATA:DB\_TYPE}_%{DATA:DB\_TYPE\_NEW}_%{DATA:DB\_ID}_%{NUMBER:REV\_NO})\n#+\n\ntimestamp\s\*|\s_module name\s_|\s_result\n\_+\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s_(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s\*(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s\*(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s\*(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n%{GREEDYDATA:message}Finishing workflow at %{YEAR:workflow\_finish\_year}.%{WORD:workflow\_finish\_month}.%{WORD:workflow\_finish\_day} %{TIME:workflow\_finish\_time}"

\</\>

The problem is, the number of entries in the table (logtimestamp | module name | result) varies from log to log and i can not set endless entries of the same search pattern one after another to capture all entries.

Here ist the complete filter section:

\</\>

filter {  
if [fields][logtype] == 'importer' {  
grok {  
match =\> {  
"message" =\> [  
"Started 'DB2 to PostgreSQL transfer' workflow at (?\<workflow\_start\_timestamp\>%{YEAR:workflow\_start\_year}.%{MONTHNUM:workflow\_start\_month}.%{MONTHDAY:workflow\_start\_day} %{HOUR:workflow\_start\_hour}:%{MINUTE:workflow\_start\_minute}:%{SECOND:workflow\_start\_second})\n\ndatabase(?%{DATA:DB\_VERSION}_%{DATA:DB\_TYPE}_%{DATA:DB\_TYPE\_NEW}_%{DATA:DB\_ID}_%{NUMBER:REV\_NO})\n#+\n\ntimestamp\s\*|\s_module name\s_|\s_result\n\_+\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s_(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s\*(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s\*(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n(?\<log\_timestamp\>(?\<log.date\>%{YEAR:log\_year}.%{MONTHNUM:log\_month}.%{MONTHDAY:log\_day})\s\*(?\<log.time\>%{HOUR:log\_hour}:%{MINUTE:log\_minute}:%{SECOND:log\_second}))\s\*|\s\*%{WORD:module\_name}\s\*|\s\*%{WORD:result}\n%{GREEDYDATA:message}Finishing workflow at %{YEAR:workflow\_finish\_year}.%{WORD:workflow\_finish\_month}.%{WORD:workflow\_finish\_day} %{TIME:workflow\_finish\_time}"  
]  
}  
break\_on\_match =\> false  
overwrite =\> ["message"]  
}

\</\>

I also tried the search pattern with the multiline tag (?m) and „(SEARCH\_PATTERN)+“ but only one match of the table entries is captured.

How can I define the Grok filter to capture all entries in a single event? How can all entries that match the same search filter be saved in lists?

I don't want to split the multiline into individual lines. I have already tested this successfully.

Best regards

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 29, 2024, 1:22pm UTC](https://discuss.elastic.co/t/how-to-capture-all-entries-that-matches-the-same-search-pattern/360473/2 "2024-05-29T13:22:47Z")

</div>

OpenSearch/OpenDistro are AWS run products and differ from the original Elasticsearch and Kibana products that Elastic builds and maintains. You may need to contact them directly for further assistance. See [What is OpenSearch and the OpenSearch Dashboard? | Elastic](https://www.elastic.co/elasticsearch/opensearch) for more details.

(This is an automated response from your friendly Elastic bot. Please report this post if you have any suggestions or concerns :elasticheart: )
