# Catching all lines between two lines (two lines included)

**URL:** <https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [October 17, 2016, 5:52am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150 "2016-10-17T05:52:14Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 17, 2016, 5:52am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/1 "2016-10-17T05:52:14Z")

</div>

Hello,

I have not found a way to do this with multiline.  
Basically what I want to do is take all lines between two specific lines, and make a single message out of it.

Example :

_Line I don't want_  
_Line I don't want_  
_Line I don't want_  
_--- BEGINNING OF BLOCK I WANT_  
_Line of the block I want_  
_Line of the block I want_  
_Line of the block I want_  
_Line of the block I want_  
_Line of the block I want_  
_--- END OF BLOCK I WANT_  
_Line I don't want_  
_Line I don't want_  
_Line I don't want_  
_Line I don't want_

I tried to use multiline to catch any line that follow the line with _^--- [A-Z]+_ pattern but obviously doesn't work since it would ignore the _end of the block line_, and then everything is a mess.

Any tips ?  
Thanks.

---

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 18, 2016, 7:06am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/2 "2016-10-18T07:06:26Z")

</div>

I was thinking that I could maybe take everything after _--- BEGINNING OF BLOCK I WANT_ then include only the lines that are in the block (they are always the same pattern it seems), but that would be very unpractical, and not even sure it would work, since the include\_files would apply on a block (since multiline is applied before the include\_files).

Would greatly appreciate some guidance 🙂

---

<div class="post-metadata">

**Author:** ![tudor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tudor/32/3753_2.png) [@tudor](https://discuss.elastic.co/u/tudor)\
**Post date:** [October 18, 2016, 8:39am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/3 "2016-10-18T08:39:52Z")

</div>

I'm afraid we don't have a good solution for this with the current multiline implementation. For a proper solution, we'd need to start and stop patterns for multiline, which we discussed before but never implemented. Posting a ticket for an enhancement request for this would make sense.

If you post a real portion of your logs, perhaps we can think of a workaround.

---

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 18, 2016, 9:31am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/4 "2016-10-18T09:31:56Z")

</div>

Thanks for your reply, I understand.

Here's a sample (I'm sorry it's in french), what I would want is the block, or at least all the lines but the last line :

_04/10/2016 09:01:48 ANS1898I \*\*\*\*\* 5 115 000 fichiers trait▒s \*\*\*\*\*_  
_04/10/2016 09:01:49 ANS1898I \*\*\*\*\* 5 119 000 fichiers trait▒s \*\*\*\*\*_  
_04/10/2016 09:02:29 ANS1999E Le traitement de Incr▒mentale pour '/' est arr▒t▒._

**_04/10/2016 09:02:29 --- DEBUT ETAT JOURNAL DES OPERATIONS PLANIFIEES_**  
**_04/10/2016 09:02:29 Nombre total d'objets inspect▒s : 5 119 261_**  
**_04/10/2016 09:02:29 Nombre total d'objets sauvegard▒s : 163_**  
**_04/10/2016 09:02:29 Nombre total d'objets mis ▒ jour : 0_**  
**_04/10/2016 09:02:29 Nombre total d'objets reli▒s : 0_**  
**_04/10/2016 09:02:29 Nombre total d'objets supprim▒s : 0_**  
**_04/10/2016 09:02:29 Nombre total d'objets expir▒s : 0_**  
**_04/10/2016 09:02:29 Nombre total d'objets en ▒chec : 0_**  
**_04/10/2016 09:02:29 Nombre total d'objets chiffr▒s : 0_**  
**_04/10/2016 09:02:29 Le nombre total d'objets a augment▒ : 0_**  
**_04/10/2016 09:02:29 Nombre total de tentatives : 0_**  
**_04/10/2016 09:02:29 Nombre total d'octets inspect▒s : 7,99 TB_**  
**_04/10/2016 09:02:29 Nombre total d'octets transf▒r▒s : 927,09 MB_**  
**_04/10/2016 09:02:29 Dur▒e de transfert des donn▒es : 101,22 sec_**  
**_04/10/2016 09:02:29 D▒bit de transfert de donn▒es du r▒seau : 9 378,40 ko/s_**  
**_04/10/2016 09:02:29 D▒bit de transfert d'un groupe de fichiers : 960,91 ko/s_**  
**_04/10/2016 09:02:29 Taux de compression des objets : 0%_**  
**_04/10/2016 09:02:29 Rapport de r▒duction des donn▒es total : 99,99%_**  
**_04/10/2016 09:02:29 Temps de traitement ▒coul▒ : 00:16:27_**  
**_04/10/2016 09:02:29 --- FIN ETAT JOURNAL DES OPERATIONS PLANIFIEES_**  
_04/10/2016 09:02:29 ANS4023E Erreur lors du traitement de '/' : erreur d'entr▒e/sortie sur le fichier_

Thanks.

---

<div class="post-metadata">

**Author:** ![tudor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tudor/32/3753_2.png) [@tudor](https://discuss.elastic.co/u/tudor)\
**Post date:** [October 18, 2016, 9:37am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/5 "2016-10-18T09:37:01Z")

</div>

The workaround that I'm thinking is to have a prospector that:

- Filters out all the lines starting with ANS
- Filters out the FIN line
- Groups together using multiline everything after the DEBUT line.

But this assume you don't need the ANS lines. If you need those, you'll need a second Filebeat instance with a separate registry file to get them (and only them), because I think using multiple prospectors on the same file won't work well.

Just an idea.

---

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 18, 2016, 9:42am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/6 "2016-10-18T09:42:08Z")

</div>

Thanks for the suggestion. I will try something like that, though I think it may not work since multilines are applied before the exclude/include\_lines (as said in the documentation).

I think I will be able to make it work somehow anyway.

Regards.

---

<div class="post-metadata">

**Author:** ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)\
**Post date:** [October 18, 2016, 11:02am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/7 "2016-10-18T11:02:55Z")

</div>

There is not much structural difference between lines to be composed into multiline and other lines. This makes matching a little hard.

All fields seem to be some 'numeric' stats , pointing to a regex pattern like `' : \d'` for matching a collon followed by a single digit. This can be refined a little by including some keywords to reduce the chance of false-positives like: `(total|transfert|objects|Temps de).+ : \d` . The `|` operator means `or` introducing some kind of backtracking. Depending on `match` setting you use, you can modify the regex to also include `'^--- FIN'` or `'^--- DEBUT'`. For example

```auto
'(^--- FIN)|((total|transfert|objects|Temps de).+ : \d)'

```

It can be kind of tricky to write a regex pattern like this without creating false-positives and might require some refinements, but it might be a start.

Feel free to create an [enhancement request](https://github.com/elastic/beats/issues), as having a start/end-pattern like matcher would solve the problem much more robustly.

---

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 18, 2016, 5:44pm UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/8 "2016-10-18T17:44:18Z")

</div>

I can totally see this working. I will update it if does 🙂

Thanks a whole lot!

---

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 19, 2016, 5:44am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/9 "2016-10-19T05:44:27Z")

</div>

It did work, thanks again 🙂  
The lines in the block were always the same, so I just went ahead and wrote them completely in the regex, to be sure it won't catch anything else.

Then I used an _include\_lines_ option to throw away all lines that don't start with _--- DEBUT ETAT JOURNAL_

Have a nice day.

---

<div class="post-metadata">

**Author:** ![steffens](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steffens/32/79630_2.png) [@steffens](https://discuss.elastic.co/u/steffens)\
**Post date:** [October 19, 2016, 11:38am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/10 "2016-10-19T11:38:31Z")

</div>

Wow, you must have a really big regex by now. As you're filtering out anything not starting with `--- DEBUT` some false positives outside of your multiline shouldn't be bad at all.

Regex engines often require `O(n^2)` for string matching (the matcher searches for a substring) + backtracking all the patterns might be relatively expensive. I'd monitor filebeat CPU usage and check if using another regex would optimize resource usage a little. Having a pattern starting with '^' can create a one-pass regex in some cases.

e.g. `'^.{20} ((--- DEBUT)|((Nombre|Le nombre|Rapport).+ (total|transfert|objects|Temps de).+ : \d))'`

here I'm using `^.{20}` to ignore the timestamp (exactly 20 characters) + introduce some early stop words like `(Nombre|Le nombre|Rapport)` + introduce some more 'evidence' for filtering out false positives: `(total|transfert|objects|Temps de)` and `.+ : \d`.

With `()` introducing capture groups the matcher might still capture them + throw the results away (depends on actual implementation). Using `(?:)` introduces a non-capturing group:

`'^.{20} (?:(?:--- DEBUT)|(?:(?:Nombre|Le nombre|Rapport).+(?:total|transfert|objects|Temps de).+ : \d))'`

Here is a tool to analyze the generated regex: [https://github.com/urso/anareg](https://github.com/urso/anareg)

Usage:

```auto
 ./anareg '^.{20} (?:(?:--- DEBUT)|(?:(?:Nombre|Le nombre|Rapport).+(?:total|transfert|objects|Temps de).+ : \d))' | dot -Tsvg > tst.svg

```

remember to verify resource usage when changing patterns. In case of CPU not being really different for different patterns use the most maintainable one.

 ![](https://us1.discourse-cdn.com/elastic/original/2X/2/26bcd227a34dca4fa156d2418270402451a32d14.png)

---

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 19, 2016, 2:11pm UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/11 "2016-10-19T14:11:48Z")

</div>

Wow, that's some amazing work there, thanks lots!

My regexes wasn't actually really long (for either the multiline or the include\_files)

The multiline regex was something like this :

`^[0-9]{2}\/[0-9]{2}\/[0-9]{4} [0-9]{2}:[0-9]{2}:[0-9]{2}: (--- FIN ETAT JOURNAL)|(Nombre total)|(Le nombre total)|(Dur.e de transfert)` etc. (can't c/c, not at work at the moment).

This worked and gave me the correct block starting with `--- DEBUT ETAT JOURNAL` and ending with `FIN ETAT JOURNAL` (note that I **must** add `ETAT JOURNAL` because there are other blocks that starts with `DEBUT` and end with `FIN`)

I then saw a lot of lines getting caught nevertheless, and those weren't following any patterns so I just decided to use the include\_files option to trash them, like this :  
`include_files : ['^[0-9]{2}\/[0-9]{2}\/[0-9]{4} [0-9]{2}:[0-9]{2}:[0-9]{2}: --- DEBUT ETAT JOURNAL']`

I haven't seen the result of that include\_lines yet since I need to wait for the server to generate some logs (it does every morning, so I will see tomorrow).

In any case, it's true that I didn't think about the CPU usage at all. Your regex is relly neat, it's doing basically the same thing with far less characters, I will go with something like that. The graph is also really helpful, greatly appreciated.

Thanks again! Will update when I get the final/optimized version in case someone else needs it.

---

<div class="post-metadata">

**Author:** ![Houss](https://avatars.discourse-cdn.com/v4/letter/h/a88e4f/32.png) [@Houss](https://discuss.elastic.co/u/Houss)\
**Post date:** [October 24, 2016, 4:47am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/12 "2016-10-24T04:47:29Z")

</div>

Posting the pattern here in case it can help someone (who knows):

multiline pattern:  
pattern: '^[0-9]{2}/[0-9]{2}/[0-9]{4} [0-9]{2}:[0-9]{2}:[0-9]{2} (--- FIN ETAT JOURNAL)|(Nombre total)|(Le nombre total)|(Dur.e de transfert)|(D.bit de transfert)|(Taux de compression)|(Rapport de r.duction)|(Temps de traitement)'  
negate: false  
match: after

```
include_lines pattern:
include_lines: [".+ DEBUT ETAT JOURNAL"]

```

Seeing that it wors so well, I'm not gonna change it for now. If the CPU usage becomes a problem, I will change it.

Thanks again.  
Cheers.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 7, 2016, 5:52am UTC](https://discuss.elastic.co/t/catching-all-lines-between-two-lines-two-lines-included/63150/13 "2016-11-07T05:52:41Z")

</div>

This topic was automatically closed after 21 days. New replies are no longer allowed.
