# Help to refine a grok parser

**URL:** <https://discuss.elastic.co/t/help-to-refine-a-grok-parser/42671>\
**Category:** Logstash\
**Created:** [February 25, 2016, 6:31am UTC](https://discuss.elastic.co/t/help-to-refine-a-grok-parser/42671 "2016-02-25T06:31:49Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![nzHillNet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nzhillnet/32/3423_2.png) [@nzHillNet](https://discuss.elastic.co/u/nzHillNet)\
**Post date:** [February 25, 2016, 6:31am UTC](https://discuss.elastic.co/t/help-to-refine-a-grok-parser/42671/1 "2016-02-25T06:31:49Z")

</div>

Hi all,

Logstash v2.2.2

I'm not that strong with grok and regex's, but have managed to create a working grok parser as follows:

Log excerpt:  
`2016-02-25 19:06:23 msg 3/3 (4575 bytes) msgid 0000178e55c5dc45 from <observium-bounces@observium.org> delivered to MDA_external command procmail (), deleted`

Grok parser (part of logstash filter):  
`%{TIMESTAMP_ISO8601:syslog_timestamp} msg {1,2}%{NUMBER:mess_num}\/%{NUMBER:mess_count} \(%{NUMBER:mess_bytes} bytes\) msgid %{MSGID:mess_msgid} from %{EMAILADDRESS:mess_from} %{GREEDYDATA:syslog_message}`

Patterns used:  
`MSGID [a-zA-Z0-9_.+-=:]+ EMAILADDRESSPART [a-zA-Z0-9_.+-=:]+ EMAILADDRESS \<%{EMAILADDRESSPART:email_local}@%{EMAILADDRESSPART:email_remote}\>`

If anybody has the time I'd like to know if/how this could be improved please just to aid my learning.

Many thanks,

--  
Roland

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [February 26, 2016, 9:32am UTC](https://discuss.elastic.co/t/help-to-refine-a-grok-parser/42671/2 "2016-02-26T09:32:53Z")

</div>

Looks reasonable. A few comments:

- I wouldn't consider the angle brackets to be part of the email address.
- EMAILADDRESSPART will definitely work for _most_ email addresses, but not all. How strict does this expression need to be, really? I would've just used something like this:

```nohighlight
 ... from <(?<mess_from>[^>]+)>

```

---

<div class="post-metadata">

**Author:** ![nzHillNet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nzhillnet/32/3423_2.png) [@nzHillNet](https://discuss.elastic.co/u/nzHillNet)\
**Post date:** [February 27, 2016, 11:23pm UTC](https://discuss.elastic.co/t/help-to-refine-a-grok-parser/42671/3 "2016-02-27T23:23:27Z")

</div>

Thanks for the feedback Magnus.

I'll play around with your suggestions. In terms of the email address extraction, I'm not set on how this should be done and copied something I saw elsewhere to get what I currently have.

My system is just a home one which I use for teaching myself stuff, so my needs change as I learn :-).

Thanks again.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:09am UTC](https://discuss.elastic.co/t/help-to-refine-a-grok-parser/42671/4 "2017-07-06T05:09:30Z")

</div>


