# Grok for email log

**URL:** <https://discuss.elastic.co/t/grok-for-email-log/171750>\
**Category:** Logstash\
**Created:** [March 11, 2019, 11:55am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750 "2019-03-11T11:55:41Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [March 11, 2019, 11:55am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/1 "2019-03-11T11:55:42Z")

</div>

Can you please help me with building a grok pattern for SMTP log:  
Example of my smtp log:

Return-Path: [alok@XXXXXX.in](mailto:alok@XXXXXX.in)  
Delivered-To: bigbrother@XXXX.in  
Received: from n-mail-1.xxxxx.in  
by n-mail-1.xxxxx.inwith LMTP id uOV8CFLmalxdRQAADPfE8A  
for [bigbrother@xxxxx.in](mailto:bigbrother@xxxxx.in); Mon, 18 Feb 2019 22:37:30 +0530  
Received: from Spamfilter-3.xxxx.in (Spamfilter-3.rrcat.gov.in [10.11.108.104])  
by n-mail-1.rrcat.gov.in (Postfix) with ESMTP id 1B245263926;  
Mon, 18 Feb 2019 22:37:30 +0530 (IST)  
Received: from spamfilter-3.xxxx.in (localhost [127.0.0.1])  
by Spamfilter-3.xxxx.in (Postfix) with ESMTP id 04FE12DA1D2;  
Mon, 18 Feb 2019 22:37:30 +0530 (IST)  
X-Virus-Scanned: amavisd-new at rrcat.xxxx.in  
X-Spam-Flag: NO  
X-Spam-Score: 1.163  
X-Spam-Level: \*  
X-Spam-Status: No, score=1.163 tagged\_above=-999 required=6  
tests=[ALL\_TRUSTED=-1, DEAR\_SOMETHING=1.731, INVALID\_DATE=0.432]  
autolearn=no autolearn\_force=no  
Received: from Spamfilter-3.xxxx.in ([127.0.0.1])  
by spamfilter-3.xxxx.in (spamfilter-3.xxxx.in [127.0.0.1]) (amavisd-new, port 10024)  
with ESMTP id bYziy9Hx1yoW; Mon, 18 Feb 2019 22:37:29 +0530 (IST)  
Received: from localhost (localhost [127.0.0.1])  
by Spamfilter-3.rrcat.gov.in (Postfix) with ESMTP id 900F22D6BD6;  
Mon, 18 Feb 2019 22:37:29 +0530 (IST)  
From: noreply@xxxx.in  
To: [mkonline1996@gmail.com](mailto:mkonline1996@gmail.com)  
Date: Mon, 18 Feb 19 17:05:17 +0000  
Subject: Mail from Trade Apprenticeship Scheme at xxxxx (TASAR-2019)  
Message-Id: [20190218170729.900F22D6BD6@Spamfilter-3.xxxx.in](mailto:20190218170729.900F22D6BD6@Spamfilter-3.xxxx.in)

Dear Candidate,

Your password for Online Application Submission for the Apprenticeship Program of trade Electrician against  
Advertisement No.xxxxxxxx and trade code I-8 is japupu3ys

Login name is same as your Email ID.

* * *

## This is system generated mail, please do not reply to it.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 11, 2019, 12:40pm UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/2 "2019-03-11T12:40:07Z")

</div>

For each header that you want to extract use a pattern that matches zero-or-more characters that are not newline followed by a newline.

```
    grok { match => { message => "^Header: (?<Header>[^
]*)
" } }

```

Then you can give the grok filter an array of patterns to match.

---

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [March 12, 2019, 4:22am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/3 "2019-03-12T04:22:49Z")

</div>

Please elaborate little bit more.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 12, 2019, 1:14pm UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/4 "2019-03-12T13:14:38Z")

</div>

I assume that you consume the entire log as a single event using multiline codec. If you are consuming it one line at a time things will be a little different. Decide which headers you want to parse and use the pattern I described to extract them. For example

```
    grok {
        break_on_match => false
        match => {
            message => [
                "^Date: (?<Date>[^
]*)
",
                "^Subject: (?<Subject>[^
]*)
",
                "^X-Spam-Level: (?<X-Spam-Level>[^
]*)
",
                "^X-Spam-Status: (?<X-Spam-Status>[^
]*)
"
            ]
        }
    }

```

Once you have extracted them you may require additional parsing to extract fields from those headers.

For Received you cannot do it using grok because there are multiple occurences, however, it is simple to do a similar regex match in ruby

```
    ruby {
        code => 'event.set("Received", event.get("message").scan(/^Received: ([^
]+
)/).flatten)'
    }

```

That will set Received to an array of 5 strings which contain the contents of the Received headers.

---

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [March 15, 2019, 9:10am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/5 "2019-03-15T09:10:09Z")

</div>

> [@Badger](#):
>
> ["^Date: (?\<Date\>[^]_) ", "^Subject: (?\<Subject\>[^]_) ", "^X-Spam-Level: (?\<X-Spam-Level\>[^]_) ", "^X-Spam-Status: (?\<X-Spam-Status\>[^]_) " ]

actually I am newborn for elk and I am still unable to understand how to use multiline filter

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 15, 2019, 1:53pm UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/6 "2019-03-15T13:53:02Z")

</div>

> [@Vikash\_Singh1](#):
>
> actually I am newborn for elk and I am still unable to understand how to use multiline filter

How are you ingesting the log?

---

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [March 15, 2019, 2:37pm UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/7 "2019-03-15T14:37:55Z")

</div>

via FileBeat

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 15, 2019, 3:03pm UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/8 "2019-03-15T15:03:57Z")

</div>

Do you want to ingest the entire file as a single event?

---

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [March 17, 2019, 9:41am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/9 "2019-03-17T09:41:10Z")

</div>

Yes!!

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 17, 2019, 4:25pm UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/10 "2019-03-17T16:25:23Z")

</div>

If filebeat works the same way as a multiline codec, then you should be able to do that by specifying a pattern that never matches

```
multiline.pattern: '^Spalanzani'
multiline.negate: true
multiline.match: after
```

---

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [March 29, 2019, 5:43am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/11 "2019-03-29T05:43:56Z")

</div>

Is there any other way to segregate multi lines logs other than groks?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 29, 2019, 12:24pm UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/12 "2019-03-29T12:24:03Z")

</div>

You can use other filters. What are you trying to do and what problem are you having using grok?

---

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [April 1, 2019, 4:18am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/13 "2019-04-01T04:18:47Z")

</div>

Other filters!!! Like? I am not able define grok

---

<div class="post-metadata">

**Author:** ![Vikash\_Singh1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikash_singh1/32/42119_2.png) [@Vikash\_Singh1](https://discuss.elastic.co/u/Vikash_Singh1)\
**Post date:** [April 1, 2019, 4:35am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/14 "2019-04-01T04:35:05Z")

</div>

Since the mail log is unstructured how can I define one common filter which would read all the logs correctly?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 29, 2019, 4:35am UTC](https://discuss.elastic.co/t/grok-for-email-log/171750/15 "2019-04-29T04:35:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
