# Multiple grok parsing not extracting all the fields

**URL:** <https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764>\
**Category:** Logstash\
**Created:** [January 4, 2020, 1:45am UTC](https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764 "2020-01-04T01:45:46Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![nmoham](https://avatars.discourse-cdn.com/v4/letter/n/77aa72/32.png) [@nmoham](https://discuss.elastic.co/u/nmoham)\
**Post date:** [January 4, 2020, 1:45am UTC](https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764/1 "2020-01-04T01:45:46Z")

</div>

Trying to extract some fields from the msgbody field using grok , but only the first field in the gork gets extracted.

interested Fields - corId, controller, httpStatusText and uri

Sample Data -

2020-01-03 10:44:17,025 [93] ERROR MedServFileLogger corId=cf25b00d-1e37-4eb7-ab75-82ceeec7fdab - Exception **controller= Loan** action= Getmethod= GET **uri= [http://xxxxxxxxxx/v2/media/instance/xxxx/loans/cdb79433-32fa-4df8-b73a-e87aa89f2007/files/images-178ee8d0-fa48-4b9f-a8df-abcc9cfb1ac7.zip/entries/0b3e99f8-8af8-49a5-95b1-1537c715eb43.png?tokencreator=Encompass&tokenexpires=1578076775&token=pLCvT%2F1pBPhuFXiHKDIlB5F9feocqeq7Wxx%2FyhAz7B6DCcKeOP3YjO%2FnalfjTgXdieAmyFHEiW72Soym14oBuw%3D%3D](http://xxxxxxxxxx/v2/media/instance/xxxx/loans/cdb79433-32fa-4df8-b73a-e87aa89f2007/files/images-178ee8d0-fa48-4b9f-a8df-abcc9cfb1ac7.zip/entries/0b3e99f8-8af8-49a5-95b1-1537c715eb43.png?tokencreator=Encompass&tokenexpires=1578076775&token=pLCvT%2f1pBPhuFXiHKDIlB5F9feocqeq7Wxx%2fyhAz7B6DCcKeOP3YjO%2fnalfjTgXdieAmyFHEiW72Soym14oBuw%3d%3d)**  
System.UnauthorizedAccessException: MediaTokenInvalid - A valid Token must be provided for accessing Media

2020-01-03 03:58:12,822 [37] ERROR MedServFileLogger corId=5aa9b90b-9fe6-4700-aa8f-c08be2d3f0ea - Returning controller= Health action= Getmethod= GET **uri** = [http://localhost/v2/media/healthhttpStatusCode=503](http://localhost/v2/media/healthhttpStatusCode=503) **httpStatusText** =ServiceUnavailable

Logstash Filter -

filter {

if [project] == "media\_server"  
{  
grok {  
match =\> ["message", "(?m)%{TIMESTAMP\_ISO8601:logtime} [(?[\d.]+)] +%{LOGLEVEL:loglevel} %{GREEDYDATA:msgbody}" ]  
}

```
     grok {
            match => {
            break_on_match => "false"
            "msgbody" => ["corId=%{UUID:corId}", "controller=%{SPACE}%{WORD:controller}", "httpStatusText=%{WORD:httpStatusText}", "uri=%{SPACE}%{URI:uri}"]
            }
    }

       date {
        locale => "en"
        match => ["logtime", "YYYY-MM-dd HH:mm:ss,SSS"]
        timezone => "America/Los_Angeles"
        target => "@timestamp"

    }
    mutate
    {
            remove_field => ["msgbody"]
    }

```

}  
}

Using the above configuration, only the corId field is getting extracted and all other fields are dropped. I don't see any parsing errors/failures in the logstash logs.

Appreciate any help or guidance with this.

Thanks

---

<div class="post-metadata">

**Author:** ![mancharagopan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mancharagopan/32/60266_2.png) [@mancharagopan](https://discuss.elastic.co/u/mancharagopan)\
**Post date:** [January 4, 2020, 5:25am UTC](https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764/2 "2020-01-04T05:25:10Z")

</div>

I have tested below sample data,

> [@nmoham](#):
>
> 2020-01-03 03:58:12,822 [37] ERROR MedServFileLogger corId=5aa9b90b-9fe6-4700-aa8f-c08be2d3f0ea - Returning controller= Health action= Getmethod= GET **uri** = [http://localhost/v2/media/healthhttpStatusCode=503](http://localhost/v2/media/healthhttpStatusCode=503) **httpStatusText** =ServiceUnavailable

This grok filter works for me,

```
(?m)%{TIMESTAMP_ISO8601:logtime}%{SPACE}\[%{INT:num}\]%{SPACE}%{LOGLEVEL:log_level}%{SPACE}%{WORD:logger}%{SPACE}corId=(?<corId>[A-Za-z0-9-]+)%{SPACE}-%{SPACE}%{WORD:field1}%{SPACE}controller=%{SPACE}%{WORD:Controller}%{SPACE}action=%{SPACE}Getmethod=%{SPACE}%{WORD:get_method}%{SPACE}uri=%{SPACE}%{URI:url}%{SPACE}httpStatusText=%{WORD:http_Status}

```

![image](https://us1.discourse-cdn.com/elastic/original/3X/c/4/c4eb5127e90774ca9dc829098c7f795a5e367b6e.png)

There is a grok debugger available with kibana. you can use your sample data there and build grok filter that matches.

---

<div class="post-metadata">

**Author:** ![nmoham](https://avatars.discourse-cdn.com/v4/letter/n/77aa72/32.png) [@nmoham](https://discuss.elastic.co/u/nmoham)\
**Post date:** [January 6, 2020, 4:26pm UTC](https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764/3 "2020-01-06T16:26:16Z")

</div>

Thanks @mancharagopan,

The log message is not consistent, if you see the first example, there's stack trace followed after the "url" field and no"httpStatusText". Similarly, some of the logs do not contain the "httpStatusText" field at all.

Hence, I am trying to have everything in the "msgbody" field, followed by using multiple match to extract only other required fields from the body.

[https://www.elastic.co/guide/en/logstash/5.5/plugins-filters-grok.html#plugins-filters-grok-match](https://www.elastic.co/guide/en/logstash/5.5/plugins-filters-grok.html#plugins-filters-grok-match)

---

<div class="post-metadata">

**Author:** ![mancharagopan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mancharagopan/32/60266_2.png) [@mancharagopan](https://discuss.elastic.co/u/mancharagopan)\
**Post date:** [January 7, 2020, 2:27am UTC](https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764/4 "2020-01-07T02:27:03Z")

</div>

@nmoham  
You can use `%{GREEDYDATA:msgbody`} after `%{URI:url}%{SPACE}` to have all the trails into `msgbody` field then you can use if conditions to match data inside `msgbody` to apply different patterns.

---

<div class="post-metadata">

**Author:** ![nmoham](https://avatars.discourse-cdn.com/v4/letter/n/77aa72/32.png) [@nmoham](https://discuss.elastic.co/u/nmoham)\
**Post date:** [January 7, 2020, 2:28am UTC](https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764/5 "2020-01-07T02:28:51Z")

</div>

hi @mancharagopan,

got it working , by including the following "patterns\_dir" and also "break\_on\_match" inside the grok filter and before the "match" stanza.

```auto
            patterns_dir => "/etc/logstash/patterns"
            break_on_match => "false"

```

Working Filter -

```auto
filter {

   if [project] == "media_server"
        {
        grok {
            match => ["message", "(?m)%{TIMESTAMP_ISO8601:logtime} \[(?<threadid>[\d.]+)\] +%{LOGLEVEL:loglevel} %{GREEDYDATA:msgbody}" ]
        }

         grok {
                patterns_dir => "/etc/logstash/patterns"
                break_on_match => "false"
                match => {
                "msgbody" => ["corId=%{UUID:corId}", "controller=%{SPACE}%{WORD:controller}", "httpStatusText=%{WORD:httpStatusText}", "uri=%{SPACE}%{URI:uri}"]
                }
        }

           date {
            locale => "en"
            match => ["logtime", "YYYY-MM-dd HH:mm:ss,SSS"]
            timezone => "America/Los_Angeles"
            target => "@timestamp"

        }
        mutate
        {
                remove_field => ["msgbody"]
        }
  }
}

```

Thank you so much for your valuable input.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 4, 2020, 2:29am UTC](https://discuss.elastic.co/t/multiple-grok-parsing-not-extracting-all-the-fields/213764/6 "2020-02-04T02:29:12Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
