# Parsing human readable log files with grok

**URL:** https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821
**Category:** Logstash
**Created:** [August 4, 2015, 5:13pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821 "2015-08-04T17:13:05Z")
**Posts on this page:** 20
**Page:** 1

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 4, 2015, 5:13pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/1 "2015-08-04T17:13:05Z")

</div>

We have the following log sample file that was created in a human readable format. I'm interested in the lines in bold and want to capture the following event fields:

timestamp (i.e. 7/17/15 18:50:59:616 GMT)  
status (i.e. success or failure)

Can someone help me write the grok filter to filter out the two fields, please?

* * *

```
                              TraceLogMessage : NA: CurrencyExchangeRate: An unexpected error occured when retreiving the currency exchange file and processing it.

```

## [7/17/15 18:50:59:616 GMT] 00000688 CurrencyExcha I \*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\* **[7/17/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/17 18:50:59 failure.** [7/17/15 18:50:59:617 GMT] 00000688 CurrencyExcha I \*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\* TraceLogMessage : NA: CurrencyExchangeRate: An unexpected error occured when retreiving the currency exchange file and processing it. [7/18/15 18:50:59:616 GMT] 00000688 CurrencyExcha I \*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\* **[7/18/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/17 18:50:59 success.** [7/18/15 18:50:59:617 GMT] 00000688 CurrencyExcha I \*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\* TraceLogMessage : NA: CurrencyExchangeRate: An unexpected error occured when retreiving the currency exchange file and processing it. [7/19/15 18:50:59:616 GMT] 00000688 CurrencyExcha I \*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\* **[7/19/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/17 18:50:59 failure.** [7/19/15 18:50:59:617 GMT] 00000688 CurrencyExcha I \*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 4, 2015, 6:11pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/2 "2015-08-04T18:11:25Z")

</div>

So everything's on a single line, i.e.

```
[7/17/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/17 18:50:59 failure.

```

rather than

```
[7/17/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 
2015/07/17 18:50:59 failure.

```

or...?

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 4, 2015, 6:18pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/3 "2015-08-04T18:18:46Z")

</div>

Yes, everything is on one line, Magnus.

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 4, 2015, 6:22pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/4 "2015-08-04T18:22:06Z")

</div>

I was able to make a bit of progress on my own.  
Given Input:  
[7/17/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/17 18:50:59 failed.

Filter:  
%{DATESTAMP:tslice} %{GREEDYDATA} %{WORD:status}

Output:  
{  
"tslice": [  
[  
"7/17/15 18:50:59:616"  
]  
],  
"status": [  
[  
"failed"  
]  
]  
}

I'm not exactly sure why it worked... 🙂 Not sure how GREDYDATA decided to stop where it did.

This filter I have isn't specific enough. I want the filter to just capture the Success/Failures of the exchange rate retrieval and ignore everything else from that log. As it stands, it would parse any and every line that has a timestamp and that's not what I want.

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 4, 2015, 8:18pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/5 "2015-08-04T20:18:24Z")

</div>

> Not sure how GREDYDATA decided to stop where it did.

Because you said that there's a WORD at the end.

> This filter I have isn't specific enough. I want the filter to just capture the Success/Failures of the exchange rate retrieval and ignore everything else from that log.

This should be close enough:

```
filter {
  grok {
    match => [
      "message",
      "^\[%{DATESTAMP:tslice} GMT\] \d+ CurrencyExcha I Currency Exchange Rate File Retreival %{GREEDYDATA} %{WORD:status}$"
    ]
  }
  if "_grokparsefailure" in [tags] {
    drop { }
  }
}

```

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 5, 2015, 1:59pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/6 "2015-08-05T13:59:58Z")

</div>

Magnus - thank you very much. That works great.

Is \_grokparserfailure thrown when the parser can't extract one of the fields in the expression? What all is returned in [tags]?

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 5, 2015, 2:02pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/7 "2015-08-05T14:02:26Z")

</div>

The `_grokparsefailure` tag is added when none of the supplied expressions match the field they're supposed to match.

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 5, 2015, 2:13pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/8 "2015-08-05T14:13:10Z")

</div>

Thank you - lastly:

This config I now have is meant to feed only one Index type in ElasticSearch. We may need to feed other Indices from the same log. Do I need to create different config files and start separate LogStash instances for those? Or do I somehow include Input/Filter/Output for those other indices all in the same config file? Please -

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 5, 2015, 2:18pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/9 "2015-08-05T14:18:12Z")

</div>

Look into [conditionals](https://www.elastic.co/guide/en/logstash/current/configuration.html#conditionals).

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 5, 2015, 2:21pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/10 "2015-08-05T14:21:33Z")

</div>

Thanks again -

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 6, 2015, 5:17pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/11 "2015-08-06T17:17:17Z")

</div>

My logstash config file is working now, I'm parsing the input log file and extracting the right data, all good here. The event looks good in the RUBYDEBUG from logstash. The issue I'm having is that I started redirecting my extracted events to Elasticsearch and now event structure that I get into ES is different than what I defined it using \_Mappings definition. Somehow logstash is overriding the document properties.

Here is what it looks like in logstash console (all good):

```
{
       "message" => "[7/30/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/30 18:50:59 success",
      "@version" => "1",
    "@timestamp" => "2015-08-06T16:13:38.100Z",
          "host" => "IBM-EN189AKEUJ4",
          "path" => "C:/logstash-1.5.3/inputs/SystemOut1.log",
          "type" => "sysout",
        "tslice" => "7/30/15 18:50:59:616",
        "status" => "success"
}

```

When it gets to ES it looks terrible - Logstash stuffs everything into the [\_source]{message} array, seemingly ignoring the Mapping definition in ES for index 'sprint'... The doc looks like this in ES:

```
curl -XGET 'localhost:9200/sprint/sysout/AU8DyhBIErJTJM_bP4rF?pretty'
  % Total % Received % Xferd Average Speed Time Time Time Current
                                 Dload Upload Total Spent Left Speed
100 471 100 471 0 0 31400 0 --:--:-- --:--:-- --:--:-- 31400{
  "_index" : "sprint",
  "_type" : "sysout",
  "_id" : "AU8DyhBIErJTJM_bP4rF",
  "_version" : 1,
  "found" : true,
  "_source":{"message":"[7/30/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/30 18:50:59 success","@version":"1","@timestamp":"2015-08-06T16:13:38.100Z","host":"IBM-EN189AKEUJ4","path":"C:/logstash-1.5.3/inputs/SystemOut1.log","type":"sysout","tslice":"7/30/15 18:50:59:616","status":"success"}
}

```

Here is my \_Mapping that I've created for 'sprint' index are as follows:

```
 curl -XGET "localhost:9200/sprint/_mappings?pretty"
  % Total % Received % Xferd Average Speed Time Time Time Current
                                 Dload Upload Total Spent Left Speed
100 760 100 760 0 0 47500 0 --:--:-- --:--:-- --:--:-- 47500{
  "sprint" : {
    "mappings" : {
      "sysout" : {
        "properties" : {
          "@timestamp" : {
            "type" : "date",
            "format" : "dateOptionalTime"
          },
          "@version" : {
            "type" : "string"
          },
          "host" : {
            "type" : "string"
          },
          "message" : {
            "type" : "string"
          },
          "path" : {
            "type" : "string"
          },
          "status" : {
            "type" : "string"
          },
          "tslice" : {
            "type" : "date",
            "format" : "MM/dd/yy HH:mm:ss:SSS"
          },
          "type" : {
            "type" : "string"
          }
        }
      }
    }
  }
}

```

I should also paste the current config file as it exists right now:

```
input { 
    # Log file
    file {
       path => "C:/logstash-1.5.3/inputs/SystemOut1.log"
       type => "sysout"
       #start_position => "beginning"
      }
      
      # Standard input file
      #stdin { }
}  

filter {
  grok {
    match => [
      "message",
      "^\[%{DATESTAMP:tslice} GMT\] %{GREEDYDATA} (%{WORD:status}(\.)?)$"
    ]
  }
  if "_grokparsefailure" in [tags] {
    drop { }
  }
}
    
output { 
    elasticsearch {
        host => localhost
        index => "sprint"
    }
    stdout { 
        codec => rubydebug 
    } 
}

```

Why can't I get the individual fields in sprint/sysout to get populatedproperly? All I really care about in that Document are fields tslice and status...

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 6, 2015, 5:36pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/12 "2015-08-06T17:36:19Z")

</div>

Everything is working fine. When you use the [Get API](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-get.html) the top-level properties of the returned document contain metadata (`_id`, `_type`, ...) and the document itself is found in the `_source` property. See the example in the documentation.

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 6, 2015, 6:49pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/13 "2015-08-06T18:49:04Z")

</div>

I thought I was starting to understand this product....but I was wrong.

I initially created the sprint index by issuing the following command:  
curl -XPUT 'localhost:9200/sprint/sysout/1?pretty' -d '  
{  
"tslice": "15/07/10 07:04:36:321",  
"status": "failed"  
}'

I didn't even specify message in my document definition properties initially (it was created dynamically by logstash along with these other fields). I did have tslice and status fields in the document at the root level, but they were not populated?! I don't get it.

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 6, 2015, 6:53pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/14 "2015-08-06T18:53:49Z")

</div>

Why are you saying that the `tslice` and `status` fields are missing? They're right there, on the line that begins with "\_source", near the end.

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 6, 2015, 6:55pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/15 "2015-08-06T18:55:27Z")

</div>

I guess I expected them to be present at the document root, just where I defined them:

"sprint" : {  
"mappings" : {  
"sysout" : {  
"properties" : {  
"@timestamp" : {  
"type" : "date",  
"format" : "dateOptionalTime"  
},  
"@version" : {  
"type" : "string"  
},  
"host" : {  
"type" : "string"  
},  
"message" : {  
"type" : "string"  
},  
"path" : {  
"type" : "string"  
},  
" **status**" : {  
"type" : "string"  
},  
" **tslice**" : {  
"type" : "date",  
"format" : "MM/dd/yy HH:mm:ss:SSS"  
},  
"type" : {  
"type" : "string"  
}

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 6, 2015, 6:58pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/16 "2015-08-06T18:58:27Z")

</div>

But they _are_ in the document root. It's the Get API that returns the original document in the `_source` field. If you run the contents of the `_source` field through a JSON pretty printer it looks like this:

```
{
    "message": "[7/30/15 18:50:59:616 GMT] 00000688 CurrencyExcha I Currency Exchange Rate File Retreival at 2015/07/30 18:50:59 success",
    "@version": "1",
    "@timestamp": "2015-08-06T16:13:38.100Z",
    "host": "IBM-EN189AKEUJ4",
    "path": "C:/logstash-1.5.3/inputs/SystemOut1.log",
    "type": "sysout",
    "tslice": "7/30/15 18:50:59:616",
    "status": "success"
}

```

_That's_ your original, the source document.

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 6, 2015, 7:01pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/17 "2015-08-06T19:01:31Z")

</div>

Huh... I see that now - sorry, I was being difficult. Is there a way to run contents of \_source through JSON pretty print in ES?! Or do you do that in something like JSONLint?

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [August 6, 2015, 7:13pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/18 "2015-08-06T19:13:59Z")

</div>

In this case I used [jsonlint.com](http://jsonlint.com/). Since you supplied `?pretty` in the URL I would've expected ES to have pretty printed the document as well.

---

<div class="post-metadata">

### Author: ![jjdepaul](https://avatars.discourse-cdn.com/v4/letter/j/e0b2c6/32.png) [@jjdepaul](https://discuss.elastic.co/u/jjdepaul)
#### Post date: [August 6, 2015, 9:05pm UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/19 "2015-08-06T21:05:20Z")

</div>

Ok - thanks for your patience and advice.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 5:32am UTC](https://discuss.elastic.co/t/parsing-human-readable-log-files-with-grok/26821/20 "2017-07-06T05:32:42Z")

</div>


