# Is it possible to index only the matched log lines of grok in Logstash?

**URL:** <https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624>\
**Category:** Logstash\
**Created:** [February 2, 2017, 4:50am UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624 "2017-02-02T04:50:18Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kulasangar\_Gowrisang](https://avatars.discourse-cdn.com/v4/letter/k/ce73a5/32.png) [@Kulasangar\_Gowrisang](https://discuss.elastic.co/u/Kulasangar_Gowrisang)\
**Post date:** [February 2, 2017, 4:50am UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624/1 "2017-02-02T04:50:18Z")

</div>

I'm having a log file, which actually has `INFO`s' and `ERROR`s. So I tried to match only the needful `INFO`s by using the **grok** filter. So [this](http://pastebin.com/HAzCRD0Z) is how my log lines look like. Few of them from the file.

And this is how my `grok` look like in my _logstash_ conf:

```
grok {
		patterns_dir => ["D:/elk_stack_for_ideabiz/elk_from_chamith/ELK_stack/logstash-5.1.1/bin/patterns"]
		match => { 
			"message" => [
				"^TID\: \[0\] \[AM\] \[%{LOGTIMESTAMPTWO:logtimestamp}]%{REQUIREDDATAFORAPP:app_message}",
				"^TID\: \[0\] \[AM\] \[%{LOGTIMESTAMPTWO:logtimestamp}]%{REQUIREDDATAFORRESPONSESTATUS:response_message}"
			] 	
		}
	}

```

The pattern seems to be working fine. I could provide the `pattern` if required.

_I've got two questions_. One is I wanted only the _grok_ matched lines to be sent to the index, and prevent `Logstash` from indexing the non-matched ones and the second is to prevent `Logstash` from showing the **message** in every single ES record.

I tried using the _overwrite_ as such under the _match_ but still no luck:

```
overwrite => ["message"]

```

All in all what I need to see in my indice are the messages (app\_message, response\_message from the above match), which should match the above two conditions. Where as now, all the lines are getting indexed.

Is it possible do something like this? Or does **Logstash** index all of them by default?

Where am I going wrong? Any help could be appreciated.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [February 2, 2017, 6:50am UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624/2 "2017-02-02T06:50:19Z")

</div>

Add `remove_field => ["message"]` to your grok filter to remove the `message` field if the filter is successful.

To drop events where grok failed:

```nohighlight
if "_grokparsefailure" in [tags] {
  drop { }
}

```

---

<div class="post-metadata">

**Author:** ![Kulasangar\_Gowrisang](https://avatars.discourse-cdn.com/v4/letter/k/ce73a5/32.png) [@Kulasangar\_Gowrisang](https://discuss.elastic.co/u/Kulasangar_Gowrisang)\
**Post date:** [February 2, 2017, 8:46am UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624/3 "2017-02-02T08:46:29Z")

</div>

Thanks @magnusbaeck for the response 🙂

Well removing the `message` did work!

But I'm still getting the unnecessary lines from the log. I might have to drop here the pattern which I'm using for each line:

The log lines and patterns respectively:

> TID: [0] [AM] [2016-12-24 23:59:59,593] INFO {org.apache.synapse.mediators.builtin.LogMediator} - API Request URL = /subscription/v3/subscribe, Request ID = urn:uuid:70938535 {org.apache.synapse.mediators.builtin.LogMediator}

Pattern for the above:  
`REQUIREDDATAFORAPP (^.*API Request URL.*$) |[^\/]*\/([^\/]*)*\/[^\/]*\/`

> TID: [0] [AM] [2016-12-24 23:59:59,213] INFO {org.apache.synapse.mediators.builtin.LogMediator} - API Response Status = 200, Request ID = urn:uuid:a83a1760, Response Time(ms) = 436.0 {org.apache.synapse.mediators.builtin.LogMediator}

Pattern for the above:

> `REQUIREDDATAFORRESPONSESTATUS (^.*API Response Status.*$) |.+?(?=\s+\S*$)`

The log line which shouldn't get matched or indexed:

> TID: [0] [AM] [2016-12-24 23:59:59,593] INFO {org.wso2.carbon.apimgt.axiata.dialog.verifier.DialogAPIRequestHandler} - [2016-12-24 23:59:59] \>\>\>\>\> API Request id 1482604199593MI584 {org.wso2.carbon.apimgt.axiata.dialog.verifier.DialogAPIRequestHandler}

I don't what I'm missing here. I only wanted to see the matched lines from the log in my indice. Is there something wrong with the **regex**?

Thanks again.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [February 2, 2017, 10:15am UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624/4 "2017-02-02T10:15:17Z")

</div>

Does the message that slipped through have a `_grokparsefailure` tag? If no, the grok filter was successful. Is it `app_message` or `response_message` that's populated with data? That'll tell us which expression that matched.

---

<div class="post-metadata">

**Author:** ![Kulasangar\_Gowrisang](https://avatars.discourse-cdn.com/v4/letter/k/ce73a5/32.png) [@Kulasangar\_Gowrisang](https://discuss.elastic.co/u/Kulasangar_Gowrisang)\
**Post date:** [February 2, 2017, 12:44pm UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624/5 "2017-02-02T12:44:48Z")

</div>

Thanks @magnusbaeck. 🙂

> [@magnusbaeck](#):
>
> Does the message that slipped through have a \_grokparsefailure tag?

No

> [@magnusbaeck](#):
>
> Is it app\_message or response\_message that's populated with data?

It wasn't either the app\_message or response\_message.

The data has been populated with the `message` field. So what I did was:

```
grok {
	patterns_dir => ["D:/elk_stack_for_ideabiz/elk_from_chamith/ELK_stack/logstash-5.1.1/bin/patterns"]
	match => [ 
				"message", "^TID\: \[0\] \[AM\] \[%{LOGTIMESTAMPTWO:logtimestamp}\]%{REQUIREDDATAFORAPP:message}",	
				"message", "^TID\: \[0\] \[AM\] \[%{LOGTIMESTAMPTWO:logtimestamp}\]%{REQUIREDDATAFORRESPONSESTATUS:message}"		
			]
	#remove_field => ["message"]
	overwrite => ["message"]
}

```

I couldn't remove the `message` since I was using the `message` for filtering purposes.

So I had to throw in a **if condition** , for all the messages which didn't match and **drop** them.

```
if "<<<<< API Request" in [message] {
	drop { }
}

```

The above method somehow satisfied my need but then has been a pain now, since I'm having too much of _ifs_ at the moment trying to match and drop the unwanted lines.

Am I going wrong somewhere? Or how can I make it more efficient?

Thanks again 🙂

---

<div class="post-metadata">

**Author:** ![Kulasangar\_Gowrisang](https://avatars.discourse-cdn.com/v4/letter/k/ce73a5/32.png) [@Kulasangar\_Gowrisang](https://discuss.elastic.co/u/Kulasangar_Gowrisang)\
**Post date:** [February 5, 2017, 10:00am UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624/6 "2017-02-05T10:00:33Z")

</div>

@magnusbaeck Is there any way I could go around this?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 5, 2017, 10:00am UTC](https://discuss.elastic.co/t/is-it-possible-to-index-only-the-matched-log-lines-of-grok-in-logstash/73624/7 "2017-03-05T10:00:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
