# COMBINEDAPACHELOG Outliers

**URL:** https://discuss.elastic.co/t/combinedapachelog-outliers/151928
**Category:** Logstash
**Created:** [October 10, 2018, 7:53pm UTC](https://discuss.elastic.co/t/combinedapachelog-outliers/151928 "2018-10-10T19:53:35Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![pooch](https://avatars.discourse-cdn.com/v4/letter/p/dbc845/32.png) [@pooch](https://discuss.elastic.co/u/pooch)
#### Post date: [October 10, 2018, 7:53pm UTC](https://discuss.elastic.co/t/combinedapachelog-outliers/151928/1 "2018-10-10T19:53:36Z")

</div>

Hi All,

So I have the following GROK filter going for my Apache logs.

```auto
else if "apache" in [tags] {
    grok {
      match => { "message" => '%{COMBINEDAPACHELOG}
    }

```

This is working for ~75% of ingested logs but I seem to have some Apache log files that have slightly different formats. Some contain a port number next to hostname (hostname:8080) and some have random characters sprinkled in for some reason ('/'). Examples below with the odd message being the first and the characters causing grok failures in **BOLD**.

```auto
10.101.76.157 10.101.76.157 vdpswebdev01.qualcomm.com **:8030** - [10/Oct/2018:11:45:00 -0700] **\**"GET /solr-chipcode/collection1/replication?command=indexversion&wt=javabin&qt=%2Freplication&version=2 HTTP/1.1\" 200 76 **\**"- **\**" **\**"Solr[org.apache.solr.client.solrj.impl.HttpSolrServer] 1.0 **\**"

10.101.76.147 10.101.76.147 vdpswebdev01.qualcomm.com - [09/Oct/2018:13:45:55 -0700] "GET /solr-cpip-global/cpip-directory-path/replication?command=indexversion&wt=javabin&qt=%2Freplication&version=2 HTTP/1.1" 200 80 "-" "Solr[org.apache.solr.client.solrj.impl.HttpSolrClient] 1.0"

172.30.42.114, 10.49.16.6 10.49.16.6 prismsearch-dev.qualcomm.com - [09/Oct/2018:13:52:09 -0700] "GET /solr-prism/collection1/select?q=*%3A*&wt=json&indent=true HTTP/1.1" 200 4200 "-" "Mozilla/4.0 (compatible; MSIE 4.01; Windows NT)"

```

My question is do I need to create customized filters to capture every log line variation or is there a more sane approach?

Cheers!

---

<div class="post-metadata">

### Author: ![pooch](https://avatars.discourse-cdn.com/v4/letter/p/dbc845/32.png) [@pooch](https://discuss.elastic.co/u/pooch)
#### Post date: [October 10, 2018, 8:00pm UTC](https://discuss.elastic.co/t/combinedapachelog-outliers/151928/2 "2018-10-10T20:00:40Z")

</div>

Note the characters that are causing the parsing to choke are marked with \*\*.

---

<div class="post-metadata">

### Author: ![pooch](https://avatars.discourse-cdn.com/v4/letter/p/dbc845/32.png) [@pooch](https://discuss.elastic.co/u/pooch)
#### Post date: [October 12, 2018, 6:28pm UTC](https://discuss.elastic.co/t/combinedapachelog-outliers/151928/3 "2018-10-12T18:28:25Z")

</div>

Is the simple answer, yes I need to create customized grok filters for every log line variation?

---

<div class="post-metadata">

### Author: ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)
#### Post date: [October 12, 2018, 6:36pm UTC](https://discuss.elastic.co/t/combinedapachelog-outliers/151928/4 "2018-10-12T18:36:39Z")

</div>

Yes. You will need to create patterns that match the variations as field are only extracted on successful match.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [November 9, 2018, 6:39pm UTC](https://discuss.elastic.co/t/combinedapachelog-outliers/151928/5 "2018-11-09T18:39:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
