# Logstash configuration for Cloudfront logs

**URL:** https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660
**Category:** Logstash
**Created:** [March 29, 2016, 8:28am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660 "2016-03-29T08:28:39Z")
**Posts on this page:** 9
**Page:** 1

<div class="post-metadata">

### Author: ![manishr](https://avatars.discourse-cdn.com/v4/letter/m/c57346/32.png) [@manishr](https://discuss.elastic.co/u/manishr)
#### Post date: [March 29, 2016, 8:28am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/1 "2016-03-29T08:28:40Z")

</div>

Hi Guys,

Help me getting cloudfront logs parsed in logstash. I want each filed to be searchable even parameters like aid, bid, cid etc. See below

# Sample cloudfront log

2016-03-29 04:02:08 ABC1 461 22.20.17.8 GET [afsaGdhfxghxgh.cloudfront.net](http://afsaGdhfxghxgh.cloudfront.net) /1.gif - Mozilla/5.0%2520(Linux;%2520Android%25205.1.1;%2520SM-G920I%2520Build/LMY47X;%2520wv)%2520AppleWebKit/537.36%2520(KHTML,%2520like%2520Gecko)%2520Version/4.0%2520Chrome/48.0.2564.106%2520Mobile%2520Safari/537.36 aid=fsdggg25346&bid=fsdgagsexfdhg&cid=1423690744601076&cb=fdfsdggg&did=fsagdsgg&eid=fDSGzsgdfhdsh - Miss jAk9duSOoOPVfssDGZdfhgxxxghfghpzK35tRuujwuQ== [afsaGdhfxghxgh.cloudfront.net](http://afsaGdhfxghxgh.cloudfront.net) https 558 0.715 - TLSv1.2 ECDHE-RSA-AES128-GCM-SHA256 Miss

# Working one but I want parameters as well to be searchable

match =\> { "message" =\> "%{DATE\_EU:date}\t%{TIME:time}\t%{WORD:x\_edge\_location}\t(?:%{NUMBER:sc\_bytes}|-)\t%{IPORHOST:c\_ip}\t%{WORD:cs\_method}\t%{HOSTNAME:cs\_host}\t%{NOTSPACE:cs\_uri\_stem}\t%{NUMBER:sc\_status}\t%{GREEDYDATA:referrer}\t%{GREEDYDATA:User\_Agent}\t%{GREEDYDATA:cs\_uri\_stem}\t%{GREEDYDATA:cookies}\t%{WORD:x\_edge\_result\_type}\t%{NOTSPACE:x\_edge\_request\_id}\t%{HOSTNAME:x\_host\_header}\t%{URIPROTO:cs\_protocol}\t%{INT:cs\_bytes}\t%{GREEDYDATA:time\_taken}\t%{GREEDYDATA:x\_forwarded\_for}\t%{GREEDYDATA:ssl\_protocol}\t%{GREEDYDATA:ssl\_cipher}\t%{GREEDYDATA:x\_edge\_response\_result\_type}" }

# Not working

match =\> { "message" =\> "%{DATE\_EU:date}\t%{TIME:time}\t%{WORD:x\_edge\_location}\t(?:%{NUMBER:sc\_bytes}|-)\t%{IPORHOST:c\_ip}\t%{WORD:cs\_method}\t%{HOSTNAME:cs\_host}\t%{NOTSPACE:cs\_uri\_stem}\t%{NUMBER:sc\_status}\t%{GREEDYDATA:referrer}\t%{GREEDYDATA:User\_Agent}\t(?[A-Za-z0-9$.+!_'|(){},~@#%&/=:;\_?-[]\<\>^`]_)?)?)\t%{GREEDYDATA:cookies}\t%{WORD:x\_edge\_result\_type}\t%{NOTSPACE:x\_edge\_request\_id}\t%{HOSTNAME:x\_host\_header}\t%{URIPROTO:cs\_protocol}\t%{INT:cs\_bytes}\t%{GREEDYDATA:time\_taken}\t%{GREEDYDATA:x\_forwarded\_for}\t%{GREEDYDATA:ssl\_protocol}\t%{GREEDYDATA:ssl\_cipher}\t%{GREEDYDATA:x\_edge\_response\_result\_type}" }

# Logstash Configuration

input {  
file {  
path =\> "/opt/cloudfront/E2I53NO2J8KEJZ\*"  
type =\> "cloudfront"  
start\_position =\> "beginning"  
sincedb\_path =\> "log\_sincedb"  
}  
}

filter {  
if [type] == "cloudfront" {  
if ( ("#Version: 1.0" in [message]) or ("#Fields: date" in [message])) {  
drop {}  
}

```
            grok {

                                            match => { "message" => "%{DATE_EU:date}\t%{TIME:time}\t%{WORD:x_edge_location}\t(?:%{NUMBER:sc_bytes}|-)\t%{IPORHOST:c_ip}\t%{WORD:cs_method}\t%{HOSTNAME:cs_host}\t%{NOTSPACE:cs_uri_stem}\t%{NUMBER:sc_status}\t%{GREEDYDATA:referrer}\t%{GREEDYDATA:User_Agent}\t(<params>\?[A-Za-z0-9$.+!*'|(){},~@#%&/=:;_?\-\[\]<>\^\`]*)?)?)\t%{GREEDYDATA:cookies}\t%{WORD:x_edge_result_type}\t%{NOTSPACE:x_edge_request_id}\t%{HOSTNAME:x_host_header}\t%{URIPROTO:cs_protocol}\t%{INT:cs_bytes}\t%{GREEDYDATA:time_taken}\t%{GREEDYDATA:x_forwarded_for}\t%{GREEDYDATA:ssl_protocol}\t%{GREEDYDATA:ssl_cipher}\t%{GREEDYDATA:x_edge_response_result_type}" }

            }
    }
            mutate {
                    add_field => ["received_at", "%{@timestamp}"]
                    add_field => ["listener_timestamp", "%{date} %{time}"]
            }

            date {
                    match => ["listener_timestamp", "yy-MM-dd HH:mm:ss"]
            }
                    if [params] {
            mutate {
                    rename => { "params" => "params[request]" }
            }
            urldecode {
                    field => "params[request]"
            }
            kv {
                    source => "params[request]"
                    field_split => "?&"
                    target => "params"
            }
    ruby {
            code => "
            arguments = Array.new
            event['params'].to_hash.each {|k,v|
    if k == 'request' then
    next
  end
  arguments << { 'key' => k, 'value' => v }
}
unless arguments.empty?
  event['[arguments]'] = arguments
end

```

"  
remove\_field =\> ["params"]  
}

```
    }

```

}

output {  
stdout { codec =\> rubydebug }

}

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [March 29, 2016, 8:46am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/2 "2016-03-29T08:46:37Z")

</div>

So it's the `cs_uri_stem` field that contains the data you want to parse further? Keep the origin expression that works and use a kv filter to parse `cs_uri_stem`.

And try to avoid having multiple GREEDYDATA patterns in the same expression. It might seem to work but it can easily blow up later. If the fields are tab-delimited why not use the csv filter to extract the fields instead of grok?

---

<div class="post-metadata">

### Author: ![manishr](https://avatars.discourse-cdn.com/v4/letter/m/c57346/32.png) [@manishr](https://discuss.elastic.co/u/manishr)
#### Post date: [March 29, 2016, 10:07am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/3 "2016-03-29T10:07:06Z")

</div>

Yes cs\_uri\_stem contains the data that I want to parse but if you look my above configuration you will find that there are two cs\_uri\_stem. so I changed one cs\_uri\_stem to cs\_uri and used kv as suggested by you as below but after the logs got loaded to elasticsearch I am not able search the parameters for example aid="xxxxxxxx" AND bid="yyyyyyyyyyy"

input {  
file {  
path =\> "/opt/cloudfront/E2I53NO2J8KEJZ\*"  
type =\> "cloudfront"  
start\_position =\> "beginning"  
sincedb\_path =\> "log\_sincedb"  
}  
}

filter {  
if [type] == "cloudfront" {  
if ( ("#Version: 1.0" in [message]) or ("#Fields: date" in [message])) {  
drop {}  
}

```
            grok {
                    match => { "message" => "%{DATE_EU:date}\t%{TIME:time}\t%{WORD:x_edge_location}\t(?:%{NUMBER:sc_bytes}|-)\t%{IPORHOST:c_ip}\t%{WORD:cs_method}\t%{HOSTNAME:cs_host}\t%{NOTSPACE:cs_uri}\t%{NUMBER:sc_status}\t%{GREEDYDATA:referrer}\t%{GREEDYDATA:User_Agent}\t%{GREEDYDATA:cs_uri_stem}\t%{GREEDYDATA:cookies}\t%{WORD:x_edge_result_type}\t%{NOTSPACE:x_edge_request_id}\t%{HOSTNAME:x_host_header}\t%{URIPROTO:cs_protocol}\t%{INT:cs_bytes}\t%{GREEDYDATA:time_taken}\t%{GREEDYDATA:x_forwarded_for}\t%{GREEDYDATA:ssl_protocol}\t%{GREEDYDATA:ssl_cipher}\t%{GREEDYDATA:x_edge_response_result_type}" }
            }
    }
            mutate {
                    add_field => ["received_at", "%{@timestamp}"]
                    add_field => ["listener_timestamp", "%{date} %{time}"]
            }

            date {
                    match => ["listener_timestamp", "yy-MM-dd HH:mm:ss"]
            }
            if [cs_uri_stem] {
                    mutate {
                            rename => { "cs_uri_stem" => "cs_uri_stem[request]" }
                    }
                    urldecode {
                            field => "cs_uri_stem[request]"
                    }
                    kv {
                            source => "cs_uri_stem[request]"
                            field_split => "?&"
                            target => "cs_uri_stem"
                    }
            ruby {
                    code => "
                    arguments = Array.new
                    event['cs_uri_stem'].to_hash.each {|k,v|
                    if k == 'request' then
                            next
                    end
                    arguments << { 'key' => k, 'value' => v }
                    }
                    unless arguments.empty?
                    event['[arguments]'] = arguments
            end
            "
            remove_field => ["cs_uri_stem"]
            }
    }

```

}

output {  
stdout { codec =\> rubydebug }  
}

Do you think csv would be better than grok in this use case?

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [March 29, 2016, 11:00am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/4 "2016-03-29T11:00:46Z")

</div>

> but after the logs got loaded to elasticsearch I am not able search the parameters for example aid="xxxxxxxx" AND bid="yyyyyyyyyyy"

What do the resulting events look like? Please show the output of the `stdout { codec => rubydebug }` output.

---

<div class="post-metadata">

### Author: ![manishr](https://avatars.discourse-cdn.com/v4/letter/m/c57346/32.png) [@manishr](https://discuss.elastic.co/u/manishr)
#### Post date: [March 29, 2016, 11:14am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/5 "2016-03-29T11:14:26Z")

</div>

{  
"message" =\> "2016-03-29\t04:02:08\tABC1\t461\t22.20.17.8\tGET\[tafsaGdhfxghxgh.cloudfront.net](http://tafsaGdhfxghxgh.cloudfront.net)\t/1.gif\t-\tMozilla/5.0%2520(Linux;%2520Android%25205.1.1;%2520SM-G920I%2520Build/LMY47X;%2520wv)%2520AppleWebKit/537.36%2520(KHTML,%2520like%2520Gecko)%2520Version/4.0%2520Chrome/48.0.2564.106%2520Mobile%2520Safari/537.36\taid=fsdggg25346&bid=fsdgagsexfdhg&cid=1423690744601076&cb=fdfsdggg&did=fsagdsgg&eid=fDSGzsgdfhdsh\t\t Miss\tjAk9duSOoOPVfssDGZdfhgxxxghfghpzK35tRuujwuQ==\[tafsaGdhfxghxgh.cloudfront.net](http://tafsaGdhfxghxgh.cloudfront.net)\thttps\t558\t0.715\t\t TLSv1.2\tECDHE-RSA-AES128-GCM-SHA256\tMiss",  
"@version" =\> "1",  
"@timestamp" =\> "2016-03-29T01:04:12.000Z",  
"path" =\> "/opt/cloudfront/E2I53NO2J8KEJZ",  
"host" =\> "localhost",  
"type" =\> "cloudfront",  
"date" =\> "16-03-27",  
"time" =\> "01:04:12",  
"x\_edge\_location" =\> "ABC1",  
"sc\_bytes" =\> "461",  
"c\_ip" =\> "22.20.17.8",  
"cs\_method" =\> "GET",  
"cs\_host" =\> "[afsaGdhfxghxgh.cloudfront.net](http://afsaGdhfxghxgh.cloudfront.net)",  
"cs\_uri" =\> "/1.gif",  
"sc\_status" =\> "200",  
"referrer" =\> "-",  
"User\_Agent" =\> "Mozilla/5.0%2520(Linux;%2520U;%2520Android%25204.2.2;%2520en-gb;%2520SM-T110%2520Build/JDQ39)%2520AppleWebKit/534.30%2520(KHTML,%2520like%2520Gecko)%2520Version/4.0%2520Safari/534.30",  
"cookies" =\> "-",  
"x\_edge\_result\_type" =\> "Miss",  
"x\_edge\_request\_id" =\> "jAk9duSOoOPVfssDGZdfhgxxxghfghpzK35tRuujwuQ==",  
"x\_host\_header" =\> "[afsaGdhfxghxgh.cloudfront.net](http://afsaGdhfxghxgh.cloudfront.net)",  
"cs\_protocol" =\> "https",  
"cs\_bytes" =\> "581",  
"time\_taken" =\> "0.042",  
"x\_forwarded\_for" =\> "-",  
"ssl\_protocol" =\> "TLSv1",  
"ssl\_cipher" =\> "ECDHE-RSA-AES128-SHA",  
"x\_edge\_response\_result\_type" =\> "Miss",  
"received\_at" =\> "2016-03-29T10:16:46.815Z",  
"listener\_timestamp" =\> "16-03-27 01:04:12",  
"arguments" =\> [  
[0] {  
"key" =\> "aid",  
"value" =\> "432432546376879869"  
},  
[1] {  
"key" =\> "bid",  
"value" =\> "reawca54rsyxdfhgtf"  
},  
[2] {  
"key" =\> "cid",  
"value" =\> "gzsdfhbxdfhx35q"  
},  
[3] {  
"key" =\> "did",  
"value" =\> "35434w65474ew"  
},  
[4] {  
"key" =\> "eid",  
"value" =\> "43r536w456"  
}  
]  
}

After loading this log, I see "arguments" field as not indexed and hence not searchable.

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [March 29, 2016, 11:28am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/6 "2016-03-29T11:28:19Z")

</div>

You probably don't want `arguments` to be an array of objects. Searches are not going to work like you expect them to. Instead, I suggest you aim for

```auto
"arguments": {
  "aid": "432432546376879869",
  "bid": "reawca54rsyxdfhgtf",
  ...
}

```

which is what the kv filter should give you out of the box.

---

<div class="post-metadata">

### Author: ![manishr](https://avatars.discourse-cdn.com/v4/letter/m/c57346/32.png) [@manishr](https://discuss.elastic.co/u/manishr)
#### Post date: [April 4, 2016, 1:53pm UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/7 "2016-04-04T13:53:28Z")

</div>

I used just kv filter and it is working as expected but it looks like while searching in kibana the count of log and ES data is different. Is it because I am using kv filter? Do we any alternative of kv filter to achieve what you told in your last response.

---

<div class="post-metadata">

### Author: ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)
#### Post date: [April 4, 2016, 5:05pm UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/8 "2016-04-04T17:05:00Z")

</div>

You have to be more specific than "while searching in kibana the count of log and ES data is different".

I don't think the kv filter has anything to do with this.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 5:04am UTC](https://discuss.elastic.co/t/logstash-configuration-for-cloudfront-logs/45660/9 "2017-07-06T05:04:01Z")

</div>


