# S3 logstash conf grok filter

**URL:** <https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357>\
**Category:** Logstash\
**Created:** [October 26, 2020, 5:53pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357 "2020-10-26T17:53:29Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![deep9300](https://avatars.discourse-cdn.com/v4/letter/d/c89c15/32.png) [@deep9300](https://discuss.elastic.co/u/deep9300)\
**Post date:** [October 26, 2020, 5:53pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357/1 "2020-10-26T17:53:29Z")

</div>

Hello.  
I want to collect s3 access logs from an s3 bucket and process them to logstash and elasticsearch. I have it working properly but the filter in the logstash conf is not working properly. Currently its giving me too much information when I only specific parts in the message.

So on kibana I'm getting logs like:  
message:

{"Records":[{"eventVersion":"1.05","userIdentity":{"type":"AssumedRole","principalId":"AROAJKLFDKRWTVGOAWDHWH:i-0276d215093829d49","arn":"arn:aws:sts::204324406053:assumed-role/EC2forSSM-Scaling/i-0276d215093829d49","accountId":"204324406053","accessKeyId":"ACCESSKEYIDFJDN3246","sessionContext":{"sessionIssuer":{"type":"Role","principalId":"AROAJKLFDKRWTVGOAWDHWH","arn":"arn:aws:iam::205915406053:role/EC2forSSM-Scaling","accountId":"204324406053","userName":"EC2forSSM-Scaling"},"webIdFederationData":{},"attributes":

message:

{"Records":[{"eventVersion":"1.05","userIdentity":{"type":"AWSService","invokedBy":"[autoscaling.amazonaws.com](http://autoscaling.amazonaws.com)"},"eventTime":"2020-10-26T16:28:51Z","eventSource":"[sts.amazonaws.com](http://sts.amazonaws.com)","eventName":"AssumeRole","awsRegion":"us-east-1","sourceIPAddress":"[autoscaling.amazonaws.com](http://autoscaling.amazonaws.com)","userAgent":"[autoscaling.amazonaws.com](http://autoscaling.amazonaws.com)","requestParameters":{"roleArn":"arn:aws:iam::204324406053:role/aws-service-

How do I createa grok filter to allow me to filter by BUCKET event NAME SOURCE IP, Username - but exclude other personal info like arn number, account id, etc. ?

Only want to make it useful for s3 access log activity. do not need too much additional information.

PS the values I have in here are not real - I changed them for the example

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [October 26, 2020, 7:55pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357/2 "2020-10-26T19:55:55Z")

</div>

Records is an array. Are you going to use a split filter to split that into multiple events?

If you are you may be able to use a prune filter with the whitelist\_names option to specify which fields to keep.

If you are not you would have to use ruby code to iterate over the array and (in effect) implement the prune yourself.

---

<div class="post-metadata">

**Author:** ![deep9300](https://avatars.discourse-cdn.com/v4/letter/d/c89c15/32.png) [@deep9300](https://discuss.elastic.co/u/deep9300)\
**Post date:** [October 26, 2020, 8:49pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357/3 "2020-10-26T20:49:21Z")

</div>

Can you give a basic structure on how to split the events and insert the prune filter in this filter?

```
filter {
   split {
     field => "Records"
        }
   prune {
        whitelist_names => [ "principalId", "arn", "accountId", "accessKeyId", etc
  }
   }
 }
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [October 26, 2020, 9:23pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357/4 "2020-10-26T21:23:22Z")

</div>

That split looks fine. However, prune is not going to work because it only operates on top level fields.

If you have a small number of fields you want to retain then you could do something like

```
    split { field => "Records" }
    mutate {
        add_field => {
            "[Record][userIdentity][sessionContext][sessionIssuer][arn]" => "%{[Records][userIdentity][sessionContext][sessionIssuer][arn]}"
            "[Record][userIdentity][sessionContext][sessionIssuer][userName]" => "%{[Records][userIdentity][sessionContext][sessionIssuer][userName]}"
            "[Record][userIdentity][accessKeyId]" => "%{[Records][userIdentity][accessKeyId]}"
        }
        remove_field => ["Records"]
    }

```

Obviously you do not have to use the same structure on the left that you have on the right. You could also do

```
    mutate {
        add_field => {
            "[arn]" => "%{[Records][userIdentity][sessionContext][sessionIssuer][arn]}"
            "[userName]" => "%{[Records][userIdentity][sessionContext][sessionIssuer][userName]}"
            "[accessKeyId]" => "%{[Records][userIdentity][accessKeyId]}"
        }
        remove_field => ["Records"]
    }
```

---

<div class="post-metadata">

**Author:** ![deep9300](https://avatars.discourse-cdn.com/v4/letter/d/c89c15/32.png) [@deep9300](https://discuss.elastic.co/u/deep9300)\
**Post date:** [October 28, 2020, 6:29pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357/5 "2020-10-28T18:29:38Z")

</div>

[2020-10-28T13:27:59,738][WARN][logstash.filters.split][main] Only String and Array types are splittable. field:Records is of type = NilClass  
[2020-10-28T13:27:59,741][WARN][logstash.filters.split][main] Only String and Array types are splittable. field:Records is of type = NilClass  
[2020-10-28T13:27:59,756][WARN][logstash.filters.split][main] Only String and Array types are splittable. field:Records is of type = NilClass  
[2020-10-28T13:27:59,775][WARN][logstash.filters.split][main] Only String and Array types are splittable. field:Records is of type = NilClass  
[2020-10-28T13:28:11,538][WARN][logstash.filters.split][main] Only String and Array types are splittable. field:Records is of type = NilClass  
[2020-10-28T13:28:11,760][WARN][logstash.filters.split][main] Only String and Array types are splittable. field:Records is of type = NilClass

I tried both filters above but get the following warning.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [October 28, 2020, 6:43pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357/6 "2020-10-28T18:43:03Z")

</div>

That is telling you that there are events that do not have a [Records] field. You could wrap the split and mutate in

```
if [Records] {
...
}
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 25, 2020, 6:43pm UTC](https://discuss.elastic.co/t/s3-logstash-conf-grok-filter/253357/7 "2020-11-25T18:43:12Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
