# Grok S3 object "path"

**URL:** <https://discuss.elastic.co/t/grok-s3-object-path/190390>\
**Category:** Logstash\
**Created:** [July 14, 2019, 5:31pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390 "2019-07-14T17:31:54Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![eclipsed450](https://avatars.discourse-cdn.com/v4/letter/e/b5e925/32.png) [@eclipsed450](https://discuss.elastic.co/u/eclipsed450)\
**Post date:** [July 14, 2019, 5:31pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/1 "2019-07-14T17:31:54Z")

</div>

Hi all,

I'm trying to figure out a way to grok out the key of an S3 object. Here is an example key:

`elasticmapreduce/j-fjtnfnfk56/containers/application_1111111111111_0169/container_1111111111111_0169_01_000029/stderr.gz`

I used [https://grokdebug.herokuapp.com/](https://grokdebug.herokuapp.com/) and came up with the following:

`%{WORD:folder}/(?<cluster>[^/]*)/(?<subfolder_name>[^/]*)/(?<application_id>[^/]*)/(?<container_id>[^/]*)/(?<file_name>[^$]*)`

But when I put it in my .conf file, as follows, it does't parse out the key into the fields:

`filter { grok { add_field => ["file", "%{[@metadata][s3][key]}" ] match => { "file" => "%{WORD:folder}\/(?<cluster_id>[^/]*)\/(?<subfolder_name>[^/]*)\/(?<application_id>[^/]*)\/(?<container_id>[^/]*)\/(?<file_name>[^$]*)" } } }`

Even though the grok appears to work via [https://grokdebug.herokuapp.com/:](https://grokdebug.herokuapp.com/:)

`{ "folder": [[ "elasticmapreduce"] ], "cluster": [[ "j-fjtnfnfk56"] ], "subfolder_name": [[ "containers"] ], "application_id": [[ "application_1111111111111_0169"] ], "container_id": [[ "container_1111111111111_0169_01_000029"] ], "file_name": [[ "stderr.gz\n"] ] }`

Also, how to get rid of the `\n` at the end of the file name?

Any help would be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 14, 2019, 8:24pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/2 "2019-07-14T20:24:49Z")

</div>

> [@eclipsed450](#):
>
> filter { grok { add\_field =\> ["file", "%{[@metadata][s3][key]}" ] match

add\_field will only get executed when the grok succesfully completes, so the file field will not exist when it tries to match it. Matching a non-existent field is a no-op but counts as a successful completion.

Use a literal newline in mutate to remove a newline.

```
mutate { gsub => [ "filename", "
", "" ] }

```

---

<div class="post-metadata">

**Author:** ![eclipsed450](https://avatars.discourse-cdn.com/v4/letter/e/b5e925/32.png) [@eclipsed450](https://discuss.elastic.co/u/eclipsed450)\
**Post date:** [July 14, 2019, 9:01pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/3 "2019-07-14T21:01:36Z")

</div>

> [@eclipsed450](#):
>
> { "folder": [["elasticmapreduce"] ], "cluster": [["j-fjtnfnfk56"] ], "subfolder\_name": [["containers"] ], "application\_id": [["application\_1111111111111\_0169"] ], "container\_id": [["container\_1111111111111\_0169\_01\_000029"] ], "file\_name": [["stderr.gz\n"] ] }

Thank you for that suggestion, however it only seems to have split the file field into a comma-separated line now. I tried adding the `add_field` option, but that didn't seem to do anything ☹

`filter { grok { match => { "message" => "%{DATESTAMP:message_timestamp} %{LOGLEVEL:severity} %{GREEDYDATA:msg}" } add_field => ["file", "%{[@metadata][s3][key]}" ] } mutate { copy => { "file" => "file_tmp" } split => ["file_tmp" , "/"] add_field => { "folder" => "%{file_tmp[0]}" "cluster" => "%{file_tmp[1]}" "subfolder_name1" => "%{file_tmp[2]}" "containers" => "%{file_tmp[3]}" "application_id" => "%{file_tmp[4]}" "container_id" => "%{file_tmp[5]}" "filename" => "%{file_tmp[6]}" } } }`

Also, using the code above, I got A BUNCH of these:  
` [2019-07-14T20:54:07,936][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference:`file[0]`"} [2019-07-14T20:54:07,936][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,936][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,936][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,936][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,936][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,936][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,937][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,937][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"} [2019-07-14T20:54:07,937][WARN][logstash.filters.mutate] Exception caught while applying mutate filter {:exception=>"Invalid FieldReference: `file[0]`"}`

Also, how do you do a multi-line code block?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 14, 2019, 10:03pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/4 "2019-07-14T22:03:41Z")

</div>

> [@eclipsed450](#):
>
> `Invalid FieldReference:` file[0]

That should be [file][0] etc.

---

<div class="post-metadata">

**Author:** ![eclipsed450](https://avatars.discourse-cdn.com/v4/letter/e/b5e925/32.png) [@eclipsed450](https://discuss.elastic.co/u/eclipsed450)\
**Post date:** [July 14, 2019, 10:23pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/5 "2019-07-14T22:23:10Z")

</div>

Thank you, the fields show up now, but the values of the fields are that value now :-/ (getting closer)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/0/3039801e466ff743dc095b1854562b4581408330.png)

It's worth asking, I saw the `dissect` filter. Would that be a better solution in this case instead?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 14, 2019, 11:22pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/6 "2019-07-14T23:22:46Z")

</div>

The dissect filter is often a better fit than grok for predictably delimited events.

```
input { generator { count => 1 lines => [''] } }
filter {
    mutate { add_field => { "path" => "elasticmapreduce/j-fjtnfnfk56/containers/application_1111111111111_0169/container_1111111111111_0169_01_000029/stderr.gz" } }
    dissect { mapping => { "path" => "%{folder}/%{cluster}/%{subfolder_name}/%{application_id}/%{container_id}/%{file_name}" } }
}
output { stdout { codec => rubydebug { metadata => true } } }

```

will generate an event with these fields

```
"subfolder_name" => "containers",
       "cluster" => "j-fjtnfnfk56",
          "path" => "elasticmapreduce/j-fjtnfnfk56/containers/application_1111111111111_0169/container_1111111111111_0169_01_000029/stderr.gz",
        "folder" => "elasticmapreduce",
     "file_name" => "stderr.gz",
"application_id" => "application_1111111111111_0169",
  "container_id" => "container_1111111111111_0169_01_000029"

```

However, it does not handle optional fields, or any unpredictability except padding on separators. That's why it is so fast.

---

<div class="post-metadata">

**Author:** ![eclipsed450](https://avatars.discourse-cdn.com/v4/letter/e/b5e925/32.png) [@eclipsed450](https://discuss.elastic.co/u/eclipsed450)\
**Post date:** [July 17, 2019, 3:08am UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/7 "2019-07-17T03:08:41Z")

</div>

Thank you for that breakdown. For now, I'm focusing on this absolute parent path, but I would like to modify it to read in all of the other parent directories, and unknown number of sub-directories. Is this possible with either dissect, or split?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 17, 2019, 1:53pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/8 "2019-07-17T13:53:02Z")

</div>

> [@eclipsed450](#):
>
> unknown number of sub-directories. Is this possible with either dissect, or split?

If you have a variable number of / in the path then if there is a constant number of / that constitute a prefix you could dissect that, then use split on whatever is left. For example,

```
    dissect { mapping => { "message" => "/%{field1}/%{field2}/%{field3}/%{restOfLine}" } }
    mutate { split => { "restOfLine" => "/" } }

```

would handle both "/a/b/c/1/2" and "/d/e/f/7/8/9".

---

<div class="post-metadata">

**Author:** ![eclipsed450](https://avatars.discourse-cdn.com/v4/letter/e/b5e925/32.png) [@eclipsed450](https://discuss.elastic.co/u/eclipsed450)\
**Post date:** [July 18, 2019, 2:29pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/9 "2019-07-18T14:29:17Z")

</div>

I'll give that a shot, thanks!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 15, 2019, 2:29pm UTC](https://discuss.elastic.co/t/grok-s3-object-path/190390/10 "2019-08-15T14:29:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
