# Read \*.json.gz from AWS S3 bucket

**URL:** <https://discuss.elastic.co/t/read-json-gz-from-aws-s3-bucket/292822>\
**Category:** Logstash\
**Created:** [December 23, 2021, 3:13pm UTC](https://discuss.elastic.co/t/read-json-gz-from-aws-s3-bucket/292822 "2021-12-23T15:13:00Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vidya\_Sagar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vidya_sagar/32/99487_2.png) [@Vidya\_Sagar](https://discuss.elastic.co/u/Vidya_Sagar)\
**Post date:** [December 23, 2021, 3:13pm UTC](https://discuss.elastic.co/t/read-json-gz-from-aws-s3-bucket/292822/1 "2021-12-23T15:13:00Z")

</div>

Hi All,

I am new to ELK Stack and trying to read data from S3 buckets. The json data is in compressed format and the folder structure in the S3 bucket is like

YYYY-MM-DD/.json.gz

2021-12-01/A.json.gz  
2021-12-02/B.json.gz  
2021-12-03/C.json.gz

Folder can have multiple files..

Looking for some code snippet to read these files using Logstash.

---

<div class="post-metadata">

**Author:** ![AquaX](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aquax/32/92006_2.png) [@AquaX](https://discuss.elastic.co/u/AquaX)\
**Post date:** [December 23, 2021, 9:27pm UTC](https://discuss.elastic.co/t/read-json-gz-from-aws-s3-bucket/292822/2 "2021-12-23T21:27:17Z")

</div>

Read the file input plugin docs [File input plugin | Logstash Reference [7.16] | Elastic](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html)  
[read](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-mode) mode supports gzip file processing but I believe you have to define a [gzip codec](https://www.elastic.co/guide/en/logstash/7.16/plugins-codecs-gzip_lines.html) then in your input.  
However, try it without the codec and see if just the read works on it's own. I haven't tried that before.

```auto
input {
   file {
       path => ["/var/log/202*/*.json.gz"]
       codec => "gzip_lines"
       mode => "read"
   }
}

```

---

<div class="post-metadata">

**Author:** ![yaauie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yaauie/32/23363_2.png) [@yaauie](https://discuss.elastic.co/u/yaauie)\
**Post date:** [December 28, 2021, 7:24pm UTC](https://discuss.elastic.co/t/read-json-gz-from-aws-s3-bucket/292822/3 "2021-12-28T19:24:45Z")

</div>

From the [S3 input's docs](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-s3.html):

> Each line from each file generates an event. Files ending in `.gz` are handled as gzip’ed files.

Since the S3 input is line-oriented, if the contents of your GZIP files are _not_ line-oriented (such as each being a JSON blob representing a single JSON object), you may need to use the multiline codec to buffer all of the lines into a single event, and then a json Filter to parse the contents into a structured object:

```auto
input {
  s3 {
    bucket => ""
    access_key_id => "1234"
    secret_access_key => "secret"
    codec => multiline {
      pattern => "." # anything
      what => "previous" # accumulate until EOF
    }
  }
}
filter {
  json {
    source => "message"
  }
}

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 25, 2022, 7:24pm UTC](https://discuss.elastic.co/t/read-json-gz-from-aws-s3-bucket/292822/4 "2022-01-25T19:24:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
