# How to get .gz files using Gzip\_lines codec plugin from a pipeline?

**URL:** <https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746>\
**Category:** Logstash\
**Created:** [June 14, 2019, 2:08am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746 "2019-06-14T02:08:36Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![TsuWeiQuan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsuweiquan/32/46252_2.png) [@TsuWeiQuan](https://discuss.elastic.co/u/TsuWeiQuan)\
**Post date:** [June 14, 2019, 2:08am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/1 "2019-06-14T02:08:37Z")

</div>

How to get .gz files using Gzip\_lines codec plugin from a pipeline?

```
input {
  pipeline {
    address => testgz #has been configured to send .gz files into this pipe
    codec => gzip_lines
  }
}
output {
  stdout { codec => rubydebug }
}

```

But the message output is still garbled...  
Thanks

---

<div class="post-metadata">

**Author:** ![cknoell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cknoell/32/47353_2.png) [@cknoell](https://discuss.elastic.co/u/cknoell)\
**Post date:** [June 14, 2019, 8:40am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/2 "2019-06-14T08:40:38Z")

</div>

Are you using the correct charset ? I think UTF-8 is default. Maybe you need to set the correct one, if the default doesn't fit.

---

<div class="post-metadata">

**Author:** ![TsuWeiQuan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsuweiquan/32/46252_2.png) [@TsuWeiQuan](https://discuss.elastic.co/u/TsuWeiQuan)\
**Post date:** [June 14, 2019, 8:41am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/3 "2019-06-14T08:41:49Z")

</div>

Oh! i just checked and my .gz file is using Binary charset. i will change n try now.

-Update.

> [root@linuxclient adsm]# file -i gacs\_event\_hdlr.log.2019-05-28-03.gz  
> gacs\_event\_hdlr.log.2019-05-28-03.gz: application/x-gzip; charset=binary

and setting the charset to "BINARY" does not work. Message output is still garble.

---

<div class="post-metadata">

**Author:** ![yaauie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yaauie/32/23363_2.png) [@yaauie](https://discuss.elastic.co/u/yaauie)\
**Post date:** [June 14, 2019, 11:21pm UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/4 "2019-06-14T23:21:22Z")

</div>

The pipeline input ignores codecs (the Event objects are passed in memory without intermediate serialisation); this is a known issue in previous releases of Logstash and will be resolved before pipeline-to-pipeline graduates from beta.

---

<div class="post-metadata">

**Author:** ![TsuWeiQuan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsuweiquan/32/46252_2.png) [@TsuWeiQuan](https://discuss.elastic.co/u/TsuWeiQuan)\
**Post date:** [June 15, 2019, 12:54am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/5 "2019-06-15T00:54:13Z")

</div>

Thank you for notifying me of this! I guess I have to do some file manipulation.  
Hmm isit actually possible to gunzip files at the filter portion before processing the data?

---

<div class="post-metadata">

**Author:** ![yaauie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yaauie/32/23363_2.png) [@yaauie](https://discuss.elastic.co/u/yaauie)\
**Post date:** [June 15, 2019, 1:16am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/6 "2019-06-15T01:16:55Z")

</div>

What is the original source? You may be able to use the gzip\_lines codec at that point.

---

<div class="post-metadata">

**Author:** ![TsuWeiQuan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsuweiquan/32/46252_2.png) [@TsuWeiQuan](https://discuss.elastic.co/u/TsuWeiQuan)\
**Post date:** [June 15, 2019, 1:43am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/7 "2019-06-15T01:43:20Z")

</div>

Actually the source is on the same machine itself at /var/log/adsm/\*  
Because I have tons of logs there with .log & .out & .log.gz & .out.gz extensions, I replicate the distributor pattern for pipelining the data. Where 1 main pipe will distribute logs with \*.log to 1 pipeline and \*.out to another pipeline. This is done based on [fields][type] tagging.

Hence this portion is the last pipeline to elasticsearch.

---

<div class="post-metadata">

**Author:** ![TsuWeiQuan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsuweiquan/32/46252_2.png) [@TsuWeiQuan](https://discuss.elastic.co/u/TsuWeiQuan)\
**Post date:** [June 17, 2019, 2:22am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/8 "2019-06-17T02:22:32Z")

</div>

What about the gzip plugin working with

```
input {
  beats {
    port => 5046
    codec => gzip_lines { charset => "BINARY" }
  }
}

```

Where i ingest .gz files using filebeat on the client side and it doesn't seem to work either.  
In any case, i have done it via the manual extraction. But it will be good to know if this works too 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 15, 2019, 2:22am UTC](https://discuss.elastic.co/t/how-to-get-gz-files-using-gzip-lines-codec-plugin-from-a-pipeline/185746/9 "2019-07-15T02:22:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
