# Charset not considered on HTTP input plugin

**URL:** <https://discuss.elastic.co/t/charset-not-considered-on-http-input-plugin/62765>\
**Category:** Logstash\
**Created:** [October 11, 2016, 11:51pm UTC](https://discuss.elastic.co/t/charset-not-considered-on-http-input-plugin/62765 "2016-10-11T23:51:17Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Javier\_Bravo\_Conde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javier_bravo_conde/32/12419_2.png) [@Javier\_Bravo\_Conde](https://discuss.elastic.co/u/Javier_Bravo_Conde)\
**Post date:** [October 11, 2016, 11:51pm UTC](https://discuss.elastic.co/t/charset-not-considered-on-http-input-plugin/62765/1 "2016-10-11T23:51:17Z")

</div>

Hi!

Before opening a ticket I would like to discuss my issue here.

I receive a event.message string encoded in UTF-8 even though I specify codec =\> plain { charset =\> "ASCII-8BIT"} in my config.

Here my config:

> input {  
> http {  
> host =\> "0.0.0.0"  
> port =\> 8080  
> codec =\> plain { charset =\> "ASCII-8BIT"}  
> }  
> }  
> filter {  
> example {  
> }  
> }  
> output {  
> stdout { codec =\> rubydebug }  
> }

Here an extract my custom filter:

> public  
> def filter(event)

> ```
> @logger.debug? && @logger.debug("The event.message size is: #{event.get("message").size()}")
> @logger.debug? && @logger.debug("The event.message encoding is: #{event.get("message").encoding}")
> 
> ```

> ```
> counter = 0;
> event.get("message").each_byte { |c| 
> 
> ```
> 
> #Increments the counter for each byte within the string  
> counter +=1   
> }  
> @logger.debug? && @logger.debug("There are #{counter} bytes in the string")
> 
> ```
> # filter_matched should go in the last line of our successful code
> filter_matched(event)
> 
> ```
> 
> end # def filter

And here the output (I expect The event.message encoding is: ASCII-8BIT)

> filter received {:event=\>{"message"=\>"H\u0000\u0002\u0001\a\u0000\u0000\u0000\u0001\u0002\u0003\u0004\u0005C&\v\u0000\u0000\u0000\u0000\u0000\u0000�F\u0000\u0000"\v\u0000", "@version"=\>"1", "@timestamp"=\>"2016-10-11T23:32:52.277Z", "host"=\>"192.168.1.130", "headers"=\>{"request\_method"=\>"POST", "request\_path"=\>"/", "request\_uri"=\>"/", "http\_version"=\>"HTTP/1.1", "http\_user\_agent"=\>"Mozilla/4.0 (compatible; AP:FiOS-Mercury/3.2.2.3.3.3.6; PL:Motorola-DCT/KA15.76.12.19AlderF.560; BX:VMS1100; UA:0000108336906021; U; en-US)", "http\_host"=\>"192.168.1.239:8080", "http\_accept"=\>"_/_", "content\_type"=\>"application/x-www-form-urlencoded", "content\_length"=\>"2880"}}, :level=\>:debug, :file=\>"(eval)", :line=\>"41", :method=\>"filter\_func"}  
> The event.message size is: 2880 {:level=\>:debug, :file=\>"logstash/filters/example.rb", :line=\>"18", :method=\>"filter"}  
> The event.message encoding is: UTF-8 {:level=\>:debug, :file=\>"logstash/filters/example.rb", :line=\>"19", :method=\>"filter"}  
> There are 2898 bytes in the string {:level=\>:debug, :file=\>"logstash/filters/example.rb", :line=\>"27", :method=\>"filter"}

---

<div class="post-metadata">

**Author:** ![Javier\_Bravo\_Conde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javier_bravo_conde/32/12419_2.png) [@Javier\_Bravo\_Conde](https://discuss.elastic.co/u/Javier_Bravo_Conde)\
**Post date:** [October 12, 2016, 2:06am UTC](https://discuss.elastic.co/t/charset-not-considered-on-http-input-plugin/62765/2 "2016-10-12T02:06:13Z")

</div>

I just figured out this is actually a feature implemented on LogStash::Util::Charset

> def convert(data)  
> data.force\_encoding(@charset\_encoding)

> # NON UTF-8 charset declared.
> 
> # Let's convert it (as cleanly as possible) into UTF-8 so we can use it with JSON, etc.
> 
> return data.encode(Encoding::UTF\_8, :invalid =\> :replace, :undef =\> :replace) unless @charset\_encoding == Encoding::UTF\_8  
> ...

I might be missing something, but it would be great if we could specify some kind of 'keep\_original\_charset', this would allow handling arbitrary binary protocols at filter level.

---

<div class="post-metadata">

**Author:** ![Javier\_Bravo\_Conde](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javier_bravo_conde/32/12419_2.png) [@Javier\_Bravo\_Conde](https://discuss.elastic.co/u/Javier_Bravo_Conde)\
**Post date:** [October 13, 2016, 4:39pm UTC](https://discuss.elastic.co/t/charset-not-considered-on-http-input-plugin/62765/3 "2016-10-13T16:39:45Z")

</div>

Just in case someone hit a similar problem, I solved the issue by coding a custom codec and

Inside there you have the original encoding and you can do pretty much whatever you want (parse it, create an string-array of hex's...)

Here an extract:

public  
def decode(data)

```
array_data = data.unpack('C*')

header_char = array_data.shift(1).pack('C*') #1 My first byte
header_version_number = getnumber_frombytes(array_data.shift(1)) #2 My second byte
header_platform_id_number = getnumber_frombytes(array_data.shift(1)) #3 My third byte
header_isextended_number = getnumber_frombytes(array_data.shift(1)) #4 My fourth byte
....
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:34am UTC](https://discuss.elastic.co/t/charset-not-considered-on-http-input-plugin/62765/4 "2017-07-06T04:34:21Z")

</div>


