# How to urldecode %uXXXX type of strings?

**URL:** <https://discuss.elastic.co/t/how-to-urldecode-uxxxx-type-of-strings/27718>\
**Category:** Logstash\
**Created:** [August 20, 2015, 1:35am UTC](https://discuss.elastic.co/t/how-to-urldecode-uxxxx-type-of-strings/27718 "2015-08-20T01:35:27Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![foresightyj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/foresightyj/32/3666_2.png) [@foresightyj](https://discuss.elastic.co/u/foresightyj)\
**Post date:** [August 20, 2015, 1:35am UTC](https://discuss.elastic.co/t/how-to-urldecode-uxxxx-type-of-strings/27718/1 "2015-08-20T01:35:27Z")

</div>

I needed a javascript equivalent of `unescape` in logstash. I know %uXXXX is not standard but this is still quite common. A lot of our users' browsers send urls in this format. I am reserving the use of `ruby` filter in logstash to decode this type of url as the last resort. Before that, I am looking for a less heavy solution. Thanks for any suggestions.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [August 20, 2015, 3:49am UTC](https://discuss.elastic.co/t/how-to-urldecode-uxxxx-type-of-strings/27718/2 "2015-08-20T03:49:52Z")

</div>

I suppose you've concluded that the [urldecode filter](https://www.elastic.co/guide/en/logstash/current/plugins-filters-urldecode.html) doesn't cut it?

---

<div class="post-metadata">

**Author:** ![foresightyj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/foresightyj/32/3666_2.png) [@foresightyj](https://discuss.elastic.co/u/foresightyj)\
**Post date:** [August 20, 2015, 5:26am UTC](https://discuss.elastic.co/t/how-to-urldecode-uxxxx-type-of-strings/27718/3 "2015-08-20T05:26:13Z")

</div>

![](https://us1.discourse-cdn.com/elastic/original/2X/4/4f83dc4a7c87cbb21ab7030087fc762fd8d91485.png) No. The `urldecode` filter cannot decode `%uXXXX` type of encoded urls. With a minimal config like this:

```
input {
    stdin {
    }
}

filter{
    urldecode {
        field => "message"
    }
}

output {
    stdout {
        codec => rubydebug
    }
}

```

Test it with two lines of url-encoded text, which, if correctly decoded, represent the same sequence of chinese characters:

![](https://us1.discourse-cdn.com/elastic/original/2X/d/d9f4671504acf0b98b46716c8b0dea07b0bccb2e.png)

I already figured out a ruby filter shown below to circumvent the limitation:

```
ruby {
    code => "
        # urldecode non-standard %uXXXX type of string
        ['cs_uri_query', 'cs_cookie', 'cs_referer'].each { |field|
            if event[field] and event[field].include? '%u'
                event[field] = event[field].gsub(/%u([0-9A-F]{4})/i){$1.hex.chr(Encoding::UTF_8)}.strip
            end
        }
    "
}

```

But I am still looking for easier ways.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:31am UTC](https://discuss.elastic.co/t/how-to-urldecode-uxxxx-type-of-strings/27718/4 "2017-07-06T05:31:29Z")

</div>


