# Unable to parse log containing UNICODE characters and ANSI colour codes using grok

**URL:** <https://discuss.elastic.co/t/unable-to-parse-log-containing-unicode-characters-and-ansi-colour-codes-using-grok/33487>\
**Category:** Logstash\
**Created:** [November 2, 2015, 7:58am UTC](https://discuss.elastic.co/t/unable-to-parse-log-containing-unicode-characters-and-ansi-colour-codes-using-grok/33487 "2015-11-02T07:58:31Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![safiyat](https://avatars.discourse-cdn.com/v4/letter/s/f0a364/32.png) [@safiyat](https://discuss.elastic.co/u/safiyat)\
**Post date:** [November 2, 2015, 7:58am UTC](https://discuss.elastic.co/t/unable-to-parse-log-containing-unicode-characters-and-ansi-colour-codes-using-grok/33487/1 "2015-11-02T07:58:31Z")

</div>

I am trying to parse the following line:

`2015-09-17 17:44:49.663 ^[[00;32mDEBUG oslo_concurrency.lockutils [^[[00;36m-^[[00;32m] ^[[01;35m^[[00;32mAcquired semaphore "singleton_lock"^[[00m ^[[00;33mfrom (pid=30534) lock /usr/local/lib/python2.7/dist-packages/oslo_concurrency/lockutils.py:198^[[00m`

The character `^[` is actually the [ESC](https://en.wikipedia.org/wiki/Escape_character#ASCII_escape_character) key whose octal code is `\033`, hex code is `\x1B`.

The sub-string `^[[00;32m` and others like that are actually [ANSI](https://en.wikipedia.org/wiki/ANSI_escape_code#Colors) colour codes, which when printed in a terminal is printed like [this](http://i.stack.imgur.com/m8wei.png).

I need to parse this log line but have been unable to do it. I tried using the [color-stripper plugin](https://github.com/mattheworiordan/fluent-plugin-color-stripper) but it won't work for me.

I am able to parse the log line in plaintext using the pattern:

`%{TIMESTAMP_ISO8601:timestamp}%{SPACE}%{LOGLEVEL:loglevel}%{SPACE}{NOTSPACE:api}%{SPACE}\[(?:%{DATA:request})\]%{SPACE}%{GREEDYDATA:message}`

How do I parse the coloured log line?  
To parse it at the character level, we need to parse the unicode character `\u001B`. Any alternate way to do it by parsing the unicode character?

The related stackoverflow question can be found [here](http://stackoverflow.com/questions/33440366/grok-pattern-to-parse-the-esc-key).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:24am UTC](https://discuss.elastic.co/t/unable-to-parse-log-containing-unicode-characters-and-ansi-colour-codes-using-grok/33487/2 "2017-07-06T05:24:34Z")

</div>


