# KV how to extract data between square brackets

**URL:** <https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678>\
**Category:** Logstash\
**Created:** [March 2, 2020, 12:48pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678 "2020-03-02T12:48:47Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![sahar](https://avatars.discourse-cdn.com/v4/letter/s/ea5d25/32.png) [@sahar](https://discuss.elastic.co/u/sahar)\
**Post date:** [March 2, 2020, 12:48pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/1 "2020-03-02T12:48:47Z")

</div>

Hello,  
Assume I have the following message:

"Some message [key1=val1] more text [key2=some [special] value], more text."

From the above message I would like to extract key1 and key2 as follows:  
key1 = val1  
key2 = some [special] value

How can I achieve that? I found only 1 topic similar to this:

> [@KV Filtering data between square brackets](https://discuss.elastic.co/t/kv-filtering-data-between-square-brackets/126124):
>
> Hi, How can i parse data which has two layers of square brackets. Inner square brackets Outer square brackets I only want to consider outer square brackets data as key value pairs. Example : [filename=0\_578cc[R2][]2 veckor f=f600re (2).doc] In the above example i want to take filename as key but it is also considering the inner square bracket which is breaking my logic. I am using grok and kv filter with the following config. match =\> [ "message", "%{SYSLOGTIMESTAMP:eventtime}\s(?[^\s…

Tried to play around with the solution given there but it never seems to be working for the case above.

A more realistic example:  
String:  
`Found cached [valuejson={"liveStreams":[]}] for [controller=Game] and [action=GetLiveStreams].`

I want to extract:

```
valuejson = {"liveStreams":[]}
controller = Game
action = GetLiveStreams

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 2, 2020, 2:23pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/2 "2020-03-02T14:23:13Z")

</div>

> [@sahar](#):
>
> From the above message I would like to extract key1 and key2 as follows:  
> key1 = val1  
> key2 = some [special] value

How do you know that it shouldn't be

```
key1 = val1] more text [key2=some [special] value

```

That's a serious question. What is a definition of the pattern that matches the RHS?

---

<div class="post-metadata">

**Author:** ![sahar](https://avatars.discourse-cdn.com/v4/letter/s/ea5d25/32.png) [@sahar](https://discuss.elastic.co/u/sahar)\
**Post date:** [March 2, 2020, 2:31pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/3 "2020-03-02T14:31:00Z")

</div>

Hi @Badger

If i would have to describe it in words i'd say there are 2 options to deal with what I want to achieve:

1. match every nearest pair of square brackets that contains = inside them (so if there's square brackets without = inside them, they will be used as a part of the value).
2. match fields only if there's an equal amount of square brackets: [https://stackoverflow.com/questions/546433/regular-expression-to-match-balanced-parentheses](https://stackoverflow.com/questions/546433/regular-expression-to-match-balanced-parentheses)

I tried to do as in that stackoverflow topic, unfortunately it seems like logstash fails to read the regexes that are suggested in the answer there, logstash doesn't even load.

Edit: here is an example from that topic that seems to work when I test it, but doesn't work in logstash itself:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/2/b2f148383168090c11e96791bc45dceff433897f.png)

---

<div class="post-metadata">

**Author:** ![sahar](https://avatars.discourse-cdn.com/v4/letter/s/ea5d25/32.png) [@sahar](https://discuss.elastic.co/u/sahar)\
**Post date:** [March 2, 2020, 4:56pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/4 "2020-03-02T16:56:12Z")

</div>

So after I understood that recursive regex isn't supported probably I tried a different regex with max 2 levels of nesting:  
`\[(?:[^\]\[]+|\[(?:[^\]\[]+|\[[^\]\[]*\])*\])*\]`

It seems to work when testing:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/a/7a684dc13610d989de4841ed4a1b5cf907fe74c0.png)

Logstash also loads successfully, but still splits the fields in the wrong way...

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/a/fa7614fe5bde90e99344d9dff9d7e72f62dbc7eb.png)

What am I missing? perhaps need a different filter for this one? or even ruby code?  
Thanks for the help.

---

<div class="post-metadata">

**Author:** ![sahar](https://avatars.discourse-cdn.com/v4/letter/s/ea5d25/32.png) [@sahar](https://discuss.elastic.co/u/sahar)\
**Post date:** [March 3, 2020, 5:27pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/5 "2020-03-03T17:27:57Z")

</div>

Anyone any ideas? ;o

@Badger can the pattern you suggested on the link I gave in the first post be modified somehow to match my needs?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [March 3, 2020, 6:27pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/6 "2020-03-03T18:27:02Z")

</div>

I do not know enough about regexps to write one that does what you want.

---

<div class="post-metadata">

**Author:** ![sahar](https://avatars.discourse-cdn.com/v4/letter/s/ea5d25/32.png) [@sahar](https://discuss.elastic.co/u/sahar)\
**Post date:** [March 3, 2020, 6:52pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/7 "2020-03-03T18:52:31Z")

</div>

@Badger can I perhaps use Ruby code to achieve this?  
For example, I already wrote a regex to capture the groups I need here: [KV how to extract data between square brackets](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/4)  
Can I write a ruby code to save them under a certain event, and then send this event to KV filter?

---

<div class="post-metadata">

**Author:** ![sahar](https://avatars.discourse-cdn.com/v4/letter/s/ea5d25/32.png) [@sahar](https://discuss.elastic.co/u/sahar)\
**Post date:** [March 3, 2020, 8:14pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/8 "2020-03-03T20:14:29Z")

</div>

For these interested, here's how I achieved it:

```
ruby {
	code => 'event.set("kv", event.get("tempMessage").scan(/\[(?:[^\]\[]+|\[(?:[^\]\[]+|\[[^\]\[]*\])*\])*\]/))'
}
kv {
	source => "kv"
	field_split_pattern => "(?:^\[|\]$)"
	trim_key => " "
	trim_value => " "
}

```

Apparently kv can also take parameter of array of strings, which makes life lot easier for my case since it treats each array value independently.

Thanks for the helpers.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 31, 2020, 8:14pm UTC](https://discuss.elastic.co/t/kv-how-to-extract-data-between-square-brackets/221678/9 "2020-03-31T20:14:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
