# Kv split everything after ";"

**URL:** <https://discuss.elastic.co/t/kv-split-everything-after/180403>\
**Category:** Logstash\
**Created:** [May 9, 2019, 4:04pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403 "2019-05-09T16:04:28Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![ghost3h](https://avatars.discourse-cdn.com/v4/letter/g/bc8723/32.png) [@ghost3h](https://discuss.elastic.co/u/ghost3h)\
**Post date:** [May 9, 2019, 4:04pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/1 "2019-05-09T16:04:28Z")

</div>

So I am trying to split the following data format examples:

> 'process\_count'=259;300;400;  
> 'cpu'=3.8%;;;  
> 'available'=27.29GiB;50;57;

I essentially want my key pair to be name and first value, e.g "process\_count" =\> "259", "available" =\> "27.59"

My code is

```
	kv {    	  
	  source => "SERVICEPERFDATA"
	  trim_key => "'"
	  remove_char_value => "GiB,%,;"
	  trim_value => ";"
	}

```

the trim\_value doesnt seem to work, and doesnt remove any of the ";" but I also want to delete everything after the first ";" . If I add the ";" into the remove\_char\_value, it does remove it, but keeps the values after/in between it.

Can anyone suggest how I can achieve this?  
Thanks  
Kyle

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 9, 2019, 5:09pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/2 "2019-05-09T17:09:12Z")

</div>

> [@ghost3h](#):
>
> Can anyone suggest how I can achieve this?

How about

```
grok { match => { "message" => "^'(?<key>[^']+)'=(?<value>[0-9\.]+).*" } }

```

I often see folks trying to use grok in use-cases where I think another filter is a better fit. This is a case where I think grok is a great fit 😃

---

<div class="post-metadata">

**Author:** ![ghost3h](https://avatars.discourse-cdn.com/v4/letter/g/bc8723/32.png) [@ghost3h](https://discuss.elastic.co/u/ghost3h)\
**Post date:** [May 10, 2019, 7:58am UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/3 "2019-05-10T07:58:34Z")

</div>

I might be using this wrong, as I don't know grok at all. Should I be putting this into a loop? As it only retrieves the first value from the field I give it. It also outputs the data into keypairs named "key" and "value". Which I cant replce with predefined one, I need these to be derived from the data.

`grok { match => { "SERVICEPERFDATA" => "^'(?<key>[^']+)'=(?<value>[0-9\.]+).*" } }`

> "SERVICEPERFDATA" =\> "'available'=5.51GiB;6;6; 'total'=7.00GiB;6;6; 'free'=5.51GiB;6;6; 'used'=1.49GiB;6;6;"  
> SERVICEPERFDATA:'process\_count'=259;300;400; 'cpu'=3.8%;;; 'memory'=18.1%;;; 'memory\_vms'=36.44GB;;; 'memory\_rss'=1.21GB;;;

Some examples of the data. I just want the value between the ' ' as my keyname, and the first value after.

I've been trying to do regex for ";\*" in the KV, but it doesnt seem to work?

```
	kv {
	  source => "SERVICEPERFDATA"
	  trim_key => "'"
	  field_split => " "
	  remove_char_value => "GiB,%,;*"
	  #trim_value => "\;*"
	}

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 10, 2019, 3:37pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/4 "2019-05-10T15:37:37Z")

</div>

> [@ghost3h](#):
>
> As it only retrieves the first value from the field I give it.

Well the sample data you included in your question only have a single key and value on each line. For that example of SERVICEPERFDATA the following would work

```
kv { source => "SERVICEPERFDATA" trim_key => "'" target => "[@metadata][spd]" }
ruby {
    code => '
        event.get("[@metadata][spd]").each { |k, v|
            m = /^[0-9\.]+/.match(v)
            event.set(k, m[0])
        }
    '
}

```

Error handling is left as an exercise for the reader.

---

<div class="post-metadata">

**Author:** ![ghost3h](https://avatars.discourse-cdn.com/v4/letter/g/bc8723/32.png) [@ghost3h](https://discuss.elastic.co/u/ghost3h)\
**Post date:** [May 13, 2019, 1:14pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/5 "2019-05-13T13:14:05Z")

</div>

This worked perfect, thanks Badger!

---

<div class="post-metadata">

**Author:** ![ghost3h](https://avatars.discourse-cdn.com/v4/letter/g/bc8723/32.png) [@ghost3h](https://discuss.elastic.co/u/ghost3h)\
**Post date:** [May 17, 2019, 2:55pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/6 "2019-05-17T14:55:53Z")

</div>

I'm trying to enhance this to put this into an array, with the name from another field. It works for the first element, but doesnt do the rest. Am I doing something stupid on a friday afternoon?

```
	ruby {
		code => '
			event.get("[@metadata][spd]").each { |k, v|
			m = /^[0-9\.]+/.match(v)
			event.set((event.get("SERVICEDESC")),[Hash[k, m[0]]])				
			}
		'
	}
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 17, 2019, 5:06pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/7 "2019-05-17T17:06:07Z")

</div>

That would overwrite that field once for each item in spd. This

```
input { generator { count => 1 lines => [''] } }
filter {
    mutate {
        add_field => {
            "SERVICEDESC" => "joe"
            "SERVICEPERFDATA" => "'available'=5.51GiB;6;6; 'total'=7.00GiB;6;6; 'free'=5.51GiB;6;6; 'used'=1.49GiB;6;6;"
        }
    }
    kv { source => "SERVICEPERFDATA" trim_key => "'" target => "[@metadata][spd]" }
    ruby {
        code => '
            a = []
            event.get("[@metadata][spd]").each { |k, v|
                m = /^[0-9\.]+/.match(v)
                a << Hash[k, m[0]]
            }
            event.set(event.get("SERVICEDESC"), a)
        '
    }
}

```

will get you

```
            "joe" => [
    [0] {
        "available" => "5.51"
    },
    [1] {
        "free" => "5.51"
    },
    [2] {
        "total" => "7.00"
    },
    [3] {
        "used" => "1.49"
    }
],

```

If that's not quite not what you want (and an array of hashes does seem an unlikely requirement) perhaps it will help you get there.

---

<div class="post-metadata">

**Author:** ![ghost3h](https://avatars.discourse-cdn.com/v4/letter/g/bc8723/32.png) [@ghost3h](https://discuss.elastic.co/u/ghost3h)\
**Post date:** [May 20, 2019, 1:22pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/8 "2019-05-20T13:22:41Z")

</div>

This does work for me, but I was doing a conversion after to chnge the values to floats. When I did this with the hash, and subsequent googling, told me this isnt possible. Myabe if I explain the sitatuons that might help.

I'm parsing nagios flat file, tab delimited log files for performance metrics. The problem being that some of the metric names (such as 'available', 'free') are used across several different metric types (SERVICEDESC). So in kibana, I was having issues using these for visulizations etc.

My intended solution was to nest these in an array of the metric type (SERVICEDESC). So they could then be referenced by metric\_type.metric\_name.

below is examples of the data

> SERVICEDESC:Memory Usage  
> SERVICEPERFDATA:'available'=27.29GiB;50;57; 'total'=62.89GiB;50;57; 'free'=1.09GiB;50;57; 'used'=34.46GiB;50;57;

> SERVICEDESC:Swap Usage  
> SERVICEPERFDATA:'total'=16.62GiB;13;15; 'used'=3.07GiB;13;15; 'free'=13.55GiB;13;15;

> SERVICEDESC:Disk Usage  
> SERVICEPERFDATA:'used'=3.58GiB;24;27; 'free'=25.91GiB;24;27; 'total'=29.49GiB;24;27;

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 20, 2019, 1:33pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/9 "2019-05-20T13:33:03Z")

</div>

I think it is more likely you will want to use

```
filter {
mutate {
    add_field => {
        "SERVICEDESC" => "joe"
        "SERVICEPERFDATA" => "'available'=5.51GiB;6;6; 'total'=7.00GiB;6;6; 'free'=5.51GiB;6;6; 'used'=1.49GiB;6;6;"
    }
}
kv { source => "SERVICEPERFDATA" trim_key => "'" target => "[@metadata][spd]" }
ruby {
    code => '
        h = {}
        event.get("[@metadata][spd]").each { |k, v|
            m = /^[0-9\.]+/.match(v)
            h[k] = m[0].to_f
        }
        event.set(event.get("SERVICEDESC"), h)
    '
}
}

```

Then you would be able to refer to [Memory Usage][available] or [Memory Usage][total] rather than [Memory Usage][0][available] and [Memory Usage][1][total].

---

<div class="post-metadata">

**Author:** ![ghost3h](https://avatars.discourse-cdn.com/v4/letter/g/bc8723/32.png) [@ghost3h](https://discuss.elastic.co/u/ghost3h)\
**Post date:** [May 20, 2019, 1:53pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/10 "2019-05-20T13:53:57Z")

</div>

That seems to work perfectly! Thanks badger

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 17, 2019, 2:07pm UTC](https://discuss.elastic.co/t/kv-split-everything-after/180403/11 "2019-06-17T14:07:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
