# Logstash gsub remove string/characters

**URL:** <https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852>\
**Category:** Logstash\
**Created:** [September 18, 2017, 10:10am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852 "2017-09-18T10:10:29Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 18, 2017, 10:10am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/1 "2017-09-18T10:10:29Z")

</div>

In my log file, I need to remove a defined characters from the RAW log.

my RAW log sample as below.

```
  Sep 18 150942 db[27414]: category=\"Event\" subcategory=\"System\"....
  Sep 18 151043 db[27464]: category=\"Event\" subcategory=\"System\" ...

```

and I want to remove "**db[xxxxx]:**" from the RAW log before it extract with KV

I tried several options, but could not get expected result.

Here is one sample regex which I have tried under gsub

```
      gsub => [ 
		"message", "db[.*]", ""
		]
	}

```

But above doesn't removed the defined string.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 18, 2017, 11:32am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/2 "2017-09-18T11:32:58Z")

</div>

You need to escape the square brackets for that to work.

But I suggest that you use a grok filter to extract the different pieces of the log messages into discrete fields (timestamp in one field, program name in one field, pid in one field, and key/value pairs in one field).

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 18, 2017, 11:59am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/3 "2017-09-18T11:59:33Z")

</div>

resolved with removing the square bracket, now I have another issue when I filter RAW log with KV, and message error string doesn't filtered.

```
  Sep 18 18:15:43 category=\"Event\" subcategory=\"System\" typeid=30909 level=\"error\" user=\"admin\" nas=\"\" action=\"\" status=\"\" FTM license activation error: unable to resolve server domain name: directregistration.abc.com:443\u0000

```

message **FTM license activation error: unable to resolve server domain name: [directregistration.abc.com:443](http://directregistration.abc.com:443)\u0000**

doesn't filter with KV, how do I filter it with KV by adding any option ?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 18, 2017, 1:25pm UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/4 "2017-09-18T13:25:44Z")

</div>

That string comes after the key/value pairs so the kv filter won't help you. You should be able to extract the string with a grok filter if you can describe when the key/value pairs end. For example, is it safe to assume that the message at the end never contains a double quote? Put differently, how would grok know that "FTM license" is where the message begins?

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 18, 2017, 1:42pm UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/5 "2017-09-18T13:42:11Z")

</div>

last kv pair is " **status=**" value of status= may anything and after that error messages comes. But we can't say that error message begins with a exact word, it may change. The only thing we have to consider is the last key/value pair.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 18, 2017, 1:51pm UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/6 "2017-09-18T13:51:44Z")

</div>

Well, it's good enough if we know that the last key is status. Then a grok expression similar to

```
(match date and program/pid here) (?<kvpairs>.* status="[^"]+" )%{GREEDYDATA:extra_message}

```

should work.

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 19, 2017, 5:13am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/7 "2017-09-19T05:13:25Z")

</div>

I tried below code

```
    grok {
		match => ["message", "%{SYSLOGTIMESTAMP:timestamp} %{GREEDYDATA:kvpairs}"]
	}

	kv { 
		source => "kvpairs"
	}

```

Above code correctly filter date and KV pairs.

sample output for above filter is listed below and full raw message has not dropped in the filtered result.

```
 {
   "nas_name" => "\"\"",
  "log_level" => "error",
     "source" => "192.168.1.1",
    "message" => "Sep 19 09:31:42 category=\"Event\" subcategory=\"System\" typeid=30909 level=\"error\" user=\"admin\" nas=\"\" action=\"\" status=\"\" FTM license activation error: unable to resolve server domain name: directregistration.abc.com:443\u0000",
       "type" => "forti-authnticator",
   "@version" => "1",
     "action" => "\"\"",
     "typeid" => "30909",
   "category" => "Event",
"subcategory" => "System",
       "user" => "admin",
  "timestamp" => "Sep 19 09:31:42",
     "status" => "\"\""
}

```

I have add additional filter grok to separate message on end of the log. But conf file gives errors and failed to execute.

As you said, I have tried to separate the extra message at the end of the RAW log. But tried scenarios gives lots of errors. So what is the best way to match, is it okey again apply grok pattern to separate ?.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 19, 2017, 5:16am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/8 "2017-09-19T05:16:34Z")

</div>

> I have add additional filter grok to separate message on end of the log. But conf file gives errors and failed to execute.

I can't help without knowing what your configuration looked like.

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 19, 2017, 5:24am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/9 "2017-09-19T05:24:40Z")

</div>

Here is the sample log

```
<11>Sep 18 12:16:42 db[19023]: category=\"Event\" subcategory=\"System\" typeid=30909 level=\"error\" user=\"admin\" nas=\"\" action=\"\" status=\"\" FTM license activation error: unable to resolve server domain name: directregistration.abc.com:443\u0000

```

here is the configuration file

```
input {
    udp {
	  port => 30001 
      type => "raw_log"
    }
}

filter {
   if [type] == "raw_log" 
   { 
    mutate {
	     gsub => [ 
		         "message", "db\[.*\]: ", "", // => remove db[xxxxx]:
		         "message", "<.*>", "" // => remove <xx>
		 ]
	 }

	grok {
		match => ["message", "%{SYSLOGTIMESTAMP:timestamp} %{GREEDYDATA:kvpairs}"]
	}

	kv { 
		source => "kvpairs"
	}

	mutate {
		rename => ["oid", "type_id"]z
		rename => ["logid", "log_id"]
	    rename => ["level", "log_level"]
		rename => ["cat", "category"]
		rename => ["subcategorty", "sub_categorty"]
		rename => ["nas", "nas_name"]
		rename => ["host", "source"]
		remove_field => ["@timestamp", kvpairs]
         }
   }
}

output {
     stdout { codec => rubydebug }
}

```

Could you help me to sort out this !

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 19, 2017, 5:46am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/10 "2017-09-19T05:46:47Z")

</div>

> ```
> rename => ["oid", "type_id"]z
> 
> ```

Do you really have a "z" there?

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 19, 2017, 5:53am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/11 "2017-09-19T05:53:09Z")

</div>

No, it is mistake while I put it here. there is no z

```
 rename => ["oid", "type_id"]
```

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 19, 2017, 6:52am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/12 "2017-09-19T06:52:51Z")

</div>

I have resolved it using grok with help of your hint

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 19, 2017, 11:17am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/13 "2017-09-19T11:17:41Z")

</div>

I need to resolve one thing again

I need to convert

SYSLOGTIMESTAMP into TIMESTAMP\_ISO8601

For an example

**Sep 19 16:41:43** in to **2017-09-19 16:41:43**

I tried below coding

```
   date {
		match => ["timestamp", " yyyy-MM-dd HH:mm:ss.SSS"]
	}

```

it doesn't convert ti required format

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 19, 2017, 11:48am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/14 "2017-09-19T11:48:19Z")

</div>

The pattern should be the format you have, not the format you want. See [https://www.elastic.co/guide/en/logstash/current/config-examples.html#\_processing\_syslog\_messages](https://www.elastic.co/guide/en/logstash/current/config-examples.html#_processing_syslog_messages) for an example.

---

<div class="post-metadata">

**Author:** ![tharu85](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@tharu85](https://discuss.elastic.co/u/tharu85)\
**Post date:** [September 20, 2017, 3:26am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/15 "2017-09-20T03:26:21Z")

</div>

Yes, that's why I need a method to convert into the format which I need

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 20, 2017, 3:49am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/16 "2017-09-20T03:49:57Z")

</div>

The date filter doesn't let you configure the output format. You'll have to write some Ruby code in a ruby filter.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 18, 2017, 3:50am UTC](https://discuss.elastic.co/t/logstash-gsub-remove-string-characters/100852/17 "2017-10-18T03:50:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
