# Parsing URL with Logstash (using ECS fields) nested!

**URL:** <https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields-nested/209953>\
**Category:** Logstash\
**Tags:** ecs-elastic-common-schema\
**Created:** [November 29, 2019, 10:12am UTC](https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields-nested/209953 "2019-11-29T10:12:13Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vincent\_Maury](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vincent_maury/32/59973_2.png) [@Vincent\_Maury](https://discuss.elastic.co/u/Vincent_Maury)\
**Post date:** [November 29, 2019, 10:12am UTC](https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields-nested/209953/1 "2019-11-29T10:12:14Z")

</div>

Following [Parsing URL with Logstash (using ECS fields)](https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields/205814)  
Several improvements like nested fields & tld parsing

```
# Parser a url field named url (according to rsa meta field name)
filter {
	mutate {
		rename => { "message" => "[url][original]"}
	}
	grok {
		match => {
      "[url][original]" => [
  			# match https://user:pwd@stuff.domain.com:8080/some/path?p1=v1&p2=v2#anchor
  			"%{URIPROTO:[url][scheme]}://(?:%{USER:[url][username]}:(?<[url][password]>[^@]*)@)?(?:%{IPORHOST:[url][address]}(?::%{POSINT:[url][port]}))?(?:%{URIPATH:[url][path]}(?:%{URIPARAM:[url][query]}))?",
  			# match stuff.domain.com:8080/some/path?p1=v1&p2=v2#anchor
  			"%{IPORHOST:[url][address]}(?::%{POSINT:[url][port]})(?:%{URIPATH:[url][path]}(?:%{URIPARAM:[url][query]}))?",
  			# match /some/path?p1=v1&p2=v2#anchor
  			"%{URIPATH:[url][path]}(?:%{URIPARAM:[url][query]})"
      ]
		}
		add_tag => ["urlparsed"]
	}
	if "urlparsed" in [tags] {
		# parse the address to distinguish domain or ip
		grok {
			match => {
				"[url][address]" => "(%{IP:[url][ip]}|%{HOSTNAME:[url][domain]})"
			}
		}
		# Requires a custom plugin here, see https://www.elastic.co/guide/en/logstash/current/plugins-filters-tld.html
		tld {
			source => "[url][domain]"
			target => "[url][tld]"
		}
		mutate {
			rename => {
				"[url][tld][domain]" => "[url][registered_domain]"
				"[url][tld][tld]" => "[url][top_level_domain]"
				"[url][tld][sld]" => "[url][second_level_domain]"
				"[url][tld][trd]" => "[url][sub_domain]"
			}
			remove_field => ["[url][tld]" ]
		}
		# parse the query to extract fragment
		grok {
			match => {
				"[url][query]" => "^\?(?<[url][query]>[A-Za-z0-9$.+!*'|(){},~@%&/=:;_?\-\[\]<>]*)(?:#(?:%{WORD:[url][fragment]}))?"
			}
      overwrite => ["[url][query]" ]
		}
		kv {
  		source => "[url][query]"
  		field_split => "&"
			value_split => "="
  		target => "[url][queryparams]"
		}
	}
}
```

---

<div class="post-metadata">

**Author:** ![webmat](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/webmat/32/46191_2.png) [@webmat](https://discuss.elastic.co/u/webmat)\
**Post date:** [November 29, 2019, 1:38pm UTC](https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields-nested/209953/2 "2019-11-29T13:38:45Z")

</div>

Haven't tested it, but it looks good 👍

A few comments:

- ECS comment: make sure to parse out the extension, when there's one 🙂
- Non-ECS comment: I've taken down a cluster once, parsing out all query params of a busy web application 😂 Make sure your custom field `url.queryparams` uses a datatype that's not going to cause a mapping explosion, like [flattened](https://www.elastic.co/guide/en/elasticsearch/reference/current/flattened.html) 🙂
- ECS Nitpick: The best match for `IPORHOST` in ECS is `.address`, which is specified as the address "when you're not sure yet if it's an IP, a domain or a unix socket". So for URLs, you'd fill .address, then copy out to .ip if it's an IP, and otherwise copy to .domain.
  - This way you end up with `.address` which is filled reliably, 100% of the time.
  - If you need to do IP-specific analysis, you have `.ip` which is the `ip` datatype, and lets you do CIDR lookups, for example. Or just looking for `exists:url.ip` may surface interesting weird stuff.
  - If you need to do analysis on domain names, you have `.domain` and all of the domain breakdown fields which won't contain IP addresses.

- This actually brings me to the domain breakdown fields. I don't recall if we have a solid way to break them down by effective TLD in Logstash . But if the data source analyses many domains (as opposed to incoming web traffic on your webserver), it would be interesting to fill the domain breakdown fields as well. So "[www.example.co.uk](http://www.example.co.uk)" becomes:
  - `.top_level_domain:co.uk` =\> to analyze broad traffic destinations
  - `.registered_domain:example.co.uk` =\> to group traffic by web properties (e.g. [assets.example.co.uk](http://assets.example.co.uk), [downloads.example.co.uk](http://downloads.example.co.uk) are all part of [example.co.uk](http://example.co.uk))

---

<div class="post-metadata">

**Author:** ![Vincent\_Maury](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vincent_maury/32/59973_2.png) [@Vincent\_Maury](https://discuss.elastic.co/u/Vincent_Maury)\
**Post date:** [November 29, 2019, 7:23pm UTC](https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields-nested/209953/3 "2019-11-29T19:23:22Z")

</div>

Thank you very much @webmat  
I followed the .address advice and did the TLD parsing (with the custom plugin) and updated the configuration up there accordingly.  
I'm still unsure how to address your second point on query params. Once kv filter done, should I convert into json and then use an es template to force the flattened type?

---

<div class="post-metadata">

**Author:** ![webmat](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/webmat/32/46191_2.png) [@webmat](https://discuss.elastic.co/u/webmat)\
**Post date:** [November 29, 2019, 8:13pm UTC](https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields-nested/209953/4 "2019-11-29T20:13:36Z")

</div>

After the kv filter is done, you should have multiple keys nested under `[url][queryparams]`.

The only thing you would need to do, in order to avoid the mapping explosion, is modify your index template so that the field `url.queryparams` itself is of type `flattened`.

So assuming you're using the sample ECS template [we provide here](https://github.com/elastic/ecs/blob/master/generated/elasticsearch/7/template.json#L1982-L1985), you could add your custom field right below the definition for the `query` field, like this:

```auto
          "query": {
            "ignore_above": 1024, 
            "type": "keyword"
          }, 
          "queryparams": {
            "type": "flattened",
            // other params for the flattened type?
          }, 

```

Another option would indeed be to do as you describe. Turn the resulting structure into a big string where perhaps the `text` datatype could help dig in there.

But I would definitely give `flattened` a try first. It behaves somewhat like a bunch of `keyword` fields, but also avoids the mapping explosion.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 27, 2019, 8:13pm UTC](https://discuss.elastic.co/t/parsing-url-with-logstash-using-ecs-fields-nested/209953/5 "2019-12-27T20:13:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
