# Parsing url in log

**URL:** <https://discuss.elastic.co/t/parsing-url-in-log/227640>\
**Category:** Elasticsearch\
**Created:** [April 11, 2020, 10:41pm UTC](https://discuss.elastic.co/t/parsing-url-in-log/227640 "2020-04-11T22:41:37Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![thedraketaylor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thedraketaylor/32/52283_2.png) [@thedraketaylor](https://discuss.elastic.co/u/thedraketaylor)\
**Post date:** [April 11, 2020, 10:41pm UTC](https://discuss.elastic.co/t/parsing-url-in-log/227640/1 "2020-04-11T22:41:37Z")

</div>

I'm storing urls in the format "[http://domain.com/showthread.php?10357-thread-title-and-such/page22](http://domain.com/showthread.php?10357-thread-title-and-such/page22)" as "text". However, when I try to match the page number It fails to find anything. If I try to match anything before the last "/'" (such as "thread" or "title") it works perfectly.

Can someone point me to a way of being able to match "page"?

Thanks!

---

<div class="post-metadata">

**Author:** ![Luca\_Belluccini](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/luca_belluccini/32/33239_2.png) [@Luca\_Belluccini](https://discuss.elastic.co/u/Luca_Belluccini)\
**Post date:** [April 11, 2020, 11:46pm UTC](https://discuss.elastic.co/t/parsing-url-in-log/227640/2 "2020-04-11T23:46:38Z")

</div>

Hello @thedraketaylor

The default analyzer for `text` fields is the [standard](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-standard-analyzer.html) one.

To see which are the tokens generated by the `standard` analyser, you can use:

```auto
POST _analyze
{
  "analyzer": "standard", 
  "text": "http://domain.com/showthread.php?10357-thread-title-and-such/page22"
}
# Result
{
  "tokens" : [
    {
      "token" : "http",
      "start_offset" : 0,
      "end_offset" : 4,
      "type" : "<ALPHANUM>",
      "position" : 0
    },
    {
      "token" : "domain.com",
      "start_offset" : 7,
      "end_offset" : 17,
      "type" : "<ALPHANUM>",
      "position" : 1
    },
    {
      "token" : "showthread.php",
      "start_offset" : 18,
      "end_offset" : 32,
      "type" : "<ALPHANUM>",
      "position" : 2
    },
    {
      "token" : "10357",
      "start_offset" : 33,
      "end_offset" : 38,
      "type" : "<NUM>",
      "position" : 3
    },
    {
      "token" : "thread",
      "start_offset" : 39,
      "end_offset" : 45,
      "type" : "<ALPHANUM>",
      "position" : 4
    },
    {
      "token" : "title",
      "start_offset" : 46,
      "end_offset" : 51,
      "type" : "<ALPHANUM>",
      "position" : 5
    },
    {
      "token" : "and",
      "start_offset" : 52,
      "end_offset" : 55,
      "type" : "<ALPHANUM>",
      "position" : 6
    },
    {
      "token" : "such",
      "start_offset" : 56,
      "end_offset" : 60,
      "type" : "<ALPHANUM>",
      "position" : 7
    },
    {
      "token" : "page22",
      "start_offset" : 61,
      "end_offset" : 67,
      "type" : "<ALPHANUM>",
      "position" : 8
    }
  ]
}

```

You can test out `simple` with:

```auto
POST _analyze
{
  "analyzer": "simple", 
  "text": "http://domain.com/showthread.php?10357-thread-title-and-such/page22"
}
# Result
{
  "tokens" : [
    {
      "token" : "http",
      "start_offset" : 0,
      "end_offset" : 4,
      "type" : "word",
      "position" : 0
    },
    {
      "token" : "domain",
      "start_offset" : 7,
      "end_offset" : 13,
      "type" : "word",
      "position" : 1
    },
    {
      "token" : "com",
      "start_offset" : 14,
      "end_offset" : 17,
      "type" : "word",
      "position" : 2
    },
    {
      "token" : "showthread",
      "start_offset" : 18,
      "end_offset" : 28,
      "type" : "word",
      "position" : 3
    },
    {
      "token" : "php",
      "start_offset" : 29,
      "end_offset" : 32,
      "type" : "word",
      "position" : 4
    },
    {
      "token" : "thread",
      "start_offset" : 39,
      "end_offset" : 45,
      "type" : "word",
      "position" : 5
    },
    {
      "token" : "title",
      "start_offset" : 46,
      "end_offset" : 51,
      "type" : "word",
      "position" : 6
    },
    {
      "token" : "and",
      "start_offset" : 52,
      "end_offset" : 55,
      "type" : "word",
      "position" : 7
    },
    {
      "token" : "such",
      "start_offset" : 56,
      "end_offset" : 60,
      "type" : "word",
      "position" : 8
    },
    {
      "token" : "page",
      "start_offset" : 61,
      "end_offset" : 65,
      "type" : "word",
      "position" : 9
    }
  ]
}

```

If you're using well known URLs (for which you know their typical structure), you might use the `pattern` analyzer.

More information about this subject can be found [in our documentation](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis.html).

It is also possible to create a [custom analyzer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-custom-analyzer.html) and use it in your index.

---

<div class="post-metadata">

**Author:** ![thedraketaylor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thedraketaylor/32/52283_2.png) [@thedraketaylor](https://discuss.elastic.co/u/thedraketaylor)\
**Post date:** [April 12, 2020, 12:20am UTC](https://discuss.elastic.co/t/parsing-url-in-log/227640/3 "2020-04-12T00:20:55Z")

</div>

Thank you so much! That's given me a LOT to learn. I appreciate it!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 10, 2020, 12:32am UTC](https://discuss.elastic.co/t/parsing-url-in-log/227640/4 "2020-05-10T00:32:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
