# URL analysis

**URL:** <https://discuss.elastic.co/t/url-analysis/12091>\
**Category:** Elasticsearch\
**Created:** [May 23, 2013, 9:44am UTC](https://discuss.elastic.co/t/url-analysis/12091 "2013-05-23T09:44:59Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![dantuff](https://avatars.discourse-cdn.com/v4/letter/d/e79b87/32.png) [@dantuff](https://discuss.elastic.co/u/dantuff)\
**Post date:** [May 23, 2013, 9:44am UTC](https://discuss.elastic.co/t/url-analysis/12091/1 "2013-05-23T09:44:59Z")

</div>

Hi

I have a URL that contains an ID. I would like extract the ID during  
analysis so that the ID part of the URL is searchable, I would like to it  
to have its own property in the index 'productId' and store it so it can be  
returned

the url:

[http://host](http://host):port/path/1234/path

all of the other parts of the URL are not numeric.

Is there a way to do this in elasticSearch?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [May 23, 2013, 10:04am UTC](https://discuss.elastic.co/t/url-analysis/12091/2 "2013-05-23T10:04:36Z")

</div>

Hey,

you may want to take a look at the pattern analyzer (if you know the the  
structure of your URLs)

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

--Alex

On Thu, May 23, 2013 at 11:44 AM, es newbie [dan.tuffery@gmail.com](mailto:dan.tuffery@gmail.com) wrote:

> Hi
> 
> I have a URL that contains an ID. I would like extract the ID during  
> analysis so that the ID part of the URL is searchable, I would like to it  
> to have its own property in the index 'productId' and store it so it can be  
> returned
> 
> the url:
> 
> [http://host](http://host):port/path/1234/path
> 
> all of the other parts of the URL are not numeric.
> 
> Is there a way to do this in elasticSearch?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dantuff](https://avatars.discourse-cdn.com/v4/letter/d/e79b87/32.png) [@dantuff](https://discuss.elastic.co/u/dantuff)\
**Post date:** [May 23, 2013, 11:09am UTC](https://discuss.elastic.co/t/url-analysis/12091/3 "2013-05-23T11:09:02Z")

</div>

Thanks Alex.

Would the following split on forward slash?

"url\_index": {  
"pattern": "/",  
"lowercase": false,  
"type": "pattern"  
}

On Thursday, May 23, 2013 11:04:36 AM UTC+1, Alexander Reelsen wrote:

> Hey,
> 
> you may want to take a look at the pattern analyzer (if you know the the  
> structure of your URLs)
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/pattern-analyzer/)
> 
> --Alex
> 
> On Thu, May 23, 2013 at 11:44 AM, es newbie \<[dan.t...@gmail.com](mailto:dan.t...@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > Hi
> > 
> > I have a URL that contains an ID. I would like extract the ID during  
> > analysis so that the ID part of the URL is searchable, I would like to it  
> > to have its own property in the index 'productId' and store it so it can be  
> > returned
> > 
> > the url:
> > 
> > [http://host](http://host):port/path/1234/path
> > 
> > all of the other parts of the URL are not numeric.
> > 
> > Is there a way to do this in elasticSearch?
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 23, 2013, 11:22am UTC](https://discuss.elastic.co/t/url-analysis/12091/4 "2013-05-23T11:22:34Z")

</div>

Given that you want to extract just the ID from the URL, rather than all of  
the parts, you can use the pattern tokenizer to capture just the first  
capture group:

curl -XPUT '[http://127.0.0.1:9200/test/?pretty=1](http://127.0.0.1:9200/test/?pretty=1)' -d '  
{  
"settings" : {  
"analysis" : {  
"tokenizer" : {  
"extract\_id" : {  
"pattern" : "/([0-9]+)(/|$)",  
"group" : 1,  
"type" : "pattern"  
}  
},  
"analyzer" : {  
"extract\_id" : {  
"tokenizer" : "extract\_id"  
}  
}  
}  
}  
}  
'

This pattern "/([0-9]+)(/|$)" matches:

- /
- followed by 1 or more numbers
- followed by a / or the end of string $

You can test this out as:

curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)' -d  
'  
[http://foo:123/path/111/](http://foo:123/path/111/)  
'

-\> 111

curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)' -d  
'  
[http://foo:123/path/111](http://foo:123/path/111)  
'

-\> 111

curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)' -d  
'  
[http://foo:123/path/111/foo](http://foo:123/path/111/foo)  
'

-\> 111

curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)' -d  
'  
[http://foo:123/path/111aa/](http://foo:123/path/111aa/)  
'

-\> no match

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dantuff](https://avatars.discourse-cdn.com/v4/letter/d/e79b87/32.png) [@dantuff](https://discuss.elastic.co/u/dantuff)\
**Post date:** [May 23, 2013, 4:18pm UTC](https://discuss.elastic.co/t/url-analysis/12091/5 "2013-05-23T16:18:54Z")

</div>

Thanks Clinton, that works.

On Thursday, May 23, 2013 12:22:34 PM UTC+1, Clinton Gormley wrote:

> Given that you want to extract just the ID from the URL, rather than all  
> of the parts, you can use the pattern tokenizer to capture just the first  
> capture group:
> 
> curl -XPUT '[http://127.0.0.1:9200/test/?pretty=1](http://127.0.0.1:9200/test/?pretty=1)' -d '  
> {  
> "settings" : {  
> "analysis" : {  
> "tokenizer" : {  
> "extract\_id" : {  
> "pattern" : "/([0-9]+)(/|$)",  
> "group" : 1,  
> "type" : "pattern"  
> }  
> },  
> "analyzer" : {  
> "extract\_id" : {  
> "tokenizer" : "extract\_id"  
> }  
> }  
> }  
> }  
> }  
> '
> 
> This pattern "/([0-9]+)(/|$)" matches:
> 
> - /
> - followed by 1 or more numbers
> - followed by a / or the end of string $
> 
> You can test this out as:
> 
> curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)'  
> -d '  
> [http://foo:123/path/111/](http://foo:123/path/111/)  
> '
> 
> -\> 111
> 
> curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)'  
> -d '  
> [http://foo:123/path/111](http://foo:123/path/111)  
> '
> 
> -\> 111
> 
> curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)'  
> -d '  
> [http://foo:123/path/111/foo](http://foo:123/path/111/foo)  
> '
> 
> -\> 111
> 
> curl '[http://127.0.0.1:9200/test/\_analyze?pretty=1&&analyzer=extract\_id](http://127.0.0.1:9200/test/_analyze?pretty=1&&analyzer=extract_id)'  
> -d '  
> [http://foo:123/path/111aa/](http://foo:123/path/111aa/)  
> '
> 
> -\> no match

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:35am UTC](https://discuss.elastic.co/t/url-analysis/12091/6 "2017-07-06T02:35:02Z")

</div>


