# Search when url is one of the terms

**URL:** https://discuss.elastic.co/t/search-when-url-is-one-of-the-terms/6999
**Category:** Elasticsearch
**Created:** [March 14, 2012, 2:35pm UTC](https://discuss.elastic.co/t/search-when-url-is-one-of-the-terms/6999 "2012-03-14T14:35:12Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![Wiki](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wiki/32/2954_2.png) [@Wiki](https://discuss.elastic.co/u/Wiki)
#### Post date: [March 14, 2012, 2:35pm UTC](https://discuss.elastic.co/t/search-when-url-is-one-of-the-terms/6999/1 "2012-03-14T14:35:12Z")

</div>

Hi,

My index holds web metadata object, like:  
{  
:id =\> { :type =\> 'string', :store =\> true },  
:title =\> { :type =\> 'string', :analyzer =\>  
'snowball', :boost =\> 5 },  
:url =\> { :type =\> 'string', :index =\>  
'not\_analyzed', :store =\> true },  
:author =\> { :type =\> 'string', :index =\>  
'not\_analyzed', :store =\> true },  
:summary =\> { :type =\> 'string', :analyzer =\>  
'snowball', :boost =\> 10 },  
:content =\> { :type =\> 'string', :analyzer =\>  
'snowball', :boost =\> 8 },  
:published =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
:updated =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
:categories =\> { :type =\> 'string', :index =\>  
'not\_analyzed', :store =\> true },  
:image =\> { :type =\> 'string', :index =\>  
'not\_analyzed', :store =\> true },  
:site =\> { :type =\> 'string', :index =\>  
'snowball', :store =\> true }  
}

the id is the url.

1. whats the right indexing for the url\id field assuming I want to  
query by url? should I add analyzer to these fields?

2. I can't find a way to search for website data by it's url, what am  
I doing wrong? also tried encoding the url, it doesn't help:  
curl -XGET [http://localhost:9200/article-dev/\_search](http://localhost:9200/article-dev/_search) -d '{  
"query" : { "term" : { "\_id": "[http://thenextweb.com/media/2012/01/02/](http://thenextweb.com/media/2012/01/02/)  
uk-music-download-sales-grew-by-26-6-in-2011-but-the-industrys-still-  
in-decline/" } } }'

Thanks,

Viki

---

<div class="post-metadata">

### Author: ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)
#### Post date: [March 16, 2012, 12:02am UTC](https://discuss.elastic.co/t/search-when-url-is-one-of-the-terms/6999/2 "2012-03-16T00:02:30Z")

</div>

1. It all depends on what type of queries you want to execute. If the  
URL is simply a key, where you do not care to search inside the value,  
then the URL should not be analyzed. Analyzing URLs is difficult since  
they are not words. If you wanted to search inside the urls, there are  
many different routes to take, all of which requiring decomposing the  
data on the client side into smaller tokens.

2. Strings are analyzed by default in Elasticsearch, therefore your  
:id field will be analyzed when indexed. Are term query is not  
analyzed, so your search will not work. Your :url shoud work however.  
Does searching against that field work? If not, gist an example  
document.

--  
Ivan

On Wed, Mar 14, 2012 at 7:35 AM, Wiki [viki.rozental@gmail.com](mailto:viki.rozental@gmail.com) wrote:

> Hi,
> 
> My index holds web metadata object, like:  
> {  
> :id =\> { :type =\> 'string', :store =\> true },  
> :title =\> { :type =\> 'string', :analyzer =\>  
> 'snowball', :boost =\> 5 },  
> :url =\> { :type =\> 'string', :index =\>  
> 'not\_analyzed', :store =\> true },  
> :author =\> { :type =\> 'string', :index =\>  
> 'not\_analyzed', :store =\> true },  
> :summary =\> { :type =\> 'string', :analyzer =\>  
> 'snowball', :boost =\> 10 },  
> :content =\> { :type =\> 'string', :analyzer =\>  
> 'snowball', :boost =\> 8 },  
> :published =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
> HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
> :updated =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
> HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
> :categories =\> { :type =\> 'string', :index =\>  
> 'not\_analyzed', :store =\> true },  
> :image =\> { :type =\> 'string', :index =\>  
> 'not\_analyzed', :store =\> true },  
> :site =\> { :type =\> 'string', :index =\>  
> 'snowball', :store =\> true }  
> }
> 
> the id is the url.
> 
> 1. whats the right indexing for the url\id field assuming I want to  
> query by url? should I add analyzer to these fields?
> 
> 2. I can't find a way to search for website data by it's url, what am  
> I doing wrong? also tried encoding the url, it doesn't help:  
> curl -XGET [http://localhost:9200/article-dev/\_search](http://localhost:9200/article-dev/_search) -d '{  
> "query" : { "term" : { "\_id": "[Latest tech news | TNW](http://thenextweb.com/media/2012/01/02/)  
> uk-music-download-sales-grew-by-26-6-in-2011-but-the-industrys-still-  
> in-decline/" } } }'
> 
> Thanks,
> 
> Viki

---

<div class="post-metadata">

### Author: ![Wiki](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wiki/32/2954_2.png) [@Wiki](https://discuss.elastic.co/u/Wiki)
#### Post date: [March 20, 2012, 8:48am UTC](https://discuss.elastic.co/t/search-when-url-is-one-of-the-terms/6999/3 "2012-03-20T08:48:26Z")

</div>

Thanks for your response.

1. I don't need tokenizing the url, just need to search for it as a  
one single string.
2. searching against this field (:url) doesn't work:

{"\_index":"aunticles-dev-thenextweb","\_type":"article","\_id":"http://  
[thenextweb.com/apple/2012/02/23/forgotten-apple-founder-takes-to-](http://thenextweb.com/apple/2012/02/23/forgotten-apple-founder-takes-to-)  
facebook-to-explain-his-decision-to-quit-after-12-days/","\_score":1.0,  
"\_source" : {"id":"[http://thenextweb.com/apple/2012/02/23/forgotten-](http://thenextweb.com/apple/2012/02/23/forgotten-)  
apple-founder-takes-to-facebook-to-explain-his-decision-to-quit-  
after-12-days/","title":"Forgotten Apple founder takes to Facebook to  
explain his decision to quit after 12 days","summary":"Yesterday,  
third Apple founder Ron Wayne published an essay on Facebook about his  
decision to leave Apple Computer after only 12 days.","image":"http://  
[cdn.thenextweb.com/wp-content/blogs.dir/1/files/2012/02/](http://cdn.thenextweb.com/wp-content/blogs.dir/1/files/2012/02/)  
Photoxpress\_23083180-300x250.jpg","categories":"","published":"2012-03-14  
13:24:43","updated":null,"type":"article","site":null}

Thanks a lot!

On Mar 16, 2:02 am, Ivan Brusic [i...@brusic.com](mailto:i...@brusic.com) wrote:

> 1. It all depends on what type of queries you want to execute. If theURLis simply a key, where you do not care to search inside the value,  
> then theURLshould not be analyzed. Analyzing URLs is difficult since  
> they are not words. If you wanted to search inside the urls, there are  
> many different routes to take, all of which requiring decomposing the  
> data on the client side into smaller tokens.
> 
> 2. Strings are analyzed by default in Elasticsearch, therefore your  
> :id field will be analyzed when indexed. Are term query is not  
> analyzed, so your search will not work. Your :urlshoud work however.  
> Does searching against that field work? If not, gist an example  
> document.
> 
> --  
> Ivan
> 
> On Wed, Mar 14, 2012 at 7:35 AM, Wiki [viki.rozen...@gmail.com](mailto:viki.rozen...@gmail.com) wrote:
> 
> > Hi,
> 
> > My index holds web metadata object, like:  
> > {  
> > :id =\> { :type =\> 'string', :store =\> true },  
> > :title =\> { :type =\> 'string', :analyzer =\>  
> > 'snowball', :boost =\> 5 },  
> > :url =\> { :type =\> 'string', :index =\>  
> > 'not\_analyzed', :store =\> true },  
> > :author =\> { :type =\> 'string', :index =\>  
> > 'not\_analyzed', :store =\> true },  
> > :summary =\> { :type =\> 'string', :analyzer =\>  
> > 'snowball', :boost =\> 10 },  
> > :content =\> { :type =\> 'string', :analyzer =\>  
> > 'snowball', :boost =\> 8 },  
> > :published =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
> > HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
> > :updated =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
> > HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
> > :categories =\> { :type =\> 'string', :index =\>  
> > 'not\_analyzed', :store =\> true },  
> > :image =\> { :type =\> 'string', :index =\>  
> > 'not\_analyzed', :store =\> true },  
> > :site =\> { :type =\> 'string', :index =\>  
> > 'snowball', :store =\> true }  
> > }
> 
> > the id is theurl.
> > 
> > 1. whats the right indexing for theurl\id field assuming I want to  
> > query byurl? should I add analyzer to these fields?
> 
> > 1. I can't find a way to search for website data by it'surl, what am  
> > I doing wrong? also tried encoding theurl, it doesn't help:  
> > curl -XGEThttp://localhost:9200/article-dev/\_search-d '{  
> > "query" : { "term" : { "\_id": "[Latest tech news | TNW](http://thenextweb.com/media/2012/01/02/)  
> > uk-music-download-sales-grew-by-26-6-in-2011-but-the-industrys-still-  
> > in-decline/" } } }'
> 
> > Thanks,
> 
> > Viki

---

<div class="post-metadata">

### Author: ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)
#### Post date: [March 20, 2012, 10:56am UTC](https://discuss.elastic.co/t/search-when-url-is-one-of-the-terms/6999/4 "2012-03-20T10:56:59Z")

</div>

Can you gist a full curl recreation, including indexing the relevant data  
nad setting up the mapping? (see [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help)).

A few notes:

1. You don't need to explicitly store each field, by default, the whole  
\_source json document you indexed is stored, so you end up storing things  
twice, once each individual field, and once the whole doc).

2. If the url field is not analyzed, you can look it up using the same url  
it was indexed with using a term query.

On Tue, Mar 20, 2012 at 10:48 AM, Wiki [viki.rozental@gmail.com](mailto:viki.rozental@gmail.com) wrote:

> Thanks for your response.
> 
> 1. I don't need tokenizing the url, just need to search for it as a  
> one single string.
> 2. searching against this field (:url) doesn't work:
> 
> {"\_index":"aunticles-dev-thenextweb","\_type":"article","\_id":"http://  
> [thenextweb.com/apple/2012/02/23/forgotten-apple-founder-takes-to-](http://thenextweb.com/apple/2012/02/23/forgotten-apple-founder-takes-to-)  
> facebook-to-explain-his-decision-to-quit-after-12-days/","\_score":1.0,  
> "\_source" : {"id":"[http://thenextweb.com/apple/2012/02/23/forgotten-](http://thenextweb.com/apple/2012/02/23/forgotten-)  
> apple-founder-takes-to-facebook-to-explain-his-decision-to-quit-  
> after-12-days/","title":"Forgotten Apple founder takes to Facebook to  
> explain his decision to quit after 12 days","summary":"Yesterday,  
> third Apple founder Ron Wayne published an essay on Facebook about his  
> decision to leave Apple Computer after only 12 days.","image":"http://  
> [cdn.thenextweb.com/wp-content/blogs.dir/1/files/2012/02/](http://cdn.thenextweb.com/wp-content/blogs.dir/1/files/2012/02/)  
> Photoxpress\_23083180-300x250.jpg","categories":"","published":"2012-03-14  
> 13:24:43","updated":null,"type":"article","site":null}
> 
> Thanks a lot!
> 
> On Mar 16, 2:02 am, Ivan Brusic [i...@brusic.com](mailto:i...@brusic.com) wrote:
> 
> > 1. It all depends on what type of queries you want to execute. If  
> > theURLis simply a key, where you do not care to search inside the value,  
> > then theURLshould not be analyzed. Analyzing URLs is difficult since  
> > they are not words. If you wanted to search inside the urls, there are  
> > many different routes to take, all of which requiring decomposing the  
> > data on the client side into smaller tokens.
> > 
> > 2. Strings are analyzed by default in Elasticsearch, therefore your  
> > :id field will be analyzed when indexed. Are term query is not  
> > analyzed, so your search will not work. Your :urlshoud work however.  
> > Does searching against that field work? If not, gist an example  
> > document.
> > 
> > --  
> > Ivan
> > 
> > On Wed, Mar 14, 2012 at 7:35 AM, Wiki [viki.rozen...@gmail.com](mailto:viki.rozen...@gmail.com) wrote:
> > 
> > > Hi,
> > 
> > > My index holds web metadata object, like:  
> > > {  
> > > :id =\> { :type =\> 'string', :store =\> true },  
> > > :title =\> { :type =\> 'string', :analyzer =\>  
> > > 'snowball', :boost =\> 5 },  
> > > :url =\> { :type =\> 'string', :index =\>  
> > > 'not\_analyzed', :store =\> true },  
> > > :author =\> { :type =\> 'string', :index =\>  
> > > 'not\_analyzed', :store =\> true },  
> > > :summary =\> { :type =\> 'string', :analyzer =\>  
> > > 'snowball', :boost =\> 10 },  
> > > :content =\> { :type =\> 'string', :analyzer =\>  
> > > 'snowball', :boost =\> 8 },  
> > > :published =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
> > > HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
> > > :updated =\> { :type =\> 'date', :format =\> 'yyyy-MM-dd  
> > > HH:mm:ss', :index =\> 'not\_analyzed', :store =\> true },  
> > > :categories =\> { :type =\> 'string', :index =\>  
> > > 'not\_analyzed', :store =\> true },  
> > > :image =\> { :type =\> 'string', :index =\>  
> > > 'not\_analyzed', :store =\> true },  
> > > :site =\> { :type =\> 'string', :index =\>  
> > > 'snowball', :store =\> true }  
> > > }
> > 
> > > the id is theurl.
> > > 
> > > 1. whats the right indexing for theurl\id field assuming I want to  
> > > query byurl? should I add analyzer to these fields?
> > 
> > > 1. I can't find a way to search for website data by it'surl, what am  
> > > I doing wrong? also tried encoding theurl, it doesn't help:  
> > > curl -XGEThttp://localhost:9200/article-dev/\_search-d '{  
> > > "query" : { "term" : { "\_id": "[Latest tech news | TNW](http://thenextweb.com/media/2012/01/02/)  
> > > uk-music-download-sales-grew-by-26-6-in-2011-but-the-industrys-still-  
> > > in-decline/" } } }'
> > 
> > > Thanks,
> > 
> > > Viki

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 3:35am UTC](https://discuss.elastic.co/t/search-when-url-is-one-of-the-terms/6999/5 "2017-07-06T03:35:29Z")

</div>


