# Refactoring a search

**URL:** <https://discuss.elastic.co/t/refactoring-a-search/6551>\
**Category:** Elasticsearch\
**Created:** [January 31, 2012, 7:44am UTC](https://discuss.elastic.co/t/refactoring-a-search/6551 "2012-01-31T07:44:45Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [January 31, 2012, 7:44am UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/1 "2012-01-31T07:44:45Z")

</div>

Hey guys. One of my queries is scoring documents strangely (IMO).  
Obviously, this means that my query and/or mapping needs some work. Would  
you mind giving me some advice on what type of query should be used here,  
please?

The query is for my web app's generic "search" bar, and returns products  
that match the search text.

Within each product document, I want to search through:

- name
- catalog.name
- items.name
- items.property\_attribs.character.analyzed
- \_all

Is a dis\_max query with field sub-queries ideal?

Here's a gist with more detail, because that usually helps:

> <https://gist.github.com/nickhoffman/df6321bdd0b6b5d599f8>
>
> There are more than three files. show original

Thanks again for your advice.  
Nick

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [January 31, 2012, 5:41pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/2 "2012-01-31T17:41:35Z")

</div>

Hard to tell what scores wrong. Did you try and boost search on fields are are more important if they match?

On Tuesday, January 31, 2012 at 9:44 AM, Nick Hoffman wrote:

> Hey guys. One of my queries is scoring documents strangely (IMO). Obviously, this means that my query and/or mapping needs some work. Would you mind giving me some advice on what type of query should be used here, please?
> 
> The query is for my web app's generic "search" bar, and returns products that match the search text.
> 
> Within each product document, I want to search through:
> 
> - name
> - catalog.name ([http://catalog.name](http://catalog.name))
> - items.name ([http://items.name](http://items.name))
> - items.property\_attribs.character.analyzed
> - \_all
> 
> Is a dis\_max query with field sub-queries ideal?
> 
> Here's a gist with more detail, because that usually helps:  
> [Is this query ideal for a generic search? · GitHub](https://gist.github.com/df6321bdd0b6b5d599f8)
> 
> Thanks again for your advice.  
> Nick

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [January 31, 2012, 6:30pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/3 "2012-01-31T18:30:22Z")

</div>

On Tuesday, 31 January 2012 12:41:35 UTC-5, kimchy wrote:

> Hard to tell what scores wrong. Did you try and boost search on fields  
> are are more important if they match?

Yeah, I'm boosting on the most important fields:

- by 4 on "name"
- by 2 on "items.name"
- by 2 on "items.property\_attribs.character.analyzed"

The 1st document in the gist has 1 occurrence of "Grimlock", while the 2nd  
document has 7 occurrences of "Grimlock". Despite this, the 1st document is  
scored higher than the 2nd document.

How can that be, considering that the 1st document matches on 1 field  
that's boosted by 4, whereas the 2nd document matches on 1 field that's  
boosted by 4, and 6 fields that're boosted by 2?

I just updated the gist with this info, and the query's explanation:

> <https://gist.github.com/nickhoffman/df6321bdd0b6b5d599f8>
>
> There are more than three files. show original

Thanks again for your help with this, Shay. I've spent hours trying to  
figure this out, but haven't made any progress.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 1, 2012, 10:01am UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/4 "2012-02-01T10:01:57Z")

</div>

I did not see you boosting it in the query you sent, use it there...

On Tuesday, January 31, 2012 at 9:44 AM, Nick Hoffman wrote:

> Hey guys. One of my queries is scoring documents strangely (IMO). Obviously, this means that my query and/or mapping needs some work. Would you mind giving me some advice on what type of query should be used here, please?
> 
> The query is for my web app's generic "search" bar, and returns products that match the search text.
> 
> Within each product document, I want to search through:
> 
> - name
> - catalog.name ([http://catalog.name](http://catalog.name))
> - items.name ([http://items.name](http://items.name))
> - items.property\_attribs.character.analyzed
> - \_all
> 
> Is a dis\_max query with field sub-queries ideal?
> 
> Here's a gist with more detail, because that usually helps:  
> [Is this query ideal for a generic search? · GitHub](https://gist.github.com/df6321bdd0b6b5d599f8)
> 
> Thanks again for your advice.  
> Nick

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 1, 2012, 3:07pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/5 "2012-02-01T15:07:04Z")

</div>

On Wednesday, 1 February 2012 05:01:57 UTC-5, kimchy wrote:

> I did not see you boosting it in the query you sent, use it there...

The mapping already boosts the fields, though. If I boost them in the  
query, too, wouldn't that apply the boost twice?

---

<div class="post-metadata">

**Author:** ![Jan\_Fiedler](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jan_fiedler/32/2518_2.png) [@Jan\_Fiedler](https://discuss.elastic.co/u/Jan_Fiedler)\
**Post date:** [February 1, 2012, 4:13pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/6 "2012-02-01T16:13:41Z")

</div>

I have not spent a lot of time on it but glancing over the gist I noticed  
the following: Your mapping has a boost of 4 for the \*analyzed \*version of  
the top level name field. Your query runs on the not-analyzed version and  
the boost will never kick in. Is this by intention ?

Your items.name is \*always \*analyzed via edge ngram (no separate analyzed  
version). This generates many tokens that will match your user input  
'Grimlock'. I would assume that these multiple hits on p1.items.name are  
rated higher than the plain exact hit on p2.name (I am ignoring the other  
fields matching for simplicity).

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 1, 2012, 9:55pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/7 "2012-02-01T21:55:38Z")

</div>

Hey Jan. Thanks for your help.

> I have not spent a lot of time on it but glancing over the gist I noticed  
> the following: Your mapping has a boost of 4 for the \*analyzed \*version  
> of the top level name field. Your query runs on the not-analyzed version  
> and the boost will never kick in. Is this by intention ?

Good catch. That was not intentional. I've swapped that around, and updated  
the gist.  
[Is this query ideal for a generic search? · GitHub](https://gist.github.com/df6321bdd0b6b5d599f8)

Your items.name is \*always \*analyzed via edge ngram (no separate analyzed

> version). This generates many tokens that will match your user input  
> 'Grimlock'. I would assume that these multiple hits on p1.items.name are  
> rated higher than the plain exact hit on p2.name (I am ignoring the other  
> fields matching for simplicity).

Ah, that makes sense. I've changed this to be:  
"index\_analyzer" : "ascii\_edge\_ngram",  
"search\_analyzer" : "ascii\_std",

Now there's only 2 occurrences of "items.name" in the search's explanation,  
but that's still one too many, right?

Also, I just noticed that the root-level "name" field isn't mentioned in  
the search's explanation. Why would that be?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 5, 2012, 10:45am UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/8 "2012-02-05T10:45:04Z")

</div>

Use boosting on the query side, indexing time boosting is usually too restrictive.

On Wednesday, February 1, 2012 at 11:55 PM, Nick Hoffman wrote:

> Hey Jan. Thanks for your help.
> 
> > I have not spent a lot of time on it but glancing over the gist I noticed the following: Your mapping has a boost of 4 for the analyzed version of the top level name field. Your query runs on the not-analyzed version and the boost will never kick in. Is this by intention ?
> 
> Good catch. That was not intentional. I've swapped that around, and updated the gist.  
> [Is this query ideal for a generic search? · GitHub](https://gist.github.com/df6321bdd0b6b5d599f8)
> 
> > Your items.name ([http://items.name/](http://items.name/)) is always analyzed via edge ngram (no separate analyzed version). This generates many tokens that will match your user input 'Grimlock'. I would assume that these multiple hits on p1.items.name ([http://p1.items.name/](http://p1.items.name/)) are rated higher than the plain exact hit on p2.name ([http://p2.name/](http://p2.name/)) (I am ignoring the other fields matching for simplicity).
> 
> Ah, that makes sense. I've changed this to be:  
> "index\_analyzer" : "ascii\_edge\_ngram",  
> "search\_analyzer" : "ascii\_std",
> 
> Now there's only 2 occurrences of "items.name ([http://items.name](http://items.name))" in the search's explanation, but that's still one too many, right?
> 
> Also, I just noticed that the root-level "name" field isn't mentioned in the search's explanation. Why would that be?

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 5, 2012, 5:50pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/9 "2012-02-05T17:50:15Z")

</div>

On Sunday, 5 February 2012 05:45:04 UTC-5, kimchy wrote:

> Use boosting on the query side, indexing time boosting is usually too  
> restrictive.

So boost values should be specified in queries instead of in mappings?

What's the difference, exactly? I looked around for info on this, but  
couldn't find anything.

Thanks, kimchy!

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 5, 2012, 5:58pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/10 "2012-02-05T17:58:16Z")

</div>

When you specify boost values at index time, they get stored in the index, but with a reduced resolution. Most times, its much better, if possible, to provide them at query time, which gives you the flexibility of changing them on the fly without needing to reindex.

On Sunday, February 5, 2012 at 7:50 PM, Nick Hoffman wrote:

> On Sunday, 5 February 2012 05:45:04 UTC-5, kimchy wrote:
> 
> > Use boosting on the query side, indexing time boosting is usually too restrictive.
> 
> So boost values should be specified in queries instead of in mappings?
> 
> What's the difference, exactly? I looked around for info on this, but couldn't find anything.
> 
> Thanks, kimchy!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 5, 2012, 7:01pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/11 "2012-02-05T19:01:21Z")

</div>

Very interesting !  
I will change my code tomorrow with this good advice !

David 😉  
@dadoonet

Le 5 févr. 2012 à 18:58, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) a écrit :

> When you specify boost values at index time, they get stored in the index, but with a reduced resolution. Most times, its much better, if possible, to provide them at query time, which gives you the flexibility of changing them on the fly without needing to reindex.  
> On Sunday, February 5, 2012 at 7:50 PM, Nick Hoffman wrote:
> 
> > On Sunday, 5 February 2012 05:45:04 UTC-5, kimchy wrote:
> > 
> > > Use boosting on the query side, indexing time boosting is usually too restrictive.
> > 
> > So boost values should be specified in queries instead of in mappings?
> > 
> > What's the difference, exactly? I looked around for info on this, but couldn't find anything.
> > 
> > Thanks, kimchy!

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 5, 2012, 10:02pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/12 "2012-02-05T22:02:50Z")

</div>

Interesting. Thanks for that insight, kimchy.

I'm boosting a field query inside of a dis\_max query, but ES is bailing. Is  
the "boost" option here not allowed?

curl -X DELETE 'localhost:9200/test?pretty=1'

curl -X POST 'localhost:9200/test/foo?pretty=1' -d '{ name: "Nick Hoffman"  
}'  
curl -X POST 'localhost:9200/test/foo?pretty=1' -d '{ name: "John Smith" }'  
curl -X POST 'localhost:9200/test/foo?pretty=1' -d '{ name: "Nick Other" }'

curl -X POST 'localhost:9200/test/\_refresh?pretty=1'

curl 'localhost:9200/test/foo/\_search?pretty=1' -d '{  
"query":{  
"dis\_max":{  
"queries":[  
{  
"field": { "name": "Nick", "boost": 4.0 }  
}  
]  
}  
}  
}'

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 5, 2012, 11:21pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/13 "2012-02-05T23:21:43Z")

</div>

For field query, you need to send the second format that allows for more options, see the second sample here: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/field-query.html). I suggest you use the text query though, a bit faster and has more options unless you want to support the Lucene query syntax.

On Monday, February 6, 2012 at 12:02 AM, Nick Hoffman wrote:

> Interesting. Thanks for that insight, kimchy.
> 
> I'm boosting a field query inside of a dis\_max query, but ES is bailing. Is the "boost" option here not allowed?
> 
> curl -X DELETE 'localhost:9200/test?pretty=1'
> 
> curl -X POST 'localhost:9200/test/foo?pretty=1' -d '{ name: "Nick Hoffman" }'  
> curl -X POST 'localhost:9200/test/foo?pretty=1' -d '{ name: "John Smith" }'  
> curl -X POST 'localhost:9200/test/foo?pretty=1' -d '{ name: "Nick Other" }'
> 
> curl -X POST 'localhost:9200/test/\_refresh?pretty=1'
> 
> curl 'localhost:9200/test/foo/\_search?pretty=1' -d '{  
> "query":{  
> "dis\_max":{  
> "queries":[  
> {  
> "field": { "name": "Nick", "boost": 4.0 }  
> }  
> ]  
> }  
> }  
> }'

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 6, 2012, 2:28am UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/14 "2012-02-06T02:28:04Z")

</div>

Thanks again, kimchy. However, when I specify the "boost" option, the score  
doesn't seem to change:

> <https://gist.github.com/nickhoffman/51059b4774ec9d33249c>

---

<div class="post-metadata">

**Author:** ![cole](https://avatars.discourse-cdn.com/v4/letter/c/53a042/32.png) [@cole](https://discuss.elastic.co/u/cole)\
**Post date:** [February 6, 2012, 11:37pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/15 "2012-02-06T23:37:18Z")

</div>

Hi Nick and Shay,

I ran into the same issue with specifying boost values on text queries  
a few days ago. As far as I can tell from digging through the mailing  
list archives, the boost value should be respected, but I haven't dug  
into the latest ES code to verify that. To unblock myself, I wrapped  
the text query I wanted to boost with a custom\_score query with an  
associated boost value. Nick, I forked your previous gist to give a  
simple custom\_score example: [Why does the "boost" not affect the score? · GitHub](https://gist.github.com/4a446763fa0c3d6c2fb0)

Shay: I noticed the text query boost was at least showing up in the  
"explain" output for some boost values when wrapped in a custom\_score  
query. The fifth entry  
("5\_queries\_with\_boost\_and\_custom\_score\_boost.txt") in the gist I  
linked to shows two examples, the first of which doesn't have any sign  
of the intended text query boost while the second at least has some  
sign of it in the explain section.

Thanks,  
Cole

On Feb 5, 6:28 pm, Nick Hoffman [n...@deadorange.com](mailto:n...@deadorange.com) wrote:

> Thanks again, kimchy. However, when I specify the "boost" option, the score  
> doesn't seem to change:[Why does the "boost" not affect the score? · GitHub](https://gist.github.com/51059b4774ec9d33249c)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 7, 2012, 11:19am UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/16 "2012-02-07T11:19:29Z")

</div>

Yes, this happens because the idf is 1 and the boosting gets normalized. We have a simply "take the score and multiple it by X" query called "custom\_boost\_factor", here is how you can use it: [gist:1759189 · GitHub](https://gist.github.com/1759189). But, for some reason its not documented in the site!, I will fix it shortly.

On Monday, February 6, 2012 at 4:28 AM, Nick Hoffman wrote:

> Thanks again, kimchy. However, when I specify the "boost" option, the score doesn't seem to change:  
> [Why does the "boost" not affect the score? · GitHub](https://gist.github.com/51059b4774ec9d33249c)

---

<div class="post-metadata">

**Author:** ![cole](https://avatars.discourse-cdn.com/v4/letter/c/53a042/32.png) [@cole](https://discuss.elastic.co/u/cole)\
**Post date:** [February 7, 2012, 6:47pm UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/17 "2012-02-07T18:47:25Z")

</div>

That's great. Thanks, Shay. The custom\_boost\_factor query looks much  
better than the custom\_score query I was using with "script" :  
"\_score". =)

-cole

On Feb 7, 3:19 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Yes, this happens because the idf is 1 and the boosting gets normalized. We have a simply "take the score and multiple it by X" query called "custom\_boost\_factor", here is how you can use it:[gist:1759189 · GitHub](https://gist.github.com/1759189). But, for some reason its not documented in the site!, I will fix it shortly.
> 
> On Monday, February 6, 2012 at 4:28 AM, Nick Hoffman wrote:
> 
> > Thanks again, kimchy. However, when I specify the "boost" option, the score doesn't seem to change:  
> > [Why does the "boost" not affect the score? · GitHub](https://gist.github.com/51059b4774ec9d33249c)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:40am UTC](https://discuss.elastic.co/t/refactoring-a-search/6551/18 "2017-07-06T03:40:12Z")

</div>


