# Partial match of sub-phrases to be scored higher?

**URL:** <https://discuss.elastic.co/t/partial-match-of-sub-phrases-to-be-scored-higher/13621>\
**Category:** Elasticsearch\
**Created:** [September 16, 2013, 2:08pm UTC](https://discuss.elastic.co/t/partial-match-of-sub-phrases-to-be-scored-higher/13621 "2013-09-16T14:08:40Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ark](https://avatars.discourse-cdn.com/v4/letter/a/94ad74/32.png) [@Ark](https://discuss.elastic.co/u/Ark)\
**Post date:** [September 16, 2013, 2:08pm UTC](https://discuss.elastic.co/t/partial-match-of-sub-phrases-to-be-scored-higher/13621/1 "2013-09-16T14:08:40Z")

</div>

Hello,

How can I model the query and/or mapping so that a partial match of a  
sub-phrase has an higher score than what a edgengram would return?

For example, If I have four documents:

1. foo bar blah
2. foo blah bar
3. bar foo blah
4. bar blah foo

If the search string is "bar bl", I would like document 1 and 4 should be  
scored higher than document 2 and 3.

If the field is indexed using edgengram, all 4 documents would match (which  
is fine for my use-case) but I think the scoring cannot yield the result I  
am looking for.

There is also a "match\_phrase\_prefix" but that would match only #4.

Thanks  
Ark

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [September 16, 2013, 6:23pm UTC](https://discuss.elastic.co/t/partial-match-of-sub-phrases-to-be-scored-higher/13621/2 "2013-09-16T18:23:40Z")

</div>

On Mon, Sep 16, 2013 at 4:08 PM, Ark [ayam12yeh34@gmail.com](mailto:ayam12yeh34@gmail.com) wrote:

> Hello,

Hi,

> How can I model the query and/or mapping so that a partial match of a  
> sub-phrase has an higher score than what a edgengram would return?
> 
> For example, If I have four documents:
> 
> 1. foo bar blah
> 2. foo blah bar
> 3. bar foo blah
> 4. bar blah foo
> 
> If the search string is "bar bl", I would like document 1 and 4 should be  
> scored higher than document 2 and 3.
> 
> If the field is indexed using edgengram, all 4 documents would match  
> (which is fine for my use-case) but I think the scoring cannot yield the  
> result I am looking for.
> 
> There is also a "match\_phrase\_prefix" but that would match only #4.

You could use the edgeNGram filter on top of the shingle[1] filter (with  
output\_unigrams=false). This would allow you to boost on prefixes and  
positions at the same time.

The fact that you are interested in prefix matches makes me wonder whether  
you are trying to implement auto-completion: if this is the case, a better  
option could be to use the completion suggest[2] (which is way faster than  
any index-based solution) and use all suffixes of your text as inputs. For  
example, the "foo bar blah" suggestion could be indexed with "input": ["foo  
bar blah", "bar blah", "blah"]. If you are not trying to implement  
auto-completion, you can safely ignore this comment. 🙂

[1]

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

[2]

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

--  
Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ark](https://avatars.discourse-cdn.com/v4/letter/a/94ad74/32.png) [@Ark](https://discuss.elastic.co/u/Ark)\
**Post date:** [September 16, 2013, 9:09pm UTC](https://discuss.elastic.co/t/partial-match-of-sub-phrases-to-be-scored-higher/13621/3 "2013-09-16T21:09:59Z")

</div>

Thank you! I am not sure if my case falls into auto-complete - as I am just  
now learning some concepts and looking at what is possible. But having said  
that, your suggestion does sound interesting and possibly something that  
may be a good fit.

Ark

On Monday, September 16, 2013 1:23:40 PM UTC-5, Adrien Grand wrote:

> On Mon, Sep 16, 2013 at 4:08 PM, Ark \<[ayam1...@gmail.com](mailto:ayam1...@gmail.com) \<javascript:\>\>wrote:
> 
> > Hello,
> 
> Hi,
> 
> > How can I model the query and/or mapping so that a partial match of a  
> > sub-phrase has an higher score than what a edgengram would return?
> > 
> > For example, If I have four documents:
> > 
> > 1. foo bar blah
> > 2. foo blah bar
> > 3. bar foo blah
> > 4. bar blah foo
> > 
> > If the search string is "bar bl", I would like document 1 and 4 should be  
> > scored higher than document 2 and 3.
> > 
> > If the field is indexed using edgengram, all 4 documents would match  
> > (which is fine for my use-case) but I think the scoring cannot yield the  
> > result I am looking for.
> > 
> > There is also a "match\_phrase\_prefix" but that would match only #4.
> 
> You could use the edgeNGram filter on top of the shingle[1] filter (with  
> output\_unigrams=false). This would allow you to boost on prefixes and  
> positions at the same time.
> 
> The fact that you are interested in prefix matches makes me wonder whether  
> you are trying to implement auto-completion: if this is the case, a better  
> option could be to use the completion suggest[2] (which is way faster than  
> any index-based solution) and use all suffixes of your text as inputs. For  
> example, the "foo bar blah" suggestion could be indexed with "input": ["foo  
> bar blah", "bar blah", "blah"]. If you are not trying to implement  
> auto-completion, you can safely ignore this comment. 🙂
> 
> [1]  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/shingle-tokenfilter/)  
> [2]  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/api/search/completion-suggest/)
> 
> --  
> Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:16am UTC](https://discuss.elastic.co/t/partial-match-of-sub-phrases-to-be-scored-higher/13621/4 "2017-07-06T02:16:12Z")

</div>


