# Match phrase prefix in set of documents

**URL:** <https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570>\
**Category:** Elasticsearch\
**Created:** [December 12, 2018, 3:28pm UTC](https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570 "2018-12-12T15:28:09Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ozymandy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ozymandy/32/44522_2.png) [@Ozymandy](https://discuss.elastic.co/u/Ozymandy)\
**Post date:** [December 12, 2018, 3:28pm UTC](https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570/1 "2018-12-12T15:28:09Z")

</div>

Hi all,  
I am using multi match phrase\_prefix for matching my documents. And it works fine when I need to extract matched documents by whole index. But it does not work properly for case when I need to query by prefix in set of documents(which is found by document field id). For examples I'm querying by all fields of three documents with prefix "b". All documents have at least one word started with letter "b": "baby", "bed" and "bored". But query hits to only one document with word "baby" I executed my query with explain flag and found out that it happens because of fuzzy search. It loads only small part of tokens which suits my query (started with "b") from whole index. Of course i can increase value of _max\_expansion_ but it gives me performance issues. So regarding the above description I have question how to make prefix\_phrase search work:

1. Can I limit the scope of documents where it search for tokens? It should be like subquery in SQL(select from (select from) where str like).  
2)May be it could be done by another analyzer? For now I am using whitespace tokenizer with lowercase filter.

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [December 13, 2018, 9:41am UTC](https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570/2 "2018-12-13T09:41:30Z")

</div>

Hi ,

Can you print your query here ?

bye,  
Xavier

---

<div class="post-metadata">

**Author:** ![Ozymandy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ozymandy/32/44522_2.png) [@Ozymandy](https://discuss.elastic.co/u/Ozymandy)\
**Post date:** [December 13, 2018, 1:13pm UTC](https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570/3 "2018-12-13T13:13:34Z")

</div>

Sure.Here is query I execute: {"from":0,"size":10000,"query":{"bool":{"filter":[{"multi\_match":{"query":"b","fields":["category^1.0","body^1.0","title^1.0"],"type":"phrase\_prefix","operator":"OR","slop":0,"prefix\_length":0,"max\_expansions":50,"zero\_terms\_query":"NONE","auto\_generate\_synonyms\_phrase\_query":true,"fuzzy\_transpositions":true,"boost":1.0}},{"terms":{"id":["1", "2", "3"],"boost":1.0}}],"boost":1.0}}],"adjust\_pure\_negative":true,"boost":1.0}},"sort":[{"published":{"order":"desc"}}]}

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [December 13, 2018, 4:02pm UTC](https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570/4 "2018-12-13T16:02:05Z")

</div>

> [@Ozymandy](#):
>
> Can I limit the scope of documents where it search for tokens? It should be like subquery in SQL(select from (select from) where str like).

That what you have in your filters with the _terms_ query on field "id", it sounds good.

Regarding the phrase\_prefix query, I'm not very aware so I don't have solution.  
Just an idea: _phrase\_prefix_ returns document having "phrase" with a "b" at the beginning ? If yes, is the case of your 2 documents not returned ?

e.g.: "My phrase with a baby" ? =\> ok or not ?

---

<div class="post-metadata">

**Author:** ![Ozymandy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ozymandy/32/44522_2.png) [@Ozymandy](https://discuss.elastic.co/u/Ozymandy)\
**Post date:** [December 13, 2018, 9:34pm UTC](https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570/5 "2018-12-13T21:34:17Z")

</div>

As I said elasticsearch tries to load ALL tokens from ALL documents which suits prefix "b" according programmed max number (max\_expansion parameter). So if I have about 100 thousands tokens started with "b" elasticsearch cannot process them all.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 10, 2019, 9:43pm UTC](https://discuss.elastic.co/t/match-phrase-prefix-in-set-of-documents/160570/6 "2019-01-10T21:43:25Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
