# The relation of the amount of indexed documents, max\_expansions and prefix\_length in text\_phrase\_prefix

**URL:** <https://discuss.elastic.co/t/the-relation-of-the-amount-of-indexed-documents-max-expansions-and-prefix-length-in-text-phrase-prefix/9883>\
**Category:** Elasticsearch\
**Created:** [November 29, 2012, 12:02pm UTC](https://discuss.elastic.co/t/the-relation-of-the-amount-of-indexed-documents-max-expansions-and-prefix-length-in-text-phrase-prefix/9883 "2012-11-29T12:02:57Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![pcdinh](https://avatars.discourse-cdn.com/v4/letter/p/76d3ee/32.png) [@pcdinh](https://discuss.elastic.co/u/pcdinh)\
**Post date:** [November 29, 2012, 12:02pm UTC](https://discuss.elastic.co/t/the-relation-of-the-amount-of-indexed-documents-max-expansions-and-prefix-length-in-text-phrase-prefix/9883/1 "2012-11-29T12:02:57Z")

</div>

Hi all,

I have an issue with ES 0.19.3 regarding to text\_phrase\_prefix query. When  
number of documents indexed in ES is small the following query  
works perfectly (type "New Y", not "New York" or "New Yo")

curl -X DELETE [http://es1:9200/cities](http://es1:9200/cities)  
curl -X POST "[http://es1:9200/cities/city](http://es1:9200/cities/city)" -d '{ "city" : "New York" }'  
curl -X POST "[http://es1:9200/cities/city](http://es1:9200/cities/city)" -d '{ "city" : "North New York"  
}'  
curl -X POST "[http://es1:9200/cities/city](http://es1:9200/cities/city)" -d '{ "city" : "East New York" }'

curl -XGET [http://es1:9200/cities/city/\_search?pretty=true](http://es1:9200/cities/city/_search?pretty=true) -d'  
{  
"fields":[  
"city"  
],  
"query":{  
"text\_phrase\_prefix": {  
"city" : {  
"query": "New Y",  
"max\_expansions": 2,  
"prefix\_length": 2  
}  
}  
},  
"from":0,  
"size":20  
}'

returns all cities or areas has "New York" in their names.

{  
"took" : 2,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : {  
"total" : 3,  
"max\_score" : 0.38356602,  
"hits" : [ {  
"\_index" : "cities",  
"\_type" : "city",  
"\_id" : "4ObkgggqS7uou1XLdwOkfA",  
"\_score" : 0.38356602,  
"fields" : {  
"city" : "New York"  
}  
}, {  
"\_index" : "cities",  
"\_type" : "city",  
"\_id" : "CZutMgvwSfa8O79Vajkshg",  
"\_score" : 0.30685282,  
"fields" : {  
"city" : "North New York"  
}  
}, {  
"\_index" : "cities",  
"\_type" : "city",  
"\_id" : "ZGA3gno9QnOIBg2MxxsPbg",  
"\_score" : 0.30685282,  
"fields" : {  
"city" : "East New York"  
}  
} ]  
}  
}

However when the number of indexed document grows up (more than 30 000  
cities or towns or areas in US), the above query does not work any more.  
I need to increase max\_expansions into a number that is greater than 17 (18  
and greater to be specific) to make it work again. Any number that  
is smaller than 17 does not work. If I don't increase max\_expansions, I  
need to use keywords like: "New Yo" or "New York"

curl -XGET [http://184.72.29.x:9200/cities/city/\_search?pretty=true](http://184.72.29.x:9200/cities/city/_search?pretty=true) -d'  
{  
"fields":[  
"area\_label"  
],  
"query":{  
"text\_phrase\_prefix": {  
"area\_label" : {  
"query": "New Y",  
"max\_expansions": 18,  
"prefix\_length": 2  
}  
}  
},  
"from":0,  
"size":20  
}  
'

returns

{  
"took" : 7,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : {  
"total" : 3,  
"max\_score" : 111.86473,  
"hits" : [ {  
"\_index" : "cities",  
"\_type" : "city",  
"\_id" : "195232",  
"\_score" : 111.86473,  
"fields" : {  
"area\_label" : "New York"  
}  
}, {  
"\_index" : "cities",  
"\_type" : "city",  
"\_id" : "46727",  
"\_score" : 89.49178,  
"fields" : {  
"area\_label" : "North New York"  
}  
}, {  
"\_index" : "cities",  
"\_type" : "city",  
"\_id" : "46772",  
"\_score" : 89.49178,  
"fields" : {  
"area\_label" : "East New York"  
}  
} ]  
}  
}

prefix\_length does not play any role in this case. I increase the value of  
prefix\_length to 20, the result is still the same.

I don't understand why the number of 18 is magic in this case. I guess that  
there is a relationship between max\_expansions and the number of  
indexed document. So when the amount of indexed documents increases, I need  
to increase max\_expansions too or the above query does not work again.

Am I missing something?

Regards,

Dinh

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [November 29, 2012, 12:49pm UTC](https://discuss.elastic.co/t/the-relation-of-the-amount-of-indexed-documents-max-expansions-and-prefix-length-in-text-phrase-prefix/9883/2 "2012-11-29T12:49:36Z")

</div>

Hi Dinh

> However when the number of indexed document grows up (more than 30 000  
> cities or towns or areas in US), the above query does not work any  
> more.  
> I need to increase max\_expansions into a number that is greater than  
> 17 (18 and greater to be specific) to make it work again. Any number  
> that  
> is smaller than 17 does not work. If I don't increase max\_expansions,  
> I need to use keywords like: "New Yo" or "New York"

> prefix\_length does not play any role in this case. I increase the  
> value of prefix\_length to 20, the result is still the same.

prefix\_length is for fuzzy queries, not phrase\_prefix

> I don't understand why the number of 18 is magic in this case. I guess  
> that there is a relationship between max\_expansions and the number of  
> indexed document. So when the amount of indexed documents increases, I  
> need to increase max\_expansions too or the above query does not work  
> again.

To build a prefix query, it looks for all terms starting with your  
prefix 'y' and adds each term to your query, up to max\_expansions. The  
terms are sorted alphabetically.

If you've indexed lots more data, than presumably you have a lot more  
terms between "ya" and "yo" than you had before.

if you want to do partial matching of words, then a better idea is to  
use ngrams or edge-ngrams to index your data up front. it is more  
efficient at search time than using a prefix query.

clint

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:02am UTC](https://discuss.elastic.co/t/the-relation-of-the-amount-of-indexed-documents-max-expansions-and-prefix-length-in-text-phrase-prefix/9883/3 "2017-07-06T03:02:13Z")

</div>


