# Speed of query with many filters

**URL:** https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665
**Category:** Elasticsearch
**Created:** [June 20, 2011, 11:06pm UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665 "2011-06-20T23:06:20Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)
#### Post date: [June 20, 2011, 11:06pm UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665/1 "2011-06-20T23:06:20Z")

</div>

Hi everybody.

We have big index "contacts" which size is about 3.5Gb (I mean "primary\_size") with 6,448,782 documents. We have performance problems with particular queries. Their execution time is \>5 seconds.

Our index configuration is 2 replicas, 3 shards. It's run on 3 EC2 m1.large servers.  
All fields in index document are not analyzed, \_source and \_all are disabled.  
Document has 2 big fields: "fields" and "reverse\_fields". Last one is reverse version of "fields", it is designed to match words from the end.

Here is part of "fields": mapping:  
...  
"fields": {  
"type": "object",  
"dynamic" : False,  
"properties" : {  
.....  
"city": {  
"type": "string",  
"index": "not\_analyzed",  
"omit\_term\_freq\_and\_positions": "true"  
},  
"state": {  
"type": "string",  
"index": "not\_analyzed",  
"omit\_term\_freq\_and\_positions": "true"  
},  
"zip": {  
"type": "string",  
"index": "not\_analyzed",  
"omit\_term\_freq\_and\_positions": "true"  
},  
....  
}  
}  
.....  
}

We looking for ways to speed up our "contain" query over all document fields. First 2 filter terms (company\_id and is\_visible) match ~43k documents.  
Example of the query:

es\_q = {'sort':  
[{'\_score': 'desc'}], 'query': {'filtered': {'filter': {'and': {  
'filters': [{'term': {'company\_id': '4b619ddffa5bd81b71000002'}}, {'term': {'is\_visible': True}}, {  
'or': [{'prefix': {'fields.last\_name': 'alexander'}}, {'prefix': {'reverse\_fields.last\_name': 'rednaxela'}},  
{'prefix': {'fields.twitter.profile': 'alexander'}},  
{'prefix': {'reverse\_fields.twitter.profile': 'rednaxela'}},  
{'prefix': {'fields.twitter.user\_name': 'alexander'}},  
{'prefix': {'reverse\_fields.twitter.user\_name': 'rednaxela'}},  
{'prefix': {'fields.twitter.user\_id': 'Alexander'}},  
{'prefix': {'reverse\_fields.twitter.user\_id': 'rednaxelA'}},  
{'prefix': {'fields.linkedin.profile': 'alexander'}},  
{'prefix': {'reverse\_fields.linkedin.profile': 'rednaxela'}},  
{'prefix': {'fields.linkedin.user\_name': 'alexander'}},  
{'prefix': {'reverse\_fields.linkedin.user\_name': 'rednaxela'}},  
{'prefix': {'fields.linkedin.user\_id': 'Alexander'}},  
{'prefix': {'reverse\_fields.linkedin.user\_id': 'rednaxelA'}}, {'prefix': {'fields.street': 'alexander'}}  
, {'prefix': {'reverse\_fields.street': 'rednaxela'}},  
{'prefix': {'fields.skype.profile': 'alexander'}},  
{'prefix': {'reverse\_fields.skype.profile': 'rednaxela'}},  
{'prefix': {'fields.skype.user\_name': 'alexander'}},  
{'prefix': {'reverse\_fields.skype.user\_name': 'rednaxela'}},  
{'prefix': {'fields.skype.user\_id': 'Alexander'}},  
{'prefix': {'reverse\_fields.skype.user\_id': 'rednaxelA'}}, {'prefix': {'fields.city': 'alexander'}},  
{'prefix': {'reverse\_fields.city': 'rednaxela'}}, {'prefix': {'fields.first\_name': 'alexander'}},  
{'prefix': {'reverse\_fields.first\_name': 'rednaxela'}}, {'prefix': {'fields.zip': 'alexander'}},  
{'prefix': {'reverse\_fields.zip': 'rednaxela'}}, {'prefix': {'fields.title': 'alexander'}},  
{'prefix': {'reverse\_fields.title': 'rednaxela'}}, {'prefix': {'fields.state': 'alexander'}},  
{'prefix': {'reverse\_fields.state': 'rednaxela'}}, {'prefix': {'fields.leadSource': 'alexander'}},  
{'prefix': {'reverse\_fields.leadSource': 'rednaxela'}}, {'prefix': {'fields.company\_name': 'alexander'}}  
, {'prefix': {'reverse\_fields.company\_name': 'rednaxela'}},  
{'prefix': {'fields.department': 'alexander'}}, {'prefix': {'reverse\_fields.department': 'rednaxela'}},  
{'prefix': {'fields.email.profile': 'alexander'}}, {'prefix':  
{  
'reverse\_fields.email.profile': 'rednaxela'}}  
, {'prefix': {'fields.email.user\_name': 'alexander'}},  
{'prefix': {'reverse\_fields.email.user\_name': 'rednaxela'}},  
{'prefix': {'fields.email.user\_id': 'Alexander'}},  
{'prefix': {'reverse\_fields.email.user\_id': 'rednaxelA'}},  
{'prefix': {'fields.website': 'alexander'}}, {'prefix': {'reverse\_fields.website': 'rednaxela'}}, {  
'prefix': {'fields.description': 'alexander'}}, {'prefix': {'reverse\_fields.description': 'rednaxela'}},  
{  
'prefix': {'fields.accountNumber': 'alexander'}},  
{'prefix': {'reverse\_fields.accountNumber': 'rednaxela'}}, {  
'prefix': {'fields.assistant': 'alexander'}}, {'prefix': {'reverse\_fields.assistant': 'rednaxela'}}, {  
'prefix': {'fields.phone': 'alexander'}}, {'prefix': {'reverse\_fields.phone': 'rednaxela'}}, {  
'prefix': {'fields.facebook.profile': 'alexander'}},  
{'prefix': {'reverse\_fields.facebook.profile': 'rednaxela'}}, {  
'prefix': {'fields.facebook.user\_name': 'alexander'}}, {  
'prefix': {'reverse\_fields.facebook.user\_name': 'rednaxela'}}, {  
'prefix': {'fields.facebook.user\_id': 'Alexander'}},  
{'prefix': {'reverse\_fields.facebook.user\_id': 'rednaxelA'}}, {  
'prefix': {'fields.leadType': 'alexander'}}, {'prefix': {'reverse\_fields.leadType': 'rednaxela'}}, {  
'prefix': {'fields.dates': 'alexander'}}, {'prefix': {'reverse\_fields.dates': 'rednaxela'}}, {  
'prefix': {'fields.name': 'alexander'}}, {'prefix': {'reverse\_fields.name': 'rednaxela'}}, {  
'prefix': {'fields.country': 'alexander'}}, {'prefix': {'reverse\_fields.country': 'rednaxela'}}, {  
'prefix': {'fields.assistantPhone': 'alexander'}}, {  
'prefix': {'reverse\_fields.assistantPhone': 'rednaxela'}}]}]}}, 'query': {'bool': {'must': [{  
'dis\_max': {'tie\_breaker': 0.7,  
'queries': [{'constant\_score': {'filter': {'term': {'fields.first\_name': 'alexander'}}, 'boost': 20.0}},  
{'constant\_score': {'filter': {'prefix': {'fields.first\_name': 'alexander'}}, 'boost': 11.0}}, {  
'constant\_score': {'filter': {'prefix': {'reverse\_fields.first\_name': 'rednaxela'}},  
'boost': 11.0}},  
{'constant\_score': {'filter': {'term': {'fields.last\_name': 'alexander'}}, 'boost': 40.0}},  
{'constant\_score': {'filter': {'prefix': {'fields.last\_name': 'alexander'}}, 'boost': 17.0}}, {  
'constant\_score': {'filter': {'prefix': {'reverse\_fields.last\_name': 'rednaxela'}},  
'boost': 17.0}},  
{'constant\_score': {'filter': {'term': {'fields.name': 'alexander'}}, 'boost': 18.0}},  
{'constant\_score': {'filter': {'prefix': {'fields.name': 'alexander'}}, 'boost': 5.0}},  
{'constant\_score': {'filter': {'prefix': {'reverse\_fields.name': 'rednaxela'}}, 'boost': 5.0}},  
{'match\_all': {}}]}}], 'should': []}}}}, 'explain': False}

Looking for any help 🙂

-- Michael Korbakov

---

<div class="post-metadata">

### Author: ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)
#### Post date: [June 21, 2011, 10:07am UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665/2 "2011-06-21T10:07:22Z")

</div>

Hi Michael

> ```
> {'prefix': {'fields.twitter.profile': 'alexander'}},
> {'prefix': {'reverse_fields.twitter.profile': 'rednaxela'}},
> {'prefix': {'fields.twitter.user_name': 'alexander'}},
> {'prefix': {'reverse_fields.twitter.user_name':
> 
> ```
> 
> 'rednaxela'}},

Prefix filters can be expensive - first they have to find all terms that  
begin with your prefix, then add separate clauses for each of those  
terms.

I think what you're looking for would be more efficiently achieved using  
edge ngrams

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

I've gisted an example of how you would create your index, and search  
for your data:

> <https://gist.github.com/clintongormley/1037563>

clint

---

<div class="post-metadata">

### Author: ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)
#### Post date: [June 21, 2011, 5:41pm UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665/3 "2011-06-21T17:41:37Z")

</div>

Thank you!

We're going to try it today. I'm concerned a little about general  
filter/query performance. I was under impression that any query will  
be slower then any filter. I guess it isn't the case here :). BTW, is  
it beneficial to wrap these query\_string into query filters?

-- Michael Korbakov

On Tue, Jun 21, 2011 at 3:07 AM, Clinton Gormley [via Elasticsearch  
Users] [ml-node+3090039-1179935648-83923@n3.nabble.com](mailto:ml-node+3090039-1179935648-83923@n3.nabble.com) wrote:

> Hi Michael
> 
> > ```
> > {'prefix': {'fields.twitter.profile': 'alexander'}},
> > {'prefix': {'reverse_fields.twitter.profile':
> > 
> > ```
> > 
> > 'rednaxela'}},  
> > {'prefix': {'fields.twitter.user\_name': 'alexander'}},  
> > {'prefix': {'reverse\_fields.twitter.user\_name':  
> > 'rednaxela'}},
> 
> Prefix filters can be expensive - first they have to find all terms that  
> begin with your prefix, then add separate clauses for each of those  
> terms.
> 
> I think what you're looking for would be more efficiently achieved using  
> edge ngrams
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/edgengram-tokenizer.html)
> 
> I've gisted an example of how you would create your index, and search  
> for your data:
> 
> [Edge ngram example for elasticsearch · GitHub](https://gist.github.com/1037563)
> 
> clint
> 
> * * *
> 
> If you reply to this email, your message will be added to the discussion  
> below:  
> [http://elasticsearch-users.115913.n3.nabble.com/Speed-of-query-with-many-filters-tp3088603p3090039.html](http://elasticsearch-users.115913.n3.nabble.com/Speed-of-query-with-many-filters-tp3088603p3090039.html)  
> To unsubscribe from Speed of query with many filters, click here.

---

<div class="post-metadata">

### Author: ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)
#### Post date: [June 22, 2011, 11:32am UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665/4 "2011-06-22T11:32:38Z")

</div>

Hiya

> We're going to try it today. I'm concerned a little about general  
> filter/query performance. I was under impression that any query will  
> be slower then any filter. I guess it isn't the case here :).

Well, queries have a second phase that filters don't have: calculating  
the score/relevance, so yes, they generally don't perform as well.  
Also, filters can be cached, while queries can't.

> BTW, is  
> it beneficial to wrap these query\_string into query filters?

Not sure what you mean here.

One thing I should have thought of yesterday was that this might not do  
exactly what you want. In your example, you were searching for  
'alexander' in its entirety.

Because I set the 'analyzer' for those fields to use edge-ngrams, it  
uses them both at index time and at search time.

So a search for 'alexander' against the twitter profile field actually  
becomes a search for a|al|ale|alex|etc

Two options here:

1. you can set the index\_analyzer to "left"|"right" and the  
search\_analyzer to "default"

2. you can use a term filter eg:  
{ and: [  
{ term: { "twitter.profile": "alexander" }},  
{ term: { "twitter.profile.reverse\_profile": "alexander" }}  
]}

Option (2) will be faster, but option(1) has the advantage that it  
handles the analysis of the search term for you.

For instance, "foo-bar" would actually be broken down into "foo" "bar",  
but if you do a term query for "foo-bar" it won't be found, it doesn't  
exist.

So if you are sure that, in your app, you are converting the query text  
(alexander) into the correct terms that are stored in ES, then method 2  
would be preferred. If not, then it may be better to rely on the query  
instead.

clint

---

<div class="post-metadata">

### Author: ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)
#### Post date: [June 23, 2011, 5:16am UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665/5 "2011-06-23T05:16:44Z")

</div>

\> BTW, is \> it beneficial to wrap these query\_string into query filters? Not sure what you mean here. I was meaning this filter: http://www.elasticsearch.org/guide/reference/query-dsl/query-filter.html One thing I should have thought of yesterday was that this might not do exactly what you want. In your example, you were searching for 'alexander' in its entirety.

Because I set the 'analyzer' for those fields to use edge-ngrams, it  
uses them both at index time and at search time.

So a search for 'alexander' against the twitter profile field actually  
becomes a search for a|al|ale|alex|etc

Two options here:

1. you can set the index\_analyzer to "left"|"right" and the  
search\_analyzer to "default"

2. you can use a term filter eg:  
{ and: [  
{ term: { "twitter.profile": "alexander" }},  
{ term: { "twitter.profile.reverse\_profile": "alexander" }}  
]}

Option (2) will be faster, but option(1) has the advantage that it  
handles the analysis of the search term for you.  
We're stopped on option (2). However this prefix substitution doesn't shown any significant speed improvement. We're still getting timeouts. Now we trying to combine all fields we're searching by into single one (pseudo \_all) and run search by it. Hope that will help.

---

<div class="post-metadata">

### Author: ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)
#### Post date: [June 23, 2011, 8:55am UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665/6 "2011-06-23T08:55:42Z")

</div>

On Wed, 2011-06-22 at 22:16 -0700, Michael Korbakov wrote:

> Clinton Gormley wrote:
> 
> > > BTW, is  
> > > it beneficial to wrap these query\_string into query filters?  
> > > Not sure what you mean here.  
> > > I was meaning this filter:  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/query-dsl/query-filter.html)

I don't know, to be honest. Not sure if wrapping a query in a  
query-filter disables the \_score calculation phase or not.

> > 1. you can use a term filter eg:  
> > { and: [  
> > { term: { "twitter.profile": "alexander" }},  
> > { term: { "twitter.profile.reverse\_profile": "alexander" }}  
> > ]}
> > 
> > Option (2) will be faster, but option(1) has the advantage that it  
> > handles the analysis of the search term for you.

We're stopped on option (2). However this prefix substitution doesn't  
shown any significant speed improvement. We're still getting timeouts. Now  
we trying to combine all fields we're searching by into single one (pseudo  
\_all) and run search by it. Hope that will help.

By "stopped" do you mean you're using option 2, or you have decided  
against using option 2?

Filters should be fast, even if there are many of them. However, you  
need to have enough memory to hold all of the terms, and the initial  
query will be slow as it needs to load all of those terms the first  
time.

After the first run, it should be significantly faster.

clint

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 4:02am UTC](https://discuss.elastic.co/t/speed-of-query-with-many-filters/4665/7 "2017-07-06T04:02:54Z")

</div>


