# Is there a way to enforce the execution order of filters?

**URL:** <https://discuss.elastic.co/t/is-there-a-way-to-enforce-the-execution-order-of-filters/81856>\
**Category:** Elasticsearch\
**Created:** [April 10, 2017, 4:49pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-enforce-the-execution-order-of-filters/81856 "2017-04-10T16:49:34Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![AndreCimander](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrecimander/32/30386_2.png) [@AndreCimander](https://discuss.elastic.co/u/AndreCimander)\
**Post date:** [April 10, 2017, 4:49pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-enforce-the-execution-order-of-filters/81856/1 "2017-04-10T16:49:34Z")

</div>

I finally got around to profile some slow running parts of our application and I found a query that is taking way too long and which also creates a good chunk of our cluster load. Profiling the query in Kibana revealed that most shards finish well below 10ms, but there are almost always a few random shards that take 10 seconds + (most of the time spent in the GlobalOrdinalsQuery.build\_scorer).

I suspect that the has\_parent query is run before the way more restrictive filters effectively joining billion of documents. Is there a way to enforce the filter order, e.g. rewrite the query to make sure the term filters are executed before the parent-child join madness?

The index refreshes every 120 seconds.

The query:

```
GET /user-v5/like/_search
{
   "query": {
      "bool": {
         "filter": [
            {
               "terms": {
                  "post_id": [
                     1489831183823275924,
                     1489206393580157727
                  ]
               }
            },
            {
               "term": {
                  "user_id": 587771206
               }
            },
            {
               "has_parent": {
                  "query": {
                     "bool": {
                        "filter": [
                           {
                              "term": {
                                 "calculated": true
                              }
                           }
                        ]
                     }
                  },
                  "score_mode": "none",
                  "parent_type": "user"
               }
            }
         ]
      }
   },
   "from": 0,
   "aggs": {
      "per_post": {
         "terms": {
            "field": "post_id",
            "size": 5
         }
      }
   },
   "size": 0
}
```

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [April 10, 2017, 5:46pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-enforce-the-execution-order-of-filters/81856/2 "2017-04-10T17:46:47Z")

</div>

`has_parent` needs the global ordinals to run at all, iirc. So it doesn't matter how selective the term filters are unless they eliminate all the documents.

---

<div class="post-metadata">

**Author:** ![AndreCimander](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrecimander/32/30386_2.png) [@AndreCimander](https://discuss.elastic.co/u/AndreCimander)\
**Post date:** [April 11, 2017, 8:30am UTC](https://discuss.elastic.co/t/is-there-a-way-to-enforce-the-execution-order-of-filters/81856/3 "2017-04-11T08:30:41Z")

</div>

Hm, so, the has\_parent query is always joining all children with all parents, regardless of reducing the children beforehand from a few billion to just 85 thousand documents?

The \_parent global ordinals are build eagerly on each refresh, so there should be no rebuild at query time, besides why are just a few random shards 400 up to 3000 times slower?

99% of the time is spent in GlobalOrdinalsQuery.build\_scorer, is this just a misleading naming, or is it really scoring there? The query is used in a filter context and with score\_mode set to none, so there shouldn't be any scoring happening, should there?

We have just 4 threads each sending around 0,5 of those queries per second to our cluster and they consume one third of all CPU resources available ...

---

<div class="post-metadata">

**Author:** ![AndreCimander](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrecimander/32/30386_2.png) [@AndreCimander](https://discuss.elastic.co/u/AndreCimander)\
**Post date:** [April 11, 2017, 10:26am UTC](https://discuss.elastic.co/t/is-there-a-way-to-enforce-the-execution-order-of-filters/81856/4 "2017-04-11T10:26:53Z")

</div>

Okay, I found a workaround that at least mitigates the slow queries, although the load on the cluster has increased thanks to the increased throughput 😰

I added a has\_child query with all bool conditions from the outer query, so in case a shard decides to use the has\_parent query first, the has\_child query inside the has\_parent query seems to limit the join carnage.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 9, 2017, 10:36am UTC](https://discuss.elastic.co/t/is-there-a-way-to-enforce-the-execution-order-of-filters/81856/5 "2017-05-09T10:36:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
