# Has\_child / has\_parent for billions of children - heavy cpu load for simple queries?

**URL:** <https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026>\
**Category:** Elasticsearch\
**Created:** [August 26, 2016, 10:20am UTC](https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026 "2016-08-26T10:20:22Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![AndreCi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andreci/32/11653_2.png) [@AndreCi](https://discuss.elastic.co/u/AndreCi)\
**Post date:** [August 26, 2016, 10:20am UTC](https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026/1 "2016-08-26T10:20:22Z")

</div>

Hey everyone,

is there a trick to speed up has\_child / has\_parent joins for small subsets of large document bodies? We have around 90 million parent documents and close to 1.4 billion child documents, seven 10-core machines with 64GB ram and three emergency-scrambled fast-clocked 4 core machines with 64GB ram.

We already had huge memory problems with random\_sort via function query, eating \>60GB of heap per query, which forced us to pre-calculate a random\_order to avoid that. Now we found the next pitfall: using has\_child/has\_parent for small subsets of children/parent documents _guzzles_ cpu time... and the response times grow from 9-15ms to 400-1500ms as soon as we touch parents/childs (depending on load).

We already use eager global ordinals, don't score the childs/parents and now I'm left wondering what we could do, except, of course, throwing more servers at the problem. Maybe like-data saving is a extremely bad use case for Elasticsearch? (I don't want to venture back to sql-land... 😅 )

Any feedback highly welcome and thanks in advance! 🙂

Oh, and here's one of the offending queries:

```
/instagram-user/instagram-like/_search
{
   "query": {
      "bool": {
     "filter": [
        {
           "term": {
              "instagram_post_id": 1312932664928280000
           }
        },
        {
           "has_parent": {
              "parent_type": "instagram-user",
              "score_mode": "none",
              "query": {
                 "bool": {
                    "must_not": [
                       {
                          "exists": {
                             "field": "calculated"
                          }
                       }
                    ]
                 }
              }
           }
        }
     ]
  }
   },
   "size": 0
}
```

---

<div class="post-metadata">

**Author:** ![abeyad](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abeyad](https://discuss.elastic.co/u/abeyad)\
**Post date:** [August 26, 2016, 3:43pm UTC](https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026/2 "2016-08-26T15:43:58Z")

</div>

Parent/child queries are slow by design. It runs a first phase that consists in identifying the matching join values, and then a second phase that consists in matching the documents that have these join values. However this second phase usually runs a linear scan in order to find matches, which is slow as it runs in linear time with the number of documents in the joined type.

You may try to avoid the parent/child paradigm by inserting as much data as possible from the child documents into the parent ones.

One thing you may try which may or may not help is to increase the number of shards (primary) in your index. That way, each shard would have a smaller set of join values to iterate through.

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [August 29, 2016, 7:40am UTC](https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026/3 "2016-08-29T07:40:40Z")

</div>

> [@abeyad](#):
>
> However this second phase usually runs a linear scan in order to find matches, which is slow as it runs in linear time with the number of documents in the joined type.

This was true for ES \< 2.0. New indices created on 2.0 and onwards the second phase is linear to the amount of parent docs that match with the rest of the query. In your case that are documents that match with the term query on the instagram\_post\_id field. How many documents do match this term query? If that number is small then I do expect reasonable performance.

---

<div class="post-metadata">

**Author:** ![AndreCi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andreci/32/11653_2.png) [@AndreCi](https://discuss.elastic.co/u/AndreCi)\
**Post date:** [August 29, 2016, 8:26am UTC](https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026/4 "2016-08-29T08:26:21Z")

</div>

The documents matching count depends on the size of the users and the age of the post, ranging from the hundreds up to a few hundred thousand.

We figured that, in our case, we shouldn't join children into parents, since this would mean joining all child documents for uncalculated users (~50 million parents) before applying the child-query filters - or does ES apply has\_child-query filters before joining the children?

If that's not the case we could try to increase the number of primary shards or start thinking about denormalizing the bool flag into the children (resulting in updating millions of documents after doing a simple calculation on the parent) or moving an aggregated liked-users into the parents (resulting in steady stream of reindexes)... decisions, decisions, yay 🙂

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [August 29, 2016, 8:55am UTC](https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026/5 "2016-08-29T08:55:34Z")

</div>

> [@AndreCi](#):
>
> We figured that, in our case, we shouldn't join children into parents, since this would mean joining all child documents for uncalculated users (~50 million parents) before applying the child-query filters - or does ES apply has\_child-query filters before joining the children?

In the first phase, the filters / queries inside the has\_child/has\_parent query are basically used to select what child document match. So these filters are applied, it is just that this isn't the most expensive part of the has\_child/has\_parent query. It is the second phase that connects the child hits back to their parent hits. This is linear, all parents need to be checked if one of its child documents has been matched in the first phase (or the other way around with has\_parent). However if the has\_child/has\_parent query are part of bigger query than not all parent queries need to be checked for whether they have child matches. In the query you shared here only child documents that match with the term query will be checked if their associated parent document were matched during the first phase of the has\_parent query. For this to work this way you should be running this query on an index created on or after ES 2.0, otherwise all child document will be checked for whether they matched. Are you on a ES 2.x version? Did you upgrade from 1.x? In that case you need to reindex the indices created before the upgrade in order for p/c to execute this way.

> [@AndreCi](#):
>
> If that's not the case we could try to increase the number of primary shards or start thinking about denormalizing the bool flag into the children (resulting in updating millions of documents after doing a simple calculation on the parent) or moving an aggregated liked-users into the parents (resulting in steady stream of reindexes)... decisions, decisions, yay 🙂

If you want super fast searches then parent/child design isn't meant to be used here. De-normalizing the data is the only way. So, yes, decisions 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:24pm UTC](https://discuss.elastic.co/t/has-child-has-parent-for-billions-of-children-heavy-cpu-load-for-simple-queries/59026/6 "2017-07-05T22:24:29Z")

</div>


