# Performance penalty for has\_child queries

**URL:** <https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794>\
**Category:** Elasticsearch\
**Created:** [February 23, 2012, 6:59pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794 "2012-02-23T18:59:44Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tim\_J](https://avatars.discourse-cdn.com/v4/letter/t/f17d59/32.png) [@Tim\_J](https://discuss.elastic.co/u/Tim_J)\
**Post date:** [February 23, 2012, 6:59pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/1 "2012-02-23T18:59:44Z")

</div>

Hey folks,  
I'm currently working with an ES index of roughly 52 million  
documents. We index approximately 10-20 new docs per second. Each  
document is broken into two pieces and indexed as a parent/child  
pair. The child contains static content and is unlikely to ever be  
updated. The parent fields are modified frequently which is why the  
child content was separated, particularly as the original source for  
the child documents is expensive to retrieve.

Documents are replicated across three nodes. No data is stored with  
the exception of a unique id for each doc. Each node is allocated 8  
GB of RAM and we occupy about 22 GB per node on disk. We use the  
routing key to "shard" our data. There are approximately 130  
different routing keys in use at the moment. Routing keys are also  
used as conditions for all searches so they should be a quick filter.

First, does anyone have a sense of the penalty we're paying for  
having this parent/child relationship? We're seeing some very long  
query times particularly when we're actively writing to the nodes.  
Sometimes a simple query with one condition on the parent and one in a  
has\_child can take 8+ minutes.

I've noticed that when we're doing a lot of writes to the child  
index in particular the times go up significantly. On the other hand  
if we only write to the parent index this is much less of a problem.  
Is this expected?

Finally, does anyone have any suggestions for tuning this  
configuration or improving our queries?

Thanks,  
-Tim

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 26, 2012, 7:11pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/2 "2012-02-26T19:11:50Z")

</div>

The main penalty that occurs when using parent/child mapping happens because the ids need to be loaded to memory in order to do an efficient join process between the child and the parent. This initial loading of the ids can be expensive, but subsequent requests will be fast, even when indexing data.

On Thursday, February 23, 2012 at 8:59 PM, Tim J wrote:

> Hey folks,  
> I'm currently working with an ES index of roughly 52 million  
> documents. We index approximately 10-20 new docs per second. Each  
> document is broken into two pieces and indexed as a parent/child  
> pair. The child contains static content and is unlikely to ever be  
> updated. The parent fields are modified frequently which is why the  
> child content was separated, particularly as the original source for  
> the child documents is expensive to retrieve.
> 
> Documents are replicated across three nodes. No data is stored with  
> the exception of a unique id for each doc. Each node is allocated 8  
> GB of RAM and we occupy about 22 GB per node on disk. We use the  
> routing key to "shard" our data. There are approximately 130  
> different routing keys in use at the moment. Routing keys are also  
> used as conditions for all searches so they should be a quick filter.
> 
> First, does anyone have a sense of the penalty we're paying for  
> having this parent/child relationship? We're seeing some very long  
> query times particularly when we're actively writing to the nodes.  
> Sometimes a simple query with one condition on the parent and one in a  
> has\_child can take 8+ minutes.
> 
> I've noticed that when we're doing a lot of writes to the child  
> index in particular the times go up significantly. On the other hand  
> if we only write to the parent index this is much less of a problem.  
> Is this expected?
> 
> Finally, does anyone have any suggestions for tuning this  
> configuration or improving our queries?
> 
> Thanks,  
> -Tim

---

<div class="post-metadata">

**Author:** ![Tim\_J](https://avatars.discourse-cdn.com/v4/letter/t/f17d59/32.png) [@Tim\_J](https://discuss.elastic.co/u/Tim_J)\
**Post date:** [February 27, 2012, 3:20pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/3 "2012-02-27T15:20:30Z")

</div>

Thanks Shay. So it sounds like there's a cache to manage the joins.  
Is there anything that would cause the cache of ids to be cleared?  
How is it managed? I'm just wondering if we're doing something silly  
that would cause it to get nuked periodically.

Thanks,  
-Tim

On Feb 26, 2:11 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> The main penalty that occurs when using parent/child mapping happens because the ids need to be loaded to memory in order to do an efficient join process between the child and the parent. This initial loading of the ids can be expensive, but subsequent requests will be fast, even when indexing data.
> 
> On Thursday, February 23, 2012 at 8:59 PM, Tim J wrote:
> 
> > Hey folks,  
> > I'm currently working with an ES index of roughly 52 million  
> > documents. We index approximately 10-20 new docs per second. Each  
> > document is broken into two pieces and indexed as a parent/child  
> > pair. The child contains static content and is unlikely to ever be  
> > updated. The parent fields are modified frequently which is why the  
> > child content was separated, particularly as the original source for  
> > the child documents is expensive to retrieve.
> 
> > Documents are replicated across three nodes. No data is stored with  
> > the exception of a unique id for each doc. Each node is allocated 8  
> > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > routing key to "shard" our data. There are approximately 130  
> > different routing keys in use at the moment. Routing keys are also  
> > used as conditions for all searches so they should be a quick filter.
> 
> > First, does anyone have a sense of the penalty we're paying for  
> > having this parent/child relationship? We're seeing some very long  
> > query times particularly when we're actively writing to the nodes.  
> > Sometimes a simple query with one condition on the parent and one in a  
> > has\_child can take 8+ minutes.
> 
> > I've noticed that when we're doing a lot of writes to the child  
> > index in particular the times go up significantly. On the other hand  
> > if we only write to the parent index this is much less of a problem.  
> > Is this expected?
> 
> > Finally, does anyone have any suggestions for tuning this  
> > configuration or improving our queries?
> 
> > Thanks,  
> > -Tim

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 27, 2012, 5:46pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/4 "2012-02-27T17:46:55Z")

</div>

On Sunday, 26 February 2012 14:11:50 UTC-5, kimchy wrote:

> The main penalty that occurs when using parent/child mapping happens  
> because the ids need to be loaded to memory in order to do an efficient  
> join process between the child and the parent. This initial loading of the  
> ids can be expensive, but subsequent requests will be fast, even when  
> indexing data.

When you said that the IDs need to be loaded into memory, am I correct in  
interpreting that as all of the parents' and childrens' IDs?

For example, if there are 10K parents and 20K children, 30K IDs will be  
loaded into memory to perform the join?

Thanks, Shay.

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 27, 2012, 6:34pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/5 "2012-02-27T18:34:48Z")

</div>

I'm in the same boat as Tim:

- Product documents are expensive to generate, and change infrequently.
- ProductHave documents (which track which users want each product) are  
cheap to generate, and change frequently.

Thus, I was planning on using a parent-child relationship, where the  
"product\_have" type is a child of the "product" type. This would make it  
cheap and trivial to update which users have a product.

However, if performance will go down the drain when there are millions of  
documents, there must be a better solution. Any ideas?

---

<div class="post-metadata">

**Author:** ![Radu\_Gheorghe1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe1/32/2688_2.png) [@Radu\_Gheorghe1](https://discuss.elastic.co/u/Radu_Gheorghe1)\
**Post date:** [February 28, 2012, 8:49am UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/6 "2012-02-28T08:49:35Z")

</div>

Hi Nick,

I'm having a similar problem, so I'll just share my thoughts here,  
without expecting them to be "solutions".

When you have documents that are related to each other, there are  
three options:

- use a "relational" database. But that would hurt scalability, of  
course
- manage relations yourself. I'm thinking about holding different  
types of documents in different indices, and then build the  
"relations" within your application's logic
- structure your data according to how your queries will look like. I  
got this idea from a book on Cassandra :D. For example, it might make  
sense to hold all the data in one document (non-nested), and update it  
with new data (eg: new fields for new customers).

On Feb 27, 8:34 pm, Nick Hoffman [n...@deadorange.com](mailto:n...@deadorange.com) wrote:

> I'm in the same boat as Tim:
> 
> - Product documents are expensive to generate, and change infrequently.
> - ProductHave documents (which track which users want each product) are  
> cheap to generate, and change frequently.
> 
> Thus, I was planning on using a parent-child relationship, where the  
> "product\_have" type is a child of the "product" type. This would make it  
> cheap and trivial to update which users have a product.
> 
> However, if performance will go down the drain when there are millions of  
> documents, there must be a better solution. Any ideas?

---

<div class="post-metadata">

**Author:** ![haarts](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/haarts/32/2972_2.png) [@haarts](https://discuss.elastic.co/u/haarts)\
**Post date:** [February 28, 2012, 11:28am UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/7 "2012-02-28T11:28:17Z")

</div>

Let me chime in as well.  
Our problem is similar. We have about 700M items and add items at a speed  
of 100/s. The performance we are seeing is not great. The query time  
required (has\_child) is dependent on the amount on new items indexed (that  
makes sense). But that time is already seconds(!) after adding a couple of  
thousand new items. And many minutes if we leave it running for a while.

We are currently investigating whether it is possible to add the ids to the  
internal memory map as soon as they are indexed.

On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:

> Hey folks,  
> I'm currently working with an ES index of roughly 52 million  
> documents. We index approximately 10-20 new docs per second. Each  
> document is broken into two pieces and indexed as a parent/child  
> pair. The child contains static content and is unlikely to ever be  
> updated. The parent fields are modified frequently which is why the  
> child content was separated, particularly as the original source for  
> the child documents is expensive to retrieve.
> 
> Documents are replicated across three nodes. No data is stored with  
> the exception of a unique id for each doc. Each node is allocated 8  
> GB of RAM and we occupy about 22 GB per node on disk. We use the  
> routing key to "shard" our data. There are approximately 130  
> different routing keys in use at the moment. Routing keys are also  
> used as conditions for all searches so they should be a quick filter.
> 
> First, does anyone have a sense of the penalty we're paying for  
> having this parent/child relationship? We're seeing some very long  
> query times particularly when we're actively writing to the nodes.  
> Sometimes a simple query with one condition on the parent and one in a  
> has\_child can take 8+ minutes.
> 
> I've noticed that when we're doing a lot of writes to the child  
> index in particular the times go up significantly. On the other hand  
> if we only write to the parent index this is much less of a problem.  
> Is this expected?
> 
> Finally, does anyone have any suggestions for tuning this  
> configuration or improving our queries?
> 
> Thanks,  
> -Tim

On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:

> Hey folks,  
> I'm currently working with an ES index of roughly 52 million  
> documents. We index approximately 10-20 new docs per second. Each  
> document is broken into two pieces and indexed as a parent/child  
> pair. The child contains static content and is unlikely to ever be  
> updated. The parent fields are modified frequently which is why the  
> child content was separated, particularly as the original source for  
> the child documents is expensive to retrieve.
> 
> Documents are replicated across three nodes. No data is stored with  
> the exception of a unique id for each doc. Each node is allocated 8  
> GB of RAM and we occupy about 22 GB per node on disk. We use the  
> routing key to "shard" our data. There are approximately 130  
> different routing keys in use at the moment. Routing keys are also  
> used as conditions for all searches so they should be a quick filter.
> 
> First, does anyone have a sense of the penalty we're paying for  
> having this parent/child relationship? We're seeing some very long  
> query times particularly when we're actively writing to the nodes.  
> Sometimes a simple query with one condition on the parent and one in a  
> has\_child can take 8+ minutes.
> 
> I've noticed that when we're doing a lot of writes to the child  
> index in particular the times go up significantly. On the other hand  
> if we only write to the parent index this is much less of a problem.  
> Is this expected?
> 
> Finally, does anyone have any suggestions for tuning this  
> configuration or improving our queries?
> 
> Thanks,  
> -Tim

---

<div class="post-metadata">

**Author:** ![Tim\_J](https://avatars.discourse-cdn.com/v4/letter/t/f17d59/32.png) [@Tim\_J](https://discuss.elastic.co/u/Tim_J)\
**Post date:** [February 28, 2012, 1:19pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/8 "2012-02-28T13:19:40Z")

</div>

Hey haarts,  
I'd be really interested to see what you come up with here. So from  
the sound of it, until a document turns up in a query it's not added  
to the memory map? Can you point me to the area of code where that  
memory map is managed?

Thanks,  
-Tim

On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:

> Let me chime in as well.  
> Our problem is similar. We have about 700M items and add items at a speed  
> of 100/s. The performance we are seeing is not great. The query time  
> required (has\_child) is dependent on the amount on new items indexed (that  
> makes sense). But that time is already seconds(!) after adding a couple of  
> thousand new items. And many minutes if we leave it running for a while.
> 
> We are currently investigating whether it is possible to add the ids to the  
> internal memory map as soon as they are indexed.
> 
> On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> 
> > Hey folks,  
> > I'm currently working with an ES index of roughly 52 million  
> > documents. We index approximately 10-20 new docs per second. Each  
> > document is broken into two pieces and indexed as a parent/child  
> > pair. The child contains static content and is unlikely to ever be  
> > updated. The parent fields are modified frequently which is why the  
> > child content was separated, particularly as the original source for  
> > the child documents is expensive to retrieve.
> 
> > Documents are replicated across three nodes. No data is stored with  
> > the exception of a unique id for each doc. Each node is allocated 8  
> > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > routing key to "shard" our data. There are approximately 130  
> > different routing keys in use at the moment. Routing keys are also  
> > used as conditions for all searches so they should be a quick filter.
> 
> > First, does anyone have a sense of the penalty we're paying for  
> > having this parent/child relationship? We're seeing some very long  
> > query times particularly when we're actively writing to the nodes.  
> > Sometimes a simple query with one condition on the parent and one in a  
> > has\_child can take 8+ minutes.
> 
> > I've noticed that when we're doing a lot of writes to the child  
> > index in particular the times go up significantly. On the other hand  
> > if we only write to the parent index this is much less of a problem.  
> > Is this expected?
> 
> > Finally, does anyone have any suggestions for tuning this  
> > configuration or improving our queries?
> 
> > Thanks,  
> > -Tim  
> > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> 
> > Hey folks,  
> > I'm currently working with an ES index of roughly 52 million  
> > documents. We index approximately 10-20 new docs per second. Each  
> > document is broken into two pieces and indexed as a parent/child  
> > pair. The child contains static content and is unlikely to ever be  
> > updated. The parent fields are modified frequently which is why the  
> > child content was separated, particularly as the original source for  
> > the child documents is expensive to retrieve.
> 
> > Documents are replicated across three nodes. No data is stored with  
> > the exception of a unique id for each doc. Each node is allocated 8  
> > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > routing key to "shard" our data. There are approximately 130  
> > different routing keys in use at the moment. Routing keys are also  
> > used as conditions for all searches so they should be a quick filter.
> 
> > First, does anyone have a sense of the penalty we're paying for  
> > having this parent/child relationship? We're seeing some very long  
> > query times particularly when we're actively writing to the nodes.  
> > Sometimes a simple query with one condition on the parent and one in a  
> > has\_child can take 8+ minutes.
> 
> > I've noticed that when we're doing a lot of writes to the child  
> > index in particular the times go up significantly. On the other hand  
> > if we only write to the parent index this is much less of a problem.  
> > Is this expected?
> 
> > Finally, does anyone have any suggestions for tuning this  
> > configuration or improving our queries?
> 
> > Thanks,  
> > -Tim

---

<div class="post-metadata">

**Author:** ![haarts](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/haarts/32/2972_2.png) [@haarts](https://discuss.elastic.co/u/haarts)\
**Post date:** [February 28, 2012, 1:39pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/9 "2012-02-28T13:39:37Z")

</div>

I'm currently looking for that area. 🙂 I'll keep you posted.  
In the meanwhile I'm also testing a hack.  
while true  
perform has\_child query  
sleep 1  
end

Seems to work surprisingly well. But we are not going to leave that in  
place of course.

On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:

> Hey haarts,  
> I'd be really interested to see what you come up with here. So from  
> the sound of it, until a document turns up in a query it's not added  
> to the memory map? Can you point me to the area of code where that  
> memory map is managed?
> 
> Thanks,  
> -Tim
> 
> On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> 
> > Let me chime in as well.  
> > Our problem is similar. We have about 700M items and add items at a  
> > speed  
> > of 100/s. The performance we are seeing is not great. The query time  
> > required (has\_child) is dependent on the amount on new items indexed  
> > (that  
> > makes sense). But that time is already seconds(!) after adding a couple  
> > of  
> > thousand new items. And many minutes if we leave it running for a while.
> > 
> > We are currently investigating whether it is possible to add the ids to  
> > the  
> > internal memory map as soon as they are indexed.
> > 
> > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim  
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim

On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:

> Hey haarts,  
> I'd be really interested to see what you come up with here. So from  
> the sound of it, until a document turns up in a query it's not added  
> to the memory map? Can you point me to the area of code where that  
> memory map is managed?
> 
> Thanks,  
> -Tim
> 
> On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> 
> > Let me chime in as well.  
> > Our problem is similar. We have about 700M items and add items at a  
> > speed  
> > of 100/s. The performance we are seeing is not great. The query time  
> > required (has\_child) is dependent on the amount on new items indexed  
> > (that  
> > makes sense). But that time is already seconds(!) after adding a couple  
> > of  
> > thousand new items. And many minutes if we leave it running for a while.
> > 
> > We are currently investigating whether it is possible to add the ids to  
> > the  
> > internal memory map as soon as they are indexed.
> > 
> > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim  
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim

---

<div class="post-metadata">

**Author:** ![haarts](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/haarts/32/2972_2.png) [@haarts](https://discuss.elastic.co/u/haarts)\
**Post date:** [February 28, 2012, 3:59pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/10 "2012-02-28T15:59:46Z")

</div>

Hi Tim,

Just an update.  
We are encountering some unexpected behaviour when running the same  
has\_child query twice consecutively. The first query takes about 30 minutes  
to complete, as does the second one. This in contrary to the believe that  
once the IDs are loaded in memory search should be fast. The index contains  
36M documents and the index size is 30GB running on an 8 core i7 with 24GB  
RAM.

Regarding your question on when an ID is loaded to the memory map; I  
believe they are all loaded all the time.  
I believe the code responsible is in  
java/org/elasticsearch/index/cache/id/simple/SimpleIdCache.java[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/cache/id/simple/SimpleIdCache.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/cache/id/simple/SimpleIdCache.java)  
.

Harm

On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:

> Hey haarts,  
> I'd be really interested to see what you come up with here. So from  
> the sound of it, until a document turns up in a query it's not added  
> to the memory map? Can you point me to the area of code where that  
> memory map is managed?
> 
> Thanks,  
> -Tim
> 
> On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> 
> > Let me chime in as well.  
> > Our problem is similar. We have about 700M items and add items at a  
> > speed  
> > of 100/s. The performance we are seeing is not great. The query time  
> > required (has\_child) is dependent on the amount on new items indexed  
> > (that  
> > makes sense). But that time is already seconds(!) after adding a couple  
> > of  
> > thousand new items. And many minutes if we leave it running for a while.
> > 
> > We are currently investigating whether it is possible to add the ids to  
> > the  
> > internal memory map as soon as they are indexed.
> > 
> > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim  
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim

On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:

> Hey haarts,  
> I'd be really interested to see what you come up with here. So from  
> the sound of it, until a document turns up in a query it's not added  
> to the memory map? Can you point me to the area of code where that  
> memory map is managed?
> 
> Thanks,  
> -Tim
> 
> On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> 
> > Let me chime in as well.  
> > Our problem is similar. We have about 700M items and add items at a  
> > speed  
> > of 100/s. The performance we are seeing is not great. The query time  
> > required (has\_child) is dependent on the amount on new items indexed  
> > (that  
> > makes sense). But that time is already seconds(!) after adding a couple  
> > of  
> > thousand new items. And many minutes if we leave it running for a while.
> > 
> > We are currently investigating whether it is possible to add the ids to  
> > the  
> > internal memory map as soon as they are indexed.
> > 
> > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim  
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim

On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:

> Hey haarts,  
> I'd be really interested to see what you come up with here. So from  
> the sound of it, until a document turns up in a query it's not added  
> to the memory map? Can you point me to the area of code where that  
> memory map is managed?
> 
> Thanks,  
> -Tim
> 
> On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> 
> > Let me chime in as well.  
> > Our problem is similar. We have about 700M items and add items at a  
> > speed  
> > of 100/s. The performance we are seeing is not great. The query time  
> > required (has\_child) is dependent on the amount on new items indexed  
> > (that  
> > makes sense). But that time is already seconds(!) after adding a couple  
> > of  
> > thousand new items. And many minutes if we leave it running for a while.
> > 
> > We are currently investigating whether it is possible to add the ids to  
> > the  
> > internal memory map as soon as they are indexed.
> > 
> > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim  
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > 
> > > Hey folks,  
> > > I'm currently working with an ES index of roughly 52 million  
> > > documents. We index approximately 10-20 new docs per second. Each  
> > > document is broken into two pieces and indexed as a parent/child  
> > > pair. The child contains static content and is unlikely to ever be  
> > > updated. The parent fields are modified frequently which is why the  
> > > child content was separated, particularly as the original source for  
> > > the child documents is expensive to retrieve.
> > 
> > > Documents are replicated across three nodes. No data is stored with  
> > > the exception of a unique id for each doc. Each node is allocated 8  
> > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > routing key to "shard" our data. There are approximately 130  
> > > different routing keys in use at the moment. Routing keys are also  
> > > used as conditions for all searches so they should be a quick filter.
> > 
> > > First, does anyone have a sense of the penalty we're paying for  
> > > having this parent/child relationship? We're seeing some very long  
> > > query times particularly when we're actively writing to the nodes.  
> > > Sometimes a simple query with one condition on the parent and one in a  
> > > has\_child can take 8+ minutes.
> > 
> > > I've noticed that when we're doing a lot of writes to the child  
> > > index in particular the times go up significantly. On the other hand  
> > > if we only write to the parent index this is much less of a problem.  
> > > Is this expected?
> > 
> > > Finally, does anyone have any suggestions for tuning this  
> > > configuration or improving our queries?
> > 
> > > Thanks,  
> > > -Tim

---

<div class="post-metadata">

**Author:** ![Tim\_J](https://avatars.discourse-cdn.com/v4/letter/t/f17d59/32.png) [@Tim\_J](https://discuss.elastic.co/u/Tim_J)\
**Post date:** [February 29, 2012, 2:11pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/11 "2012-02-29T14:11:00Z")

</div>

Thanks Harm. I think we may be looking at alternative solutions for  
this one as our deadline is fast approaching. I'll be keeping an eye  
on this thread though in case you turn up anything good!

-Tim

On Feb 28, 10:59 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:

> Hi Tim,
> 
> Just an update.  
> We are encountering some unexpected behaviour when running the same  
> has\_child query twice consecutively. The first query takes about 30 minutes  
> to complete, as does the second one. This in contrary to the believe that  
> once the IDs are loaded in memory search should be fast. The index contains  
> 36M documents and the index size is 30GB running on an 8 core i7 with 24GB  
> RAM.
> 
> Regarding your question on when an ID is loaded to the memory map; I  
> believe they are all loaded all the time.  
> I believe the code responsible is in  
> java/org/elasticsearch/index/cache/id/simple/SimpleIdCache.java[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/j...](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/j...)  
> .
> 
> Harm
> 
> On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:
> 
> > Hey haarts,  
> > I'd be really interested to see what you come up with here. So from  
> > the sound of it, until a document turns up in a query it's not added  
> > to the memory map? Can you point me to the area of code where that  
> > memory map is managed?
> 
> > Thanks,  
> > -Tim
> 
> > On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> > 
> > > Let me chime in as well.  
> > > Our problem is similar. We have about 700M items and add items at a  
> > > speed  
> > > of 100/s. The performance we are seeing is not great. The query time  
> > > required (has\_child) is dependent on the amount on new items indexed  
> > > (that  
> > > makes sense). But that time is already seconds(!) after adding a couple  
> > > of  
> > > thousand new items. And many minutes if we leave it running for a while.
> 
> > > We are currently investigating whether it is possible to add the ids to  
> > > the  
> > > internal memory map as soon as they are indexed.
> 
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> 
> > > > Thanks,  
> > > > -Tim  
> > > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> 
> > > > Thanks,  
> > > > -Tim  
> > > > On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:
> 
> > Hey haarts,  
> > I'd be really interested to see what you come up with here. So from  
> > the sound of it, until a document turns up in a query it's not added  
> > to the memory map? Can you point me to the area of code where that  
> > memory map is managed?
> 
> > Thanks,  
> > -Tim
> 
> > On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> > 
> > > Let me chime in as well.  
> > > Our problem is similar. We have about 700M items and add items at a  
> > > speed  
> > > of 100/s. The performance we are seeing is not great. The query time  
> > > required (has\_child) is dependent on the amount on new items indexed  
> > > (that  
> > > makes sense). But that time is already seconds(!) after adding a couple  
> > > of  
> > > thousand new items. And many minutes if we leave it running for a while.
> 
> > > We are currently investigating whether it is possible to add the ids to  
> > > the  
> > > internal memory map as soon as they are indexed.
> 
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> 
> > > > Thanks,  
> > > > -Tim  
> > > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> 
> > > > Thanks,  
> > > > -Tim  
> > > > On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:
> 
> > Hey haarts,  
> > I'd be really interested to see what you come up with
> 
> ...
> 
> read more »

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [February 29, 2012, 2:11pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/12 "2012-02-29T14:11:56Z")

</div>

Do you have a replica for the shard? If so, it will also need to be loaded on the replicas. See the other thread about why the cache itself is not really reloaded or managed for each change (or refresh).

On Tuesday, February 28, 2012 at 5:59 PM, haarts wrote:

> Hi Tim,
> 
> Just an update.  
> We are encountering some unexpected behaviour when running the same has\_child query twice consecutively. The first query takes about 30 minutes to complete, as does the second one. This in contrary to the believe that once the IDs are loaded in memory search should be fast. The index contains 36M documents and the index size is 30GB running on an 8 core i7 with 24GB RAM.
> 
> Regarding your question on when an ID is loaded to the memory map; I believe they are all loaded all the time.  
> I believe the code responsible is in java/org/elasticsearch/index/cache/id/simple/SimpleIdCache.java ([https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/cache/id/simple/SimpleIdCache.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/index/cache/id/simple/SimpleIdCache.java)).
> 
> Harm
> 
> On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:
> 
> > Hey haarts,  
> > I'd be really interested to see what you come up with here. So from  
> > the sound of it, until a document turns up in a query it's not added  
> > to the memory map? Can you point me to the area of code where that  
> > memory map is managed?
> > 
> > Thanks,  
> > -Tim
> > 
> > On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> > 
> > > Let me chime in as well.  
> > > Our problem is similar. We have about 700M items and add items at a speed  
> > > of 100/s. The performance we are seeing is not great. The query time  
> > > required (has\_child) is dependent on the amount on new items indexed (that  
> > > makes sense). But that time is already seconds(!) after adding a couple of  
> > > thousand new items. And many minutes if we leave it running for a while.
> > > 
> > > We are currently investigating whether it is possible to add the ids to the  
> > > internal memory map as soon as they are indexed.
> > > 
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > > 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> > > 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> > > 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> > > 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> > > 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> > > 
> > > > Thanks,  
> > > > -Tim  
> > > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > > 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> > > 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> > > 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> > > 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> > > 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> > > 
> > > > Thanks,  
> > > > -Tim  
> > > > On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:  
> > > > Hey haarts,  
> > > > I'd be really interested to see what you come up with here. So from  
> > > > the sound of it, until a document turns up in a query it's not added  
> > > > to the memory map? Can you point me to the area of code where that  
> > > > memory map is managed?
> > 
> > Thanks,  
> > -Tim
> > 
> > On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> > 
> > > Let me chime in as well.  
> > > Our problem is similar. We have about 700M items and add items at a speed  
> > > of 100/s. The performance we are seeing is not great. The query time  
> > > required (has\_child) is dependent on the amount on new items indexed (that  
> > > makes sense). But that time is already seconds(!) after adding a couple of  
> > > thousand new items. And many minutes if we leave it running for a while.
> > > 
> > > We are currently investigating whether it is possible to add the ids to the  
> > > internal memory map as soon as they are indexed.
> > > 
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > > 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> > > 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> > > 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> > > 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> > > 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> > > 
> > > > Thanks,  
> > > > -Tim  
> > > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > > 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> > > 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> > > 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> > > 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> > > 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> > > 
> > > > Thanks,  
> > > > -Tim  
> > > > On Tuesday, 28 February 2012 14:19:40 UTC+1, Tim J wrote:  
> > > > Hey haarts,  
> > > > I'd be really interested to see what you come up with here. So from  
> > > > the sound of it, until a document turns up in a query it's not added  
> > > > to the memory map? Can you point me to the area of code where that  
> > > > memory map is managed?
> > 
> > Thanks,  
> > -Tim
> > 
> > On Feb 28, 6:28 am, haarts [harmaa...@gmail.com](mailto:harmaa...@gmail.com) wrote:
> > 
> > > Let me chime in as well.  
> > > Our problem is similar. We have about 700M items and add items at a speed  
> > > of 100/s. The performance we are seeing is not great. The query time  
> > > required (has\_child) is dependent on the amount on new items indexed (that  
> > > makes sense). But that time is already seconds(!) after adding a couple of  
> > > thousand new items. And many minutes if we leave it running for a while.
> > > 
> > > We are currently investigating whether it is possible to add the ids to the  
> > > internal memory map as soon as they are indexed.
> > > 
> > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > > 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> > > 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> > > 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> > > 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> > > 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> > > 
> > > > Thanks,  
> > > > -Tim  
> > > > On Thursday, 23 February 2012 19:59:44 UTC+1, Tim J wrote:
> > > 
> > > > Hey folks,  
> > > > I'm currently working with an ES index of roughly 52 million  
> > > > documents. We index approximately 10-20 new docs per second. Each  
> > > > document is broken into two pieces and indexed as a parent/child  
> > > > pair. The child contains static content and is unlikely to ever be  
> > > > updated. The parent fields are modified frequently which is why the  
> > > > child content was separated, particularly as the original source for  
> > > > the child documents is expensive to retrieve.
> > > 
> > > > Documents are replicated across three nodes. No data is stored with  
> > > > the exception of a unique id for each doc. Each node is allocated 8  
> > > > GB of RAM and we occupy about 22 GB per node on disk. We use the  
> > > > routing key to "shard" our data. There are approximately 130  
> > > > different routing keys in use at the moment. Routing keys are also  
> > > > used as conditions for all searches so they should be a quick filter.
> > > 
> > > > First, does anyone have a sense of the penalty we're paying for  
> > > > having this parent/child relationship? We're seeing some very long  
> > > > query times particularly when we're actively writing to the nodes.  
> > > > Sometimes a simple query with one condition on the parent and one in a  
> > > > has\_child can take 8+ minutes.
> > > 
> > > > I've noticed that when we're doing a lot of writes to the child  
> > > > index in particular the times go up significantly. On the other hand  
> > > > if we only write to the parent index this is much less of a problem.  
> > > > Is this expected?
> > > 
> > > > Finally, does anyone have any suggestions for tuning this  
> > > > configuration or improving our queries?
> > > 
> > > > Thanks,  
> > > > -Tim

---

<div class="post-metadata">

**Author:** ![Nick\_Hoffman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nick_hoffman/32/1872_2.png) [@Nick\_Hoffman](https://discuss.elastic.co/u/Nick_Hoffman)\
**Post date:** [February 29, 2012, 2:42pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/13 "2012-02-29T14:42:05Z")

</div>

On Wednesday, 29 February 2012 09:11:00 UTC-5, Tim J wrote:

> Thanks Harm. I think we may be looking at alternative solutions for  
> this one as our deadline is fast approaching. I'll be keeping an eye  
> on this thread though in case you turn up anything good!
> 
> -Tim

Are you able to provide additional selection restrictions to the has\_child  
query? I would imagine that the more restrictive you can be for which  
children match the has\_child query, the faster your query search will be.

---

<div class="post-metadata">

**Author:** ![Tim\_J](https://avatars.discourse-cdn.com/v4/letter/t/f17d59/32.png) [@Tim\_J](https://discuss.elastic.co/u/Tim_J)\
**Post date:** [February 29, 2012, 5:44pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/14 "2012-02-29T17:44:53Z")

</div>

Unfortunately our query is about as restricted as it can be. The  
child documents contain a single field (text content) and a \_routing  
key. All our queries are only interested in one \_routing key so we  
use that as a filter. Otherwise, it's just a query against the  
content.

That does does raise an interesting question though. We often add a  
number of filters to the parent document which would significantly  
reduce the result set. Would it make any sense to reverse that parent/  
child relationship?

Thanks,  
-Tim

On Feb 29, 9:42 am, Nick Hoffman [n...@deadorange.com](mailto:n...@deadorange.com) wrote:

> On Wednesday, 29 February 2012 09:11:00 UTC-5, Tim J wrote:
> 
> > Thanks Harm. I think we may be looking at alternative solutions for  
> > this one as our deadline is fast approaching. I'll be keeping an eye  
> > on this thread though in case you turn up anything good!
> 
> > -Tim
> 
> Are you able to provide additional selection restrictions to the has\_child  
> query? I would imagine that the more restrictive you can be for which  
> children match the has\_child query, the faster your query search will be.

---

<div class="post-metadata">

**Author:** ![Serg\_Pilipenko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/serg_pilipenko/32/13833_2.png) [@Serg\_Pilipenko](https://discuss.elastic.co/u/Serg_Pilipenko)\
**Post date:** [October 21, 2012, 10:25pm UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/15 "2012-10-21T22:25:33Z")

</div>

I've created issue here  
[Improve implementation of SimpleIdCache · Issue #2343 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/2343) Please vote

четверг, 23 февраля 2012 г., 19:59:44 UTC+1 пользователь Tim J написал:

> Hey folks,  
> I'm currently working with an ES index of roughly 52 million  
> documents. We index approximately 10-20 new docs per second. Each  
> document is broken into two pieces and indexed as a parent/child  
> pair. The child contains static content and is unlikely to ever be  
> updated. The parent fields are modified frequently which is why the  
> child content was separated, particularly as the original source for  
> the child documents is expensive to retrieve.
> 
> Documents are replicated across three nodes. No data is stored with  
> the exception of a unique id for each doc. Each node is allocated 8  
> GB of RAM and we occupy about 22 GB per node on disk. We use the  
> routing key to "shard" our data. There are approximately 130  
> different routing keys in use at the moment. Routing keys are also  
> used as conditions for all searches so they should be a quick filter.
> 
> First, does anyone have a sense of the penalty we're paying for  
> having this parent/child relationship? We're seeing some very long  
> query times particularly when we're actively writing to the nodes.  
> Sometimes a simple query with one condition on the parent and one in a  
> has\_child can take 8+ minutes.
> 
> I've noticed that when we're doing a lot of writes to the child  
> index in particular the times go up significantly. On the other hand  
> if we only write to the parent index this is much less of a problem.  
> Is this expected?
> 
> Finally, does anyone have any suggestions for tuning this  
> configuration or improving our queries?
> 
> Thanks,  
> -Tim

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:07am UTC](https://discuss.elastic.co/t/performance-penalty-for-has-child-queries/6794/16 "2017-07-06T03:07:41Z")

</div>


