# Has\_child / has\_parent queries for a large DB

**URL:** <https://discuss.elastic.co/t/has-child-has-parent-queries-for-a-large-db/11855>\
**Category:** Elasticsearch\
**Created:** [May 7, 2013, 2:20pm UTC](https://discuss.elastic.co/t/has-child-has-parent-queries-for-a-large-db/11855 "2013-05-07T14:20:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Eran\_Eidinger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eran_eidinger/32/44831_2.png) [@Eran\_Eidinger](https://discuss.elastic.co/u/Eran_Eidinger)\
**Post date:** [May 7, 2013, 2:20pm UTC](https://discuss.elastic.co/t/has-child-has-parent-queries-for-a-large-db/11855/1 "2013-05-07T14:20:28Z")

</div>

I have a single shard with A(parent) and B (child) documents.  
At first "has\_child" queries were great.

Now, that my database has several millions of records and about 20GB of data, the queries take ALOT of time.  
regular queries work fine.

I read that there needs to be an initial loading of data to memory for has\_child/has\_parent queries to work, and so the first time should take more than the next ones.

However, the first query takes 30 minutes or more, sometimes fails miserably.

What can I do to help this?

1. I tried increasing ES\_HEAP\_SIZE, ES\_MIN\_MEM and ES\_MAX\_MEM, all to the same value. Is that wise?  
should I only have the set ES\_HEAP\_SIZE and left the others unspecified (there seems to be some confusion in the documentation as to the difference between the three).

If my machine has 2G memory, what should I set this value to, 1500m?

1. Will adding shards help ( this means reindexing, right? I can't just define more shards)

2. If I have some known queries that I keep repeating - is there a way to have them indexed or run some kind of map-reduce periodically?

If not, what would you recommend I should do?

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [May 8, 2013, 9:54am UTC](https://discuss.elastic.co/t/has-child-has-parent-queries-for-a-large-db/11855/2 "2013-05-08T09:54:26Z")

</div>

The has\_parent / has\_child queries rely on a in memory id cache, to run  
performantly. You need to have enough memory available to accomodate this  
id cache. You can view in the node stats api how much memory the id\_cache  
is taking up in the heap space.

Can you tell a bit more about your ES setup (how many nodes, how many  
indices and primary/replica shards per index)?  
2GB per machine isn't that much per machine and I think in your case is the  
cause of your problem for the has\_parent and has\_child queries. By default  
an index has 5 primary shards, this means you can add just more machines,  
this will spread the memory usage across more machines.

The id cache is loaded when the first has\_child / has\_parent query is  
executed and then reused for subsequent search requests. You can use a  
warmer with a has\_parent / has\_child query to preload the id cache, before  
actual search requests are executed:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Martijn

On 7 May 2013 16:20, eranid [eranid@gmail.com](mailto:eranid@gmail.com) wrote:

> I have a single shard with A(parent) and B (child) documents.  
> At first "has\_child" queries were great.
> 
> Now, that my database has several millions of records and about 20GB of  
> data, the queries take ALOT of time.  
> regular queries work fine.
> 
> I read that there needs to be an initial loading of data to memory for  
> has\_child/has\_parent queries to work, and so the first time should take  
> more  
> than the next ones.
> 
> However, the first query takes 30 minutes or more, sometimes fails  
> miserably.
> 
> What can I do to help this?
> 
> 1. I tried increasing ES\_HEAP\_SIZE, ES\_MIN\_MEM and ES\_MAX\_MEM, all to the  
> same value. Is that wise?  
> should I only have the set ES\_HEAP\_SIZE and left the others unspecified  
> (there seems to be some confusion in the documentation as to the difference  
> between the three).
> 
> If my machine has 2G memory, what should I set this value to, 1500m?
> 
> 1. Will adding shards help ( this means reindexing, right? I can't just  
> define more shards)
> 
> 2. If I have some known queries that I keep repeating - is there a way to  
> have them indexed or run some kind of map-reduce periodically?
> 
> If not, what would you recommend I should do?
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/has-child-has-parent-queries-for-a-large-DB-tp4034383.html](http://elasticsearch-users.115913.n3.nabble.com/has-child-has-parent-queries-for-a-large-DB-tp4034383.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Met vriendelijke groet,

Martijn van Groningen

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Eran\_Eidinger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eran_eidinger/32/44831_2.png) [@Eran\_Eidinger](https://discuss.elastic.co/u/Eran_Eidinger)\
**Post date:** [May 8, 2013, 12:17pm UTC](https://discuss.elastic.co/t/has-child-has-parent-queries-for-a-large-db/11855/3 "2013-05-08T12:17:14Z")

</div>

Thanks!  
I have one node with one shard and no replication (I was planning on expanding it as I understand ES better...)  
I'm using an amazon m1.medium or m1.large machine for this.

Looking at the id\_cache size, it's zero, both before and during the query i try to run:  
id\_cache\_size: 0b

(btw, what is the id\_cache size, and how does this differ from the heap size? or is this the same...)

So, what you suggest is adding more machines?

What i mainly feel is missing, is my understanding of how to monitor the cluster and understand what the problem is.

I attach a printout of querying with ...:9200/\_cluster/nodes/stats?process=true&os=true&fs=true&network=true

{  
cluster\_name: elasticsearch  
nodes: {  
rvPqcyaxRkKCqTHs3T2qBA: {  
timestamp: 1368014136488  
name: Rainbow  
transport\_address: inet[ip-10-137-40-231.ec2.internal/10.137.40.231:9300]  
hostname: ip-10-137-40-231  
attributes: { aws\_availability\_zone: us-east-1c}  
indices: {  
store: {  
size: 26.4gb  
size\_in\_bytes: 28389539108  
throttle\_time: 0s  
throttle\_time\_in\_millis: 0  
}  
docs: {  
count: 16795436  
deleted: 260670  
}  
indexing: {  
index\_total: 3478423  
index\_time: 1.4h  
index\_time\_in\_millis: 5346619  
index\_current: 3  
delete\_total: 210748  
delete\_time: 4.9m  
delete\_time\_in\_millis: 298083  
delete\_current: 0  
}  
get: {  
total: 263490  
time: 1.6m  
time\_in\_millis: 96962  
exists\_total: 263353  
exists\_time: 1.6m  
exists\_time\_in\_millis: 96944  
missing\_total: 137  
missing\_time: 18ms  
missing\_time\_in\_millis: 18  
current: 0  
}  
search: {  
query\_total: 2494237  
query\_time: 8.9h  
query\_time\_in\_millis: 32390632  
query\_current: 4  
fetch\_total: 2494236  
fetch\_time: 14.5m  
fetch\_time\_in\_millis: 872967  
fetch\_current: 1  
}  
cache: {  
field\_evictions: 0  
field\_size: 189mb  
field\_size\_in\_bytes: 198210516  
filter\_count: 10  
filter\_evictions: 0  
filter\_size: 13.4mb  
filter\_size\_in\_bytes: 14136296  
bloom\_size: 28.8mb  
bloom\_size\_in\_bytes: 30230528  
id\_cache\_size: 0b  
id\_cache\_size\_in\_bytes: 0  
}  
merges: {  
current: 2  
current\_docs: 22961  
current\_size: 19.4mb  
current\_size\_in\_bytes: 20388239  
total: 11509  
total\_time: 3h  
total\_time\_in\_millis: 10914669  
total\_docs: 40842823  
total\_size: 44.6gb  
total\_size\_in\_bytes: 47969339439  
}  
refresh: {  
total: 83971  
total\_time: 46.7m  
total\_time\_in\_millis: 2805647  
}  
flush: {  
total: 925  
total\_time: 19.8m  
total\_time\_in\_millis: 1190485  
}  
}  
os: {  
timestamp: 1368014136489  
uptime: 18 hours, 7 minutes and 52 seconds  
uptime\_in\_millis: 65272000  
load\_average: [  
2.92  
3.38  
2.83  
]  
cpu: {  
sys: 1  
user: 97  
idle: 0  
}  
mem: {  
free: 607.7mb  
free\_in\_bytes: 637239296  
used: 3gb  
used\_in\_bytes: 3295408128  
free\_percent: 59  
used\_percent: 40  
actual\_free: 2.1gb  
actual\_free\_in\_bytes: 2331492352  
actual\_used: 1.4gb  
actual\_used\_in\_bytes: 1601155072  
}  
swap: {  
used: 0b  
used\_in\_bytes: 0  
free: 0b  
free\_in\_bytes: 0  
}  
}  
process: {  
timestamp: 1368014136490  
open\_file\_descriptors: 1139  
cpu: {  
percent: 99  
sys: 1 hour, 35 minutes, 40 seconds and 410 milliseconds  
sys\_in\_millis: 5740410  
user: 14 hours, 29 minutes, 20 seconds and 440 milliseconds  
user\_in\_millis: 52160440  
total: 16 hours, 5 minutes and 850 milliseconds  
total\_in\_millis: 57900850  
}  
mem: {  
resident: 1.3gb  
resident\_in\_bytes: 1451016192  
share: 5.4mb  
share\_in\_bytes: 5746688  
total\_virtual: 2.1gb  
total\_virtual\_in\_bytes: 2266382336  
}  
}  
network: {  
tcp: {  
active\_opens: 50  
passive\_opens: 2971419  
curr\_estab: 181  
in\_segs: 20468767  
out\_segs: 19617991  
retrans\_segs: 6901  
estab\_resets: 15  
attempt\_fails: 0  
in\_errs: 9  
out\_rsts: 106  
}  
}  
fs: {  
timestamp: 1368014136506  
data: [  
{  
path: /var/lib/elasticsearch/elasticsearch/nodes/0  
mount: /  
dev: /dev/xvda1  
total: 98.4gb  
total\_in\_bytes: 105689415680  
free: 65.6gb  
free\_in\_bytes: 70445109248  
available: 60.6gb  
available\_in\_bytes: 65077608448  
disk\_reads: 487455  
disk\_writes: 1402229  
disk\_read\_size: 10.9gb  
disk\_read\_size\_in\_bytes: 11780604928  
disk\_write\_size: 45.3gb  
disk\_write\_size\_in\_bytes: 48698900480  
}  
]  
}  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:37am UTC](https://discuss.elastic.co/t/has-child-has-parent-queries-for-a-large-db/11855/4 "2017-07-06T02:37:39Z")

</div>


