# Parent/child not works on some data

**URL:** <https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358>\
**Category:** Elasticsearch\
**Created:** [October 15, 2012, 4:21pm UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358 "2012-10-15T16:21:51Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Serg\_Pilipenko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/serg_pilipenko/32/13833_2.png) [@Serg\_Pilipenko](https://discuss.elastic.co/u/Serg_Pilipenko)\
**Post date:** [October 15, 2012, 4:21pm UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358/1 "2012-10-15T16:21:51Z")

</div>

Hi all!  
I have experienced problem with searching using of parent/child relation  
between document.

I have the following mappings:

> <https://gist.github.com/serj-p/3893320>

And I'm running the following request:  
{  
"timeout": 180000,  
"query": {  
"filtered": {  
"filter": {  
"and": {  
"filters": [  
{  
"term": {  
"company\_id": "4d07f9c8775911968cab4a80"  
}  
},  
{  
"has\_child": {  
"query": {  
"filtered": {  
"filter": {  
"and": [  
{  
"term": {  
"company\_id": "4d07f9c8775911968cab4a80"  
}  
},  
{  
"term": {  
"name": "twitter"  
}  
}  
]  
},  
"query": {  
"match\_all": {}  
}  
}  
},  
"type": "tag"  
}  
}  
]  
}  
},  
"query": {  
"bool": {  
"must": [  
{  
"match\_all": {}  
}  
],  
"should": []  
}  
}  
}  
}  
}

You can see timeout option in request.  
When I have indexed about 2Gb of data (all data for "company\_id":  
"4d07f9c8775911968cab4a80") this query works OK. Request takes about 50ms.  
But when I have indexed all required data (about 35Gb that includes  
additional companies) after starting executing this query cluster hungs for  
few hours. No errors in log for this time. After some time I can see that  
some shards gone to "not initialized" state. Cluster becomes available only  
after restart of all nodes.

I'm suspecting that "has\_child" query doesn't work either on some specific  
documents or on bigger datasets.

Used ES version 19.10  
Cluster consists of 3 aws m1.xlarge instances. ES works with the following  
jvm options:  
-Xms14g -Xmx14g -Xss256k -Djava.awt.headless=true -XX:+UseParNewGC  
-XX:+UseConcMarkSweepGC -XX:CMSInitiatingOccupancyFraction=75  
-XX:+UseCMSInitiatingOccupancyOnly -XX:+HeapDumpOnOutOfMemoryError  
-Delasticsearch -Des.foreground=yes -Djava.net.preferIPv4Stack=true

--

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [October 15, 2012, 9:05pm UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358/2 "2012-10-15T21:05:44Z")

</div>

Hi Serg,

That doesn't look good. Can you share with us your nodes stats  
([http://localhost:9200/\_nodes/stats?all](http://localhost:9200/_nodes/stats?all))?  
Also when the cluster hangs, can you query the hot threads api  
([http://localhost:9200/\_nodes/hot\_threads](http://localhost:9200/_nodes/hot_threads))?  
How many primary and replica shards you have configured for your index?

The has\_child (and others like has\_parent and top\_children) need an  
in-memory data structure to perform efficiently. The size of this data  
structure depends on the number of unique parent ids and the number of  
documents. It is possible that you need to increase the heap space  
size for ES or increase the number of nodes.

Martijn

On 15 October 2012 18:21, Serg Pilipenko [cloun.rules@gmail.com](mailto:cloun.rules@gmail.com) wrote:

> Hi all!  
> I have experienced problem with searching using of parent/child relation  
> between document.
> 
> I have the following mappings:  
> [ES query and mapping for failing parent/child request · GitHub](https://gist.github.com/3893320)
> 
> And I'm running the following request:  
> {  
> "timeout": 180000,  
> "query": {  
> "filtered": {  
> "filter": {  
> "and": {  
> "filters": [  
> {  
> "term": {  
> "company\_id": "4d07f9c8775911968cab4a80"  
> }  
> },  
> {  
> "has\_child": {  
> "query": {  
> "filtered": {  
> "filter": {  
> "and": [  
> {  
> "term": {  
> "company\_id": "4d07f9c8775911968cab4a80"  
> }  
> },  
> {  
> "term": {  
> "name": "twitter"  
> }  
> }  
> ]  
> },  
> "query": {  
> "match\_all": {}  
> }  
> }  
> },  
> "type": "tag"  
> }  
> }  
> ]  
> }  
> },  
> "query": {  
> "bool": {  
> "must": [  
> {  
> "match\_all": {}  
> }  
> ],  
> "should":   
> }  
> }  
> }  
> }  
> }
> 
> You can see timeout option in request.  
> When I have indexed about 2Gb of data (all data for "company\_id":  
> "4d07f9c8775911968cab4a80") this query works OK. Request takes about 50ms.  
> But when I have indexed all required data (about 35Gb that includes  
> additional companies) after starting executing this query cluster hungs for  
> few hours. No errors in log for this time. After some time I can see that  
> some shards gone to "not initialized" state. Cluster becomes available only  
> after restart of all nodes.
> 
> I'm suspecting that "has\_child" query doesn't work either on some specific  
> documents or on bigger datasets.
> 
> Used ES version 19.10  
> Cluster consists of 3 aws m1.xlarge instances. ES works with the following  
> jvm options:  
> -Xms14g -Xmx14g -Xss256k -Djava.awt.headless=true -XX:+UseParNewGC  
> -XX:+UseConcMarkSweepGC -XX:CMSInitiatingOccupancyFraction=75  
> -XX:+UseCMSInitiatingOccupancyOnly -XX:+HeapDumpOnOutOfMemoryError  
> -Delasticsearch -Des.foreground=yes -Djava.net.preferIPv4Stack=true
> 
> --

--  
Met vriendelijke groet,

Martijn van Groningen

--

---

<div class="post-metadata">

**Author:** ![Serg\_Pilipenko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/serg_pilipenko/32/13833_2.png) [@Serg\_Pilipenko](https://discuss.elastic.co/u/Serg_Pilipenko)\
**Post date:** [October 15, 2012, 9:53pm UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358/3 "2012-10-15T21:53:09Z")

</div>

> That doesn't look good. Can you share with us your nodes stats  
> ([http://localhost:9200/\_nodes/stats?all](http://localhost:9200/_nodes/stats?all))?  
> Also when the cluster hangs, can you query the hot threads api  
> ([http://localhost:9200/\_nodes/hot\_threads](http://localhost:9200/_nodes/hot_threads))?  
> How many primary and replica shards you have configured for your index?

I'll provide this stats tomorrow.

The has\_child (and others like has\_parent and top\_children) need an

> in-memory data structure to perform efficiently. The size of this data  
> structure depends on the number of unique parent ids and the number of  
> documents. It is possible that you need to increase the heap space  
> size for ES or increase the number of nodes.

Does this implementation take into account filters from main query and  
subquery?

--

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [October 15, 2012, 9:58pm UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358/4 "2012-10-15T21:58:41Z")

</div>

> Does this implementation take into account filters from main query and  
> subquery?

No. The cached data structure is reused across search requests and  
therefore doesn't take into account a search request's query and  
filter.

Martijn

--

---

<div class="post-metadata">

**Author:** ![Serg\_Pilipenko](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/serg_pilipenko/32/13833_2.png) [@Serg\_Pilipenko](https://discuss.elastic.co/u/Serg_Pilipenko)\
**Post date:** [October 17, 2012, 12:51am UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358/5 "2012-10-17T00:51:01Z")

</div>

> That doesn't look good. Can you share with us your nodes stats  
> ([http://localhost:9200/\_nodes/stats?all](http://localhost:9200/_nodes/stats?all))?

> <https://gist.github.com/serj-p/3902994>

> Also when the cluster hangs, can you query the hot threads api  
> ([http://localhost:9200/\_nodes/hot\_threads](http://localhost:9200/_nodes/hot_threads))?

for i in {1..250}; do curl -XGET  
'[http://127.0.0.1:9200/\_nodes/hot\_threads?interval=2000](http://127.0.0.1:9200/_nodes/hot_threads?interval=2000)' \>\>  
hot\_threads.txt; done;  
[hot\_threads · GitHub](https://gist.github.com/3903012)

And no still no response for a long time.

> How many primary and replica shards you have configured for your index?

15 primary shards, 1 replica set

Are there any other possible approaches to implement relation between  
different docs and to do not perform 2 requests with computing on client  
side?  
The main problem for me is that adding/deletion of tag will cause  
reindexation of thousands of contacts if I include tag doc into contact doc.

--

---

<div class="post-metadata">

**Author:** ![mvg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mvg/32/98890_2.png) [@mvg](https://discuss.elastic.co/u/mvg)\
**Post date:** [October 17, 2012, 8:18am UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358/6 "2012-10-17T08:18:08Z")

</div>

Hi Serg,

Did you ran the nodes stats request before or after you started  
testing your queries? Seems that only a fraction of the allocated heap  
space is actually used.

I see that the m1.xlarge instance has 15GB of memory available. In  
your case you allocated 14GB of that to ES's heap space. ES depends a  
lot on the OS file system cache. Right now the OS has only 1GB left.  
This can make any kind of query slow. Usually a healthy balance is 50%  
of the available memory to ES and the other 50% to OS. I'd set the  
ES\_HEAP\_SIZE to 7GB.

From what I can see from the hot threads output is that it is loading  
the data structures used by the has\_child query. During the first  
has\_child query execution on a fresh index, the data structure it  
needs is loaded from disk into memory. This can make the first  
execution of the has\_child query slow. Subsequent has\_child queries  
should be much faster.

Martijn

On 17 October 2012 02:51, Serg Pilipenko [cloun.rules@gmail.com](mailto:cloun.rules@gmail.com) wrote:

> > That doesn't look good. Can you share with us your nodes stats  
> > ([http://localhost:9200/\_nodes/stats?all](http://localhost:9200/_nodes/stats?all))?
> 
> [nodes stats · GitHub](https://gist.github.com/3902994)
> 
> > Also when the cluster hangs, can you query the hot threads api  
> > ([http://localhost:9200/\_nodes/hot\_threads](http://localhost:9200/_nodes/hot_threads))?
> 
> for i in {1..250}; do curl -XGET  
> '[http://127.0.0.1:9200/\_nodes/hot\_threads?interval=2000](http://127.0.0.1:9200/_nodes/hot_threads?interval=2000)' \>\> hot\_threads.txt;  
> done;  
> [hot\_threads · GitHub](https://gist.github.com/3903012)
> 
> And no still no response for a long time.
> 
> > How many primary and replica shards you have configured for your index?
> 
> 15 primary shards, 1 replica set
> 
> Are there any other possible approaches to implement relation between  
> different docs and to do not perform 2 requests with computing on client  
> side?  
> The main problem for me is that adding/deletion of tag will cause  
> reindexation of thousands of contacts if I include tag doc into contact doc.
> 
> --

--  
Met vriendelijke groet,

Martijn van Groningen

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:08am UTC](https://discuss.elastic.co/t/parent-child-not-works-on-some-data/9358/7 "2017-07-06T03:08:22Z")

</div>


