# Parent/Child use case

**URL:** <https://discuss.elastic.co/t/parent-child-use-case/5401>\
**Category:** Elasticsearch\
**Created:** [September 19, 2011, 12:41am UTC](https://discuss.elastic.co/t/parent-child-use-case/5401 "2011-09-19T00:41:56Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Paul\_Smith](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_smith/32/1323_2.png) [@Paul\_Smith](https://discuss.elastic.co/u/Paul_Smith)\
**Post date:** [September 19, 2011, 12:41am UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/1 "2011-09-19T00:41:56Z")

</div>

Just wanted to check whether this scenario fit properly the parent/child  
mapping feature.

We currently index just meta-data of documents (dozens of fields), however  
we want to index file contents too as that's sometimes useful for our  
customers (our use case, the meta data is the primary mechanism). Since we  
have hundreds of millions of document records, and 100Tb+ filesize, it's a  
non-trivial exercise we've managed to put off for a while.

Since any reindex requires indexing both meta-data and file content, which  
for us is kept separately in DB & fileserver respectively, I didn't want any  
meta-data update to also require a seek of the filestore to get the text  
content for indexing especially since the text-content of a file never  
changes (for us). I was hoping to find a way to keep meta-data and text  
content separate in the index and meta- updates update independently.

I was thinking of having a parent/child relationship between the  
meta(parent) and the full text (child), allowing the parent to update  
(frequent) leaving the child pretty much alone once text extracted and  
indexed. Text extraction of newly uploaded files can be done async, and a  
new child record added in ES independent of the registration of the document  
meta data record.

Does this sound the right use case for parent/child in ES?

As I understand it, if we needed to reindex (say, new fields, or changed  
values or something) then we'd also have to reindex the children, but we  
could do these 2 reindex operations separately, mark the full text bit  
'offline' until that's completed, allowing the meta-data to be searched much  
earlier.

Shay, it would help too in the Docs if the 'parent/chi/d' bit referred to  
frequently is easy to find in the docs, I'm presuming it's the 'nested'  
mapping type.. ? I see reference to parent/child in the forums etc, and  
took me a while to bump into it in the docs when I went looking. I could be  
blind though!

thanks,

Paul

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [September 20, 2011, 10:20pm UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/2 "2011-09-20T22:20:04Z")

</div>

Hey Paul,  
First off, the parent/child support and the nested support are  
actually two different features. The nested support seems to be the  
more feature rich variant, though.

I was interested in both these features for storing a frequently  
changing popularity score without needing to re-index the document.  
Unfortunately, neither of these features fit my use case. Parent/child  
can be indexed separately, but it isn't possible to join the child  
document for the sorting of the parent. For nested, all data needs to  
be re-indexed due to how the data is stored internally in ES.

We avoid having to hit our backend datastore for meta updates by  
pulling the data in ES, updating it and resubmitting. Although, there  
is a new plugin that does exactly this. Either way, the whole document  
needs to get re-indexed.

Don't take this as the definitive answer, though, I've only briefly  
played around with both these features.

Best Regards,  
Paul

On Sep 18, 6:41 pm, Paul Smith [tallpsm...@gmail.com](mailto:tallpsm...@gmail.com) wrote:

> Just wanted to check whether this scenario fit properly the parent/child  
> mapping feature.
> 
> We currently index just meta-data of documents (dozens of fields), however  
> we want to index file contents too as that's sometimes useful for our  
> customers (our use case, the meta data is the primary mechanism). Since we  
> have hundreds of millions of document records, and 100Tb+ filesize, it's a  
> non-trivial exercise we've managed to put off for a while.
> 
> Since any reindex requires indexing both meta-data and file content, which  
> for us is kept separately in DB & fileserver respectively, I didn't want any  
> meta-data update to also require a seek of the filestore to get the text  
> content for indexing especially since the text-content of a file never  
> changes (for us). I was hoping to find a way to keep meta-data and text  
> content separate in the index and meta- updates update independently.
> 
> I was thinking of having a parent/child relationship between the  
> meta(parent) and the full text (child), allowing the parent to update  
> (frequent) leaving the child pretty much alone once text extracted and  
> indexed. Text extraction of newly uploaded files can be done async, and a  
> new child record added in ES independent of the registration of the document  
> meta data record.
> 
> Does this sound the right use case for parent/child in ES?
> 
> As I understand it, if we needed to reindex (say, new fields, or changed  
> values or something) then we'd also have to reindex the children, but we  
> could do these 2 reindex operations separately, mark the full text bit  
> 'offline' until that's completed, allowing the meta-data to be searched much  
> earlier.
> 
> Shay, it would help too in the Docs if the 'parent/chi/d' bit referred to  
> frequently is easy to find in the docs, I'm presuming it's the 'nested'  
> mapping type.. ? I see reference to parent/child in the forums etc, and  
> took me a while to bump into it in the docs when I went looking. I could be  
> blind though!
> 
> thanks,
> 
> Paul

---

<div class="post-metadata">

**Author:** ![Paul\_Smith](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_smith/32/1323_2.png) [@Paul\_Smith](https://discuss.elastic.co/u/Paul_Smith)\
**Post date:** [September 20, 2011, 11:21pm UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/3 "2011-09-20T23:21:43Z")

</div>

Thanks for the reply!

On 21 September 2011 08:20, ppearcy [ppearcy@gmail.com](mailto:ppearcy@gmail.com) wrote:

> Hey Paul,  
> First off, the parent/child support and the nested support are  
> actually two different features. The nested support seems to be the  
> more feature rich variant, though.

_scratches head_ So is there a web page on [elasticsearch.com](http://elasticsearch.com) that details  
the parent/child? I'm going blind, I could only find the 'nested' one  
then.. ?

> I was interested in both these features for storing a frequently  
> changing popularity score without needing to re-index the document.  
> Unfortunately, neither of these features fit my use case. Parent/child  
> can be indexed separately, but it isn't possible to join the child  
> document for the sorting of the parent. For nested, all data needs to  
> be re-indexed due to how the data is stored internally in ES.

By the '... it isn't possible to join the child document for the sorting of  
the parent' part. I don't need to sort by any child value in this case, I  
just need to be able to match on text in the child value sometimes and  
return the parent as the hit.

So, if the parent ES document has a field called "documentnumber", and the  
child is the text contents of the file attached to this parent and the child  
has a field "contents", then if I search for:

documentnumber:ABC-123 OR content:foo

then the results should return any parent which has the field documentnumber  
with that match, PLUS any parent's whose children have 'foo' in the content.  
I then sort by one of the parent's fields.

Would this work? I'm hoping to optionally allow the customer to search the  
contents of the file, but return it 'inline' with other matches of the  
parent meta-record.

thanks,

Paul

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [September 21, 2011, 12:27am UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/4 "2011-09-21T00:27:46Z")

</div>

Ah... I _think_ that might work. Here is the best overall description  
I can find on things:

> <https://github.com/elastic/elasticsearch/issues/553>
>
> The parent/child documents support allows to define a parent relationship from a… child type to a parent type. 
> \## Mapping
> 
> The relationship is defined using a simple mapping definition at the child level mapping. For example, in case of a \`blog\` type and a \`blog\_tag\` type child document, the mapping for \`blog\_tag\` should be:
> 
> \`\`\`
> {
> "blog\_tag" : {
> "\_parent" : {
> "type" : "blog"
> }
> }
> }
> \`\`\`
> 
> The above defines a parent mapping, and the type of the parent.
> \## Indexing
> 
> When indexing a child document, it is important that it will be routed to the same shard as the parent. This uses the routing capability. When indexing a doc with a parent id, it is automatically set as the routing value (unless the routing value is explicitly defined). Indexing a document with a parent id is simple:
> 
> \`\`\`
> curl -XPUT localhost:9200/blogs/blog\_tag/1122?parent=1111 -d '
> {
> "tag" : "something"
> }
> '
> \`\`\`
> 
> There is an option to set \`\_parent\` in each bulk index item as well.
> \## Querying
> 
> There are several mechanisms to query child documents. The idea of child filter / query is that its inner query is run against the child documents, and the result of it are parent docs matching those child documents.
> 
> The way it is implemented is that the child queries are first run on their own, with the results "joining" the parent documents. Then, the main query runs with the results of the child query, which includes the parent docs.
> \# \`has\_child\`
> 
> The first is the \`has\_child\` filter and \`has\_child\` query (which is a simple \`constant\_score\` query wrapping the \`has\_child\` filter):
> 
> \`\`\`
> {
> "has\_child" : {
> "type" : "blog\_tag"
> "query" : {
> "term" : {
> "tag" : "something"
> }
> }
> }
> }
> \`\`\`
> 
> The \`type\` is the child type to query against. The parent type to return is automatically detected based on the mappings.
> 
> The query (and filter), do no scoring, and the "join" process of matching which parent doc the child doc matches is done \_on each matching child doc\_.
> \# \`top\_children\`
> 
> The \`top\_children\` query basically runs the child query with an estimated hits size, and out of this hit docs, aggregates it into parent docs. If there aren't enough parent docs matching the requested from/size search request, then it is run again with a wider (more hits) search.
> 
> The \`top\_children\` also provide scoring capabilities, with the ability to specify \`max\`, \`sum\` or \`avg\` as the \`score\` type.
> 
> One downside of using the \`top\_children\` is that if there are more child docs matching the required hits when executing the child query, then the \`total\_hits\` result of the search response will be incorrect.
> 
> How many hits are asked for in the first child query run is controlled using the \`factor\` parameter (defaults to \`5\`). For example, when asking for 10 docs with from 0, then the child query will execute with 50 hits expected. If not enough parents are found (in our example, 10), and there are still more child docs to query, then the search hits are expanded my multiplying by the \`incremental\_factor\` (defaults to \`2\`).
> 
> The required parameters are the \`query\` and \`type\` (the child type to execute the query on). Here is an example with all different parameters, including the default values:
> 
> \`\`\`
> {
> "top\_children" : {
> "type": "blog\_tag",
> "query" : {
> "term" : {
> "tag" : "something"
> }
> }
> "score" : "max",
> "factor" : 5,
> "incremental\_factor" : 2
> }
> }
> \`\`\`
> \## Faceting
> 
> Faceting on the child query phase (on the results of the query executed) can be done by specifying a \`scope\` with a custom name in the query / filter. All facets now accept a \`scope\` to run on (similar to global set to \`true\`), and can now be executed on docs matching the child query.
> \## Query Performance
> 
> In general, the \`top\_children\` performance will be much better than the \`has\_child\` performance. This is because joining the child to its parent is done in the \`top\_children\` case against the expected number of hits returned, while in the \`has\_child\` case, it is executed against \_all\_ child docs matching the child query.
> \## Memory Considerations
> 
> With the current implementation, all \`\_id\` values are loaded to memory (heap) in order to support fast lookups, so make sure there is enough mem for it.

And here are the relevant query types that can be run:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

-\> I think you'd want this one since it scores

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Each one seems to query the children and return details on the parent.  
Since you have a 1 to 1 mapping of parent to child, I think  
top\_children would work correctly.

I'd be curious to know if this ends up working for you, since I've  
always had issues wrapping my head around the primary use cases for  
this feature 🙂

Best Regards,  
Paul

On Sep 20, 5:21 pm, Paul Smith [tallpsm...@gmail.com](mailto:tallpsm...@gmail.com) wrote:

> Thanks for the reply!
> 
> On 21 September 2011 08:20, ppearcy [ppea...@gmail.com](mailto:ppea...@gmail.com) wrote:
> 
> > Hey Paul,  
> > First off, the parent/child support and the nested support are  
> > actually two different features. The nested support seems to be the  
> > more feature rich variant, though.
> 
> _scratches head_ So is there a web page on [elasticsearch.com](http://elasticsearch.com) that details  
> the parent/child? I'm going blind, I could only find the 'nested' one  
> then.. ?
> 
> > I was interested in both these features for storing a frequently  
> > changing popularity score without needing to re-index the document.  
> > Unfortunately, neither of these features fit my use case. Parent/child  
> > can be indexed separately, but it isn't possible to join the child  
> > document for the sorting of the parent. For nested, all data needs to  
> > be re-indexed due to how the data is stored internally in ES.
> 
> By the '... it isn't possible to join the child document for the sorting of  
> the parent' part. I don't need to sort by any child value in this case, I  
> just need to be able to match on text in the child value sometimes and  
> return the parent as the hit.
> 
> So, if the parent ES document has a field called "documentnumber", and the  
> child is the text contents of the file attached to this parent and the child  
> has a field "contents", then if I search for:
> 
> documentnumber:ABC-123 OR content:foo
> 
> then the results should return any parent which has the field documentnumber  
> with that match, PLUS any parent's whose children have 'foo' in the content.  
> I then sort by one of the parent's fields.
> 
> Would this work? I'm hoping to optionally allow the customer to search the  
> contents of the file, but return it 'inline' with other matches of the  
> parent meta-record.
> 
> thanks,
> 
> Paul

---

<div class="post-metadata">

**Author:** ![micuenta99](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@micuenta99](https://discuss.elastic.co/u/micuenta99)\
**Post date:** [November 15, 2011, 4:31pm UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/5 "2011-11-15T16:31:35Z")

</div>

Hi Paul,  
I also am interested in this scenario ( 'doc' (parent) and a  
'filecontent' (child)) At the end how you did it do?  
Thanks,  
On 21 sep, 00:21, Paul Smith [tallpsm...@gmail.com](mailto:tallpsm...@gmail.com) wrote:

> Thanks for the reply!
> 
> On 21 September 2011 08:20, ppearcy [ppea...@gmail.com](mailto:ppea...@gmail.com) wrote:
> 
> > Hey Paul,  
> > First off, the parent/child support and the nested support are  
> > actually two different features. The nested support seems to be the  
> > more feature rich variant, though.
> 
> _scratches head_ So is there a web page on [elasticsearch.com](http://elasticsearch.com) that details  
> the parent/child? I'm going blind, I could only find the 'nested' one  
> then.. ?
> 
> > I was interested in both these features for storing a frequently  
> > changing popularity score without needing to re-index the document.  
> > Unfortunately, neither of these features fit my use case. Parent/child  
> > can be indexed separately, but it isn't possible to join the child  
> > document for the sorting of the parent. For nested, all data needs to  
> > be re-indexed due to how the data is stored internally in ES.
> 
> By the '... it isn't possible to join the child document for the sorting of  
> the parent' part. I don't need to sort by any child value in this case, I  
> just need to be able to match on text in the child value sometimes and  
> return the parent as the hit.
> 
> So, if the parent ES document has a field called "documentnumber", and the  
> child is the text contents of the file attached to this parent and the child  
> has a field "contents", then if I search for:
> 
> documentnumber:ABC-123 OR content:foo
> 
> then the results should return any parent which has the field documentnumber  
> with that match, PLUS any parent's whose children have 'foo' in the content.  
> I then sort by one of the parent's fields.
> 
> Would this work? I'm hoping to optionally allow the customer to search the  
> contents of the file, but return it 'inline' with other matches of the  
> parent meta-record.
> 
> thanks,
> 
> Paul

---

<div class="post-metadata">

**Author:** ![Paul\_Smith](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_smith/32/1323_2.png) [@Paul\_Smith](https://discuss.elastic.co/u/Paul_Smith)\
**Post date:** [November 16, 2011, 5:02am UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/6 "2011-11-16T05:02:20Z")

</div>

On 16 November 2011 03:31, micu99 [micuenta99@gmail.com](mailto:micuenta99@gmail.com) wrote:

> Hi Paul,  
> I also am interested in this scenario ( 'doc' (parent) and a  
> 'filecontent' (child)) At the end how you did it do?

I haven't gotten around to really seriously giving this a try. I had a  
quick go with the Top Children but got stuck with a syntax error (my own  
fault I'm certain) and then was involved in some other things, so haven't  
gotten back to this one as yet sorry!

---

<div class="post-metadata">

**Author:** ![micuenta99](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@micuenta99](https://discuss.elastic.co/u/micuenta99)\
**Post date:** [November 16, 2011, 6:45am UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/7 "2011-11-16T06:45:47Z")

</div>

Thanks Paul,

Can anyone help?  
What is the best way to implement a scenario like the one that comments  
Paul?

2011/11/16 Paul Smith [tallpsmith@gmail.com](mailto:tallpsmith@gmail.com)

> On 16 November 2011 03:31, micu99 [micuenta99@gmail.com](mailto:micuenta99@gmail.com) wrote:
> 
> > Hi Paul,  
> > I also am interested in this scenario ( 'doc' (parent) and a  
> > 'filecontent' (child)) At the end how you did it do?
> 
> I haven't gotten around to really seriously giving this a try. I had a  
> quick go with the Top Children but got stuck with a syntax error (my own  
> fault I'm certain) and then was involved in some other things, so haven't  
> gotten back to this one as yet sorry!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:48am UTC](https://discuss.elastic.co/t/parent-child-use-case/5401/8 "2017-07-06T03:48:29Z")

</div>


