# Querying other documents as nested content

**URL:** <https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170>\
**Category:** Elasticsearch\
**Created:** [December 15, 2011, 7:04pm UTC](https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170 "2011-12-15T19:04:31Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nick\_Sellen](https://avatars.discourse-cdn.com/v4/letter/n/f0a364/32.png) [@Nick\_Sellen](https://discuss.elastic.co/u/Nick_Sellen)\
**Post date:** [December 15, 2011, 7:04pm UTC](https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170/1 "2011-12-15T19:04:31Z")

</div>

I'm looking at keeping two documents in the backing database (CouchDB) for  
some items - one would be small, storing the main information, and the  
other large[er] (text content). The reason being to allow revisions to be  
stored for the small document without duplicating the large document each  
time. I'm using the CouchDB \_changes based river.

I'd like to be able to do a query and included nested ids to also match on,  
e.g. something like: name:Bill nested: { ids: [42,53,35], term: "books" }

The small document would know which large documents it referred to at the  
time of indexing (but not the other way round, probably).

This seems to map quite closely to the nested document concept  
([http://www.elasticsearch.org/guide/reference/mapping/nested-type.html](http://www.elasticsearch.org/guide/reference/mapping/nested-type.html)) as  
the nested content is stored as a separate document. However that  
implementation does something which ensures the nested document is indexed  
in the same "block" which would cause trouble here as the large documents could  
be stored anywhere.

I imagine there is no way to do this at the moment but perhaps either :  
a) it's an interesting idea that could be implemented by someone at some  
point  
b) there is another way to achieve a similar result  
c) it's a stupid idea, go away

Thanks,

Nick

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 16, 2011, 3:46pm UTC](https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170/2 "2011-12-16T15:46:34Z")

</div>

If you can index the "small" document as parent, and large document as  
child, then you can use the parent child support and has\_child query /  
filter. Will that help?

On Thu, Dec 15, 2011 at 9:04 PM, Nick Sellen [talktome@nicksellen.co.uk](mailto:talktome@nicksellen.co.uk)wrote:

> I'm looking at keeping two documents in the backing database (CouchDB) for  
> some items - one would be small, storing the main information, and the  
> other large[er] (text content). The reason being to allow revisions to be  
> stored for the small document without duplicating the large document each  
> time. I'm using the CouchDB \_changes based river.
> 
> I'd like to be able to do a query and included nested ids to also match  
> on, e.g. something like: name:Bill nested: { ids: [42,53,35], term: "books"  
> }
> 
> The small document would know which large documents it referred to at the  
> time of indexing (but not the other way round, probably).
> 
> This seems to map quite closely to the nested document concept (  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/nested-type.html)) as  
> the nested content is stored as a separate document. However that  
> implementation does something which ensures the nested document is indexed  
> in the same "block" which would cause trouble here as the large documents could  
> be stored anywhere.
> 
> I imagine there is no way to do this at the moment but perhaps either :  
> a) it's an interesting idea that could be implemented by someone at some  
> point  
> b) there is another way to achieve a similar result  
> c) it's a stupid idea, go away
> 
> Thanks,
> 
> Nick

---

<div class="post-metadata">

**Author:** ![Nick\_Sellen](https://avatars.discourse-cdn.com/v4/letter/n/f0a364/32.png) [@Nick\_Sellen](https://discuss.elastic.co/u/Nick_Sellen)\
**Post date:** [December 17, 2011, 12:36am UTC](https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170/3 "2011-12-17T00:36:31Z")

</div>

Ah I missed the parent/child feature - I read  
[http://www.elasticsearch.org/guide/reference/mapping/parent-field.html](http://www.elasticsearch.org/guide/reference/mapping/parent-field.html) but  
it sounded like a way to relate a mapping inside another mapping (rather  
than a document inside another document). Unfortunately though it's the  
wrong way round - when indexing the child document I wouldn't know the  
parent it related to. it's just a blob of content and multiple parent  
documents could refer to it after the child document has been indexed (like  
revisions of the parent document).

I'm not sure it's going to be solvable though - the routing would seem  
tricky as the parent/child relationship is actually many-to-many (revisions  
of a parent all point to the same child, and each parent could have  
multiple children). Perhaps if it was limited to ONE child per parent (like  
in China) the routing could be based on the child id alone, however this  
doesn't seem a good policy.

Another approach would to be able to insert the child content into the  
parent in the river process - it would have to go and fetch the child  
content and pop it into a field in the parent. This would work fine for me  
as I'm more concerned about not storing multiple copies of unchanging child  
content in the backing database rather then re-indexing it when the parent  
changes.

Does the lang-javascript plugin allow me to do things like make an HTTP  
request and pull in more content? (using the "script" field in the couchdb  
river).

---

<div class="post-metadata">

**Author:** ![Nick\_Sellen](https://avatars.discourse-cdn.com/v4/letter/n/f0a364/32.png) [@Nick\_Sellen](https://discuss.elastic.co/u/Nick_Sellen)\
**Post date:** [December 22, 2011, 2:24pm UTC](https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170/4 "2011-12-22T14:24:42Z")

</div>

Another approach that might work for me is if I could change the "id"  
inside the river script - it looks l like it would just involve getting the  
"id" from the ctx variable after the script has run in this bit of code  
inside the couchdb river: [https://gist.github.com/1510453](https://gist.github.com/1510453).

Would you be interested in having that change?

---

<div class="post-metadata">

**Author:** ![Nick\_Sellen](https://avatars.discourse-cdn.com/v4/letter/n/f0a364/32.png) [@Nick\_Sellen](https://discuss.elastic.co/u/Nick_Sellen)\
**Post date:** [January 8, 2012, 2:12pm UTC](https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170/5 "2012-01-08T14:12:47Z")

</div>

I made a version of the couchdb river which lets you update the id inside  
the script. However, ultimately I've concluded using the river causes more  
difficulties than just indexing directly (sending documents to es and  
couchdb at the same time) so I have abandoned it.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:43am UTC](https://discuss.elastic.co/t/querying-other-documents-as-nested-content/6170/6 "2017-07-06T03:43:29Z")

</div>


