# Best method of handling arbitrary document joins

**URL:** <https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457>\
**Category:** Elasticsearch\
**Created:** [May 3, 2018, 1:31pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457 "2018-05-03T13:31:05Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Erik\_Miller](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/erik_miller/32/14913_2.png) [@Erik\_Miller](https://discuss.elastic.co/u/Erik_Miller)\
**Post date:** [May 3, 2018, 1:31pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/1 "2018-05-03T13:31:05Z")

</div>

I've done some research and everywhere I look it sounds like nosql dbs, including elastic search use application side joins. For ES I understand the nested/parent-child options and have both implemented in my existing stack. But these are limited in functionality to a traditional join.

I'm wondering if there is a good solution to handling joins in a generic sense. For example, I have two document types with no nested/parent-child relationship set up in ES. But they do have ID's that can be used to join them. Is there a plugin/workflow/solution to do this other than a custom application side join?

I understand this isn't something ES is intended to handle and "this is what a relation DB is for".

I'm wondering is there is another pattern/plugin/option to do arbitrary document joins from ES?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 3, 2018, 2:42pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/2 "2018-05-03T14:42:41Z")

</div>

See the [entity centric indexing](https://www.youtube.com/watch?v=yBf7oeJKH2Y) approach which shifts the join compute cost to index-time rather than query-time.

---

<div class="post-metadata">

**Author:** ![darkmoon](https://avatars.discourse-cdn.com/v4/letter/d/ecd19e/32.png) [@darkmoon](https://discuss.elastic.co/u/darkmoon)\
**Post date:** [May 3, 2018, 2:55pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/3 "2018-05-03T14:55:44Z")

</div>

Is there a whitepaper or something we can look at? Hopping around a video to find the info you're looking for just doesn't work well.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 3, 2018, 3:26pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/4 "2018-05-03T15:26:30Z")

</div>

There's example data and scripts here: [http://bit.ly/entcent\_painless](http://bit.ly/entcent_painless)

---

<div class="post-metadata">

**Author:** ![darkmoon](https://avatars.discourse-cdn.com/v4/letter/d/ecd19e/32.png) [@darkmoon](https://discuss.elastic.co/u/darkmoon)\
**Post date:** [May 3, 2018, 3:53pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/5 "2018-05-03T15:53:14Z")

</div>

Let me try again.

Please define 'entity centric' and what it means to database structure.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 3, 2018, 4:01pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/6 "2018-05-03T16:01:47Z")

</div>

Many elasticsearch systems capture "events" - a timestamped record of some activity. These indices are "event-centric" - one document per event and often organised into time-based indices e.g. an index per day or week or month.  
Events are generated by "entities" - nouns in the real-world such as people, ip addresses, cars etc. Each entity typically generates many events. Anyone attempting to analyse the behaviour of entities (e.g. length of time spent on a website, people-who-bought-x-also-bought-?) typically have a hard time doing this on an event-centric index where the data is not centered around an entity. An entity-centric index brings a summary of an entity's activity into a single document.

---

<div class="post-metadata">

**Author:** ![Erik\_Miller](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/erik_miller/32/14913_2.png) [@Erik\_Miller](https://discuss.elastic.co/u/Erik_Miller)\
**Post date:** [May 3, 2018, 5:12pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/7 "2018-05-03T17:12:28Z")

</div>

My index is actually entity centric to begin with. My specific use case is the ability to perform custom queries against a complex document type, and then another query against a different document type and find the intersection.

The queries used are dynamic in nature so we can't tune an additional join document as each query changes frequently.

---

<div class="post-metadata">

**Author:** ![darkmoon](https://avatars.discourse-cdn.com/v4/letter/d/ecd19e/32.png) [@darkmoon](https://discuss.elastic.co/u/darkmoon)\
**Post date:** [May 3, 2018, 7:29pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/8 "2018-05-03T19:29:52Z")

</div>

We've been experimenting with this, under a different name, as things pop  
up where it would be useful. We're trying to concentrate on flows, not  
exports, so when we get an event that mentions an ip, it not only gets  
indexed into the log index, we also perform an update to an IP state  
table. That way we have the history (of all IPs) in the event logs, and  
the current (last obseved) state (of all IPs) in the state table. We're  
growing more state tables as more entity types become needed.

Thank you.

---

<div class="post-metadata">

**Author:** ![ecc256](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ecc256/32/59815_2.png) [@ecc256](https://discuss.elastic.co/u/ecc256)\
**Post date:** [May 4, 2018, 9:31pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/9 "2018-05-04T21:31:53Z")

</div>

Guys,  
I do have somewhat similar question.  
We do have set of servers behind load balancer.  
There are several event streams like web, error and performance logs.  
Event across streams can be linked/joined by time interval and server name.  
What would be the right way to do it?  
I.e. if we see CPU spike on a server, we want to know what web requests were executed and if there are any errors around this time.  
Seems like filtering all streams by server name and time interval should do it?  
So the only requirement is server name field should be named exactly the same in all event streams, right?

I didn’t mean to hijack thread. If it not related, I can repost it as separate question.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 1, 2018, 9:31pm UTC](https://discuss.elastic.co/t/best-method-of-handling-arbitrary-document-joins/130457/10 "2018-06-01T21:31:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
