# Joins within plugin or script

**URL:** <https://discuss.elastic.co/t/joins-within-plugin-or-script/84619>\
**Category:** Elasticsearch\
**Created:** [May 4, 2017, 7:06pm UTC](https://discuss.elastic.co/t/joins-within-plugin-or-script/84619 "2017-05-04T19:06:22Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![mastayoda](https://avatars.discourse-cdn.com/v4/letter/m/a183cd/32.png) [@mastayoda](https://discuss.elastic.co/u/mastayoda)\
**Post date:** [May 4, 2017, 7:06pm UTC](https://discuss.elastic.co/t/joins-within-plugin-or-script/84619/1 "2017-05-04T19:06:22Z")

</div>

Hello all,

For our project([trueno](https://truenodb.github.io/documentation/latest/)) which is a Graph Database, we need to do Joins in order to traverse the graph. We are using **Elasticsearch** as our distributed storage engine.

Basically, we have done many tweaks to get pretty good performance, such as TransportClient instead of rest, and pure WebSockets to extract and index Vertices and Edges from the storage.

We have been evaluating how to better do traversals. We where using the [SIREn Join plugin](https://github.com/sirensolutions/siren-join), but it is pretty slow and can only do 2 step joins. We are very aware that Join operations in a distributed system are very expensive but we need to do the traversal anyways. Our options seems to be the following:

- Application Side Joins(too many round trips, therefore slow)
- Use the parent or nested features to store the vertices and edges(If graph is fully connected, it wont be scalable since all documents needs to reside in the same shard).

Is there a way to use the **Script** or **Plugin** interfaces to do something like the following?:

1- The user define the graph traversal query.  
2- Sends the query to a plugin or script.  
3- The plugin or script optimizes the traversal.  
4- Within the plugin or script, the search is performed internally by steps, example:  
vertex(1)-\>neighbors()-\>neighbors(where a\>5)

**Explanation** : Start from vertex with id 1. Get it's neighbors, then get those neighbor's neighbors where property a \> 5

5- Return the whole set of documents resulting from the traversal to the client.

The aims from this are:

1- Minimize roundtrips.  
2- Optimize search within ElasticSearch and not from the client side.

We have the following structure in the storage.

- Every graph is an **Index**
- Every index has two types (vertices and edges)

Any suggestion on how to approach this in a performant way? Also, is possible to invoke an plugin endpoint from the TransportClient?

Thanks in advanced for your help, we really appreciate it.

PD: We have been all over the latest documentation but we where unable to find answers to these questions.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [May 4, 2017, 7:24pm UTC](https://discuss.elastic.co/t/joins-within-plugin-or-script/84619/2 "2017-05-04T19:24:18Z")

</div>

> [@mastayoda](#):
>
> such as TransportClient instead of rest

We've been slowly replacing TransportClient because it causes a ton of coupling with the internals of Elasticsearch. We've been replacing it with a REST based client.

> [@mastayoda](#):
>
> Is there a way to use the Script or Plugin interfaces to do something like the following?:

Scripts have to be synchronous and network communication for the join wouldn't be. Plugins can do what they like, though the Lucene API for queries is synchronous so you wouldn't want to use that without some major reworking.

---

<div class="post-metadata">

**Author:** ![mastayoda](https://avatars.discourse-cdn.com/v4/letter/m/a183cd/32.png) [@mastayoda](https://discuss.elastic.co/u/mastayoda)\
**Post date:** [May 4, 2017, 7:31pm UTC](https://discuss.elastic.co/t/joins-within-plugin-or-script/84619/3 "2017-05-04T19:31:00Z")

</div>

> [@nik9000](#):
>
> We've been slowly replacing TransportClient because it causes a ton of coupling with the internals of Elasticsearch. We've been replacing it with a REST based client.

I understand... Why not provide a webSocket(RFC 6455) also? REST calls are extremely slow and have a lot of overhead. WebSockets are available for pretty much every language and platform. We started with REST, but moving to TransportClient got us a 10X speedup.

> [@nik9000](#):
>
> Scripts have to be synchronous and network communication for the join wouldn't be. Plugins can do what they like, though the Lucene API for queries is synchronous so you wouldn't want to use that without some major reworking.

I see. Technically we can use threads within the plugin to access the index in parallel, right? Is not there an interface which lets Elasticsearch do the work? Such an internal interface which accepts queries. I'm afraid we would need to get into shard routings etc.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 1, 2017, 7:44pm UTC](https://discuss.elastic.co/t/joins-within-plugin-or-script/84619/4 "2017-06-01T19:44:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
