# Opinions on ES with highly related data

**URL:** <https://discuss.elastic.co/t/opinions-on-es-with-highly-related-data/10761>\
**Category:** Elasticsearch\
**Created:** [February 16, 2013, 2:47am UTC](https://discuss.elastic.co/t/opinions-on-es-with-highly-related-data/10761 "2013-02-16T02:47:55Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![knacktus](https://avatars.discourse-cdn.com/v4/letter/k/a5b964/32.png) [@knacktus](https://discuss.elastic.co/u/knacktus)\
**Post date:** [February 16, 2013, 2:47am UTC](https://discuss.elastic.co/t/opinions-on-es-with-highly-related-data/10761/1 "2013-02-16T02:47:55Z")

</div>

Hi guys,

I've just discovered the potential of ES as scalable multi-purpose cache or  
even only data store. So far, I've been using RDBMS with MemcacheD or Redis  
(for simple queries in the application layer). I've decided to give ES a  
try by building a Prototype, but before I dive in I'd much appreciate your  
opinions about how I plan to get the data from ES.

The issue might be that my data is highly related and I need to work mainly  
with large structures. ES's main task in this regard would be to support a  
server process collecting all data items which are within the large  
structures. These items are send to a rich client, where the actual  
structured views are build.

Here's some example data:

root = {id = 1,  
name = "Plane",  
subassemblies = [2, 3, 4]}

body = {id = 2,  
name = "Body",  
subassemblies = [5, 6]}

left\_wing = {id = 3,  
name = "Wing"  
subassemlies = []}

right\_wing = {id = 4,  
name = "Wing"  
subassemlies = []}

uppder\_body\_structure = {id = 5,  
name = "Upper Body"  
subassemlies = []}

lower\_body\_structure = {id = 6,  
name = "Lower Body"  
subassemlies = []}

So, I would query ES iteratively to get all items, starting with the root  
item. About like this in Python pseudocode:

all\_item\_ids = []  
current\_root\_id = 1  
all\_item\_ids.append(current\_root\_id)  
current\_item\_ids = [1]

while len(current\_item\_ids) \> 0:  
current\_item\_ids =  
query\_ES\_for\_items\_by\_given\_ids\_and\_return\_given\_field(current\_item\_ids,  
"subassemlies") # here would come some more advanced query options  
all\_item\_ids.extend(current\_item\_ids)

send\_ids\_to\_client(all\_item\_ids) # there's a client cache for the item  
data, so I send the ids only

The amount of data is quite large. Up to 100000 rows with up to 50 levels.  
So I would possibly end up with queries with 10000 arguments (however only  
exact matches need to be considered). Those could be split up into batches,  
but that's where I hope to get your opinions. (Hitting ES 50 times wouldn't  
make me nervous, but when it comes to a couple of thousand times, something  
seems not right. But then, if it took only a couple of seconds overall, I  
wouldn't complain :-)).

Is this the right approach to handle large structures? Do you see any  
general showstoppers or flaws? (Like limits in query-size ...)

Another question is about storing Thrift oder Protocol Buffers encoded  
data. How would you store those for simple get, mget operations? (Those  
formats are used for transport and in the client cache, which is basically  
a key-value store.)

On top of that I would use fulltext search and general combinded searches  
within the whole data. But I have no doubt that ES is the right choice  
there. So, if I'd be able to retrieve the structured data in a performant  
way, ES would be an awesome powerfull all-in-one solution.

Cheers and thanks for any comments and opinions,

Jan

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![phill](https://avatars.discourse-cdn.com/v4/letter/p/779978/32.png) [@phill](https://discuss.elastic.co/u/phill)\
**Post date:** [February 20, 2013, 3:55pm UTC](https://discuss.elastic.co/t/opinions-on-es-with-highly-related-data/10761/2 "2013-02-20T15:55:47Z")

</div>

ES is for storing _documents_ it appears what you are trying to do is  
store a hierarchy as you would represent it in a RDMS. Try to build the  
document that you want to retrieve, because that is what you will get,  
one or more documents. I doubt you ever want the query to return just  
body ={id =2,  
name ="Body",  
subassemblies =[5,6]}

-Paul

On 2/15/2013 6:47 PM, [knacktus@googlemail.com](mailto:knacktus@googlemail.com) wrote:

> Hi guys,
> 
> I've just discovered the potential of ES as scalable multi-purpose  
> cache or even only data store. So far, I've been using RDBMS with  
> MemcacheD or Redis (for simple queries in the application layer). I've  
> decided to give ES a try by building a Prototype, but before I dive in  
> I'd much appreciate your opinions about how I plan to get the data  
> from ES.
> 
> The issue might be that my data is highly related and I need to work  
> mainly with large structures. ES's main task in this regard would be  
> to support a server process collecting all data items which are within  
> the large structures. These items are send to a rich client, where the  
> actual structured views are build.
> 
> Here's some example data:
> 
> ||  
> root ={id =1,  
> name ="Plane",  
> subassemblies =[2,3,4]}
> 
> body ={id =2,  
> name ="Body",  
> subassemblies =[5,6]}
> 
> left\_wing ={id =3,  
> name ="Wing"  
> subassemlies =}
> 
> |right\_wing ={id =4,  
> name ="Wing"  
> subassemlies =}
> 
> ||uppder\_body\_structure ={id =5,  
> name ="Upper Body"  
> subassemlies =}
> 
> |||lower\_body\_structure ={id =6,  
> name ="Lower Body"  
> subassemlies =}||
> 
> So, I would query ES iteratively to get all items, starting with the  
> root item. About like this in Python pseudocode:
> 
> ||  
> all\_item\_ids =   
> current\_root\_id = 1  
> all\_item\_ids.append(current\_root\_id)  
> current\_item\_ids = [1]
> 
> while len(current\_item\_ids) \> 0:  
> current\_item\_ids =  
> query\_ES\_for\_items\_by\_given\_ids\_and\_return\_given\_field(current\_item\_ids,  
> "subassemlies") # here would come some more advanced query options  
> all\_item\_ids.extend(current\_item\_ids)
> 
> send\_ids\_to\_client(all\_item\_ids) # there's a client cache for the item  
> data, so I send the ids only
> 
> The amount of data is quite large. Up to 100000 rows with up to 50  
> levels. So I would possibly end up with queries with 10000 arguments  
> (however only exact matches need to be considered). Those could be  
> split up into batches, but that's where I hope to get your opinions.  
> (Hitting ES 50 times wouldn't make me nervous, but when it comes to a  
> couple of thousand times, something seems not right. But then, if it  
> took only a couple of seconds overall, I wouldn't complain :-)).
> 
> Is this the right approach to handle large structures? Do you see any  
> general showstoppers or flaws? (Like limits in query-size ...)
> 
> Another question is about storing Thrift oder Protocol Buffers encoded  
> data. How would you store those for simple get, mget operations?  
> (Those formats are used for transport and in the client cache, which  
> is basically a key-value store.)
> 
> On top of that I would use fulltext search and general combinded  
> searches within the whole data. But I have no doubt that ES is the  
> right choice there. So, if I'd be able to retrieve the structured data  
> in a performant way, ES would be an awesome powerfull all-in-one solution.
> 
> Cheers and thanks for any comments and opinions,
> 
> ## Jan
> 
> You received this message because you are subscribed to the Google  
> Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send  
> an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:50am UTC](https://discuss.elastic.co/t/opinions-on-es-with-highly-related-data/10761/3 "2017-07-06T02:50:31Z")

</div>


