# Index design for user's activities

**URL:** <https://discuss.elastic.co/t/index-design-for-users-activities/23196>\
**Category:** Elasticsearch\
**Created:** [April 11, 2015, 5:40pm UTC](https://discuss.elastic.co/t/index-design-for-users-activities/23196 "2015-04-11T17:40:09Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chen\_Wang](https://avatars.discourse-cdn.com/v4/letter/c/b5a626/32.png) [@Chen\_Wang](https://discuss.elastic.co/u/Chen_Wang)\
**Post date:** [April 11, 2015, 5:40pm UTC](https://discuss.elastic.co/t/index-design-for-users-activities/23196/1 "2015-04-11T17:40:09Z")

</div>

I am maintaining a years of user's activity including browse, purchase  
data. Each entry in browse/purchase is a json object:{item\_id: id1,  
item\_name, name1, category: c1, brand:b1, event\_time: t1} .

I would like to compose different queries such like getting all customers  
who browsed item A, and or purchased item B within time range t1 to t2.  
There are tens of millions customers.

My current design is to use nested object for each customer:  
customer1:  
customer\_id,id1,  
name: name1,  
country: US,  
browse: [{browseentry1\_json},{browseentry2\_json},...],  
purchase: [{purchase entry1\_json},{purchase entry2\_json},...]

With this design, I can easily compose all kinds of queries with nested  
query. The only problem is that it is hard to expire older browse/purchase  
data: I only wanna keep, for example, a years of browse/purchase data. In  
this design, I will have to at some point, read the entire index out,  
delete the expired browse/purchase data, and write them back.

Another design is to use parent/child structure.  
type: user is the parent of type browse and purchase.  
type browse will contain each browse entry.  
Although deleting old data seems easier with delete by query, for the  
above query, I will have to do multiple and/or has\_child queries,and it  
would be much less performant. In fact, initially i was using parent/child  
structure, but the query time seemed really long. I thus gave it up and  
tried to switch to nested object.

I am also thinking about using nested object, but break the data into  
different index(like monthly index) so that I can easily expire old data.  
The problem with this approach is that I have to query across those  
multiple indexes, and do aggregation on that to get the distinct users,  
which I assume will be much slower.(havn't tried yet). One requirement of  
this project is to be able to give the count of the queries in acceptable  
time frame.(like seconds) and I am afraid this approach may not be  
acceptable.

The ES cluster is 7 machines, each 8 cores and 32G memory.  
Any suggestions?

Thanks in advance!  
Chen

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/e1279e50-4ec7-4292-8ef3-49bc187498c1%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e1279e50-4ec7-4292-8ef3-49bc187498c1%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:20am UTC](https://discuss.elastic.co/t/index-design-for-users-activities/23196/2 "2017-07-06T00:20:14Z")

</div>


