# Question about best practices - should I create a separate index when objects can within a document be in the thousands

**URL:** <https://discuss.elastic.co/t/question-about-best-practices-should-i-create-a-separate-index-when-objects-can-within-a-document-be-in-the-thousands/133527>\
**Category:** Elasticsearch\
**Created:** [May 28, 2018, 11:32am UTC](https://discuss.elastic.co/t/question-about-best-practices-should-i-create-a-separate-index-when-objects-can-within-a-document-be-in-the-thousands/133527 "2018-05-28T11:32:32Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dave\_Clissold](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dave_clissold/32/30296_2.png) [@Dave\_Clissold](https://discuss.elastic.co/u/Dave_Clissold)\
**Post date:** [May 28, 2018, 11:32am UTC](https://discuss.elastic.co/t/question-about-best-practices-should-i-create-a-separate-index-when-objects-can-within-a-document-be-in-the-thousands/133527/1 "2018-05-28T11:32:33Z")

</div>

I am working with a combination of Neo4J and ES in a recommendation engine, I have an algorithm that generates a score between user and product, in Neo4J. I also want to add 3 similar product objects , similar1, similar2, similar3, which can contain up to a 30k items.

What is the best way of storing the score and the similar products in ES?

1 - Create a an object within the product doc with the user\_id as the key and the score as the value like this  
{  
"productName":'product1",  
"similar1": ["product\_id\_1","product\_id\_2","product\_id\_3"],  
"similar2": ["product\_id\_4","product\_id\_5","product\_id\_6"]  
"user\_scores:  
{"80cc5fe7-1110-44b1-ae74-51511008f5f2": 86},  
{"47dc1f69-c9bd-448a-ae22-1ff288c82fe4": 36}  
}

2 - Create a a separate index for each user with scores besides each product and again separate indices for each similar set where the doc \_id matches that of the product.

My concern with the first is that when the application goes starts to scale if I reach a couple of million users, the documents could end up reaching the 2gb Lucene size limit.

With the 2nd option and the similar product objects from option 1, how would I build a query that could return similar product data from the multiple indices as child objects within the product doc and score the product doc by the user score index

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 25, 2018, 11:32am UTC](https://discuss.elastic.co/t/question-about-best-practices-should-i-create-a-separate-index-when-objects-can-within-a-document-be-in-the-thousands/133527/2 "2018-06-25T11:32:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
