# Index management

**URL:** <https://discuss.elastic.co/t/index-management/254724>\
**Category:** Elasticsearch\
**Created:** [November 9, 2020, 9:08am UTC](https://discuss.elastic.co/t/index-management/254724 "2020-11-09T09:08:49Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![antoteo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/antoteo/32/88393_2.png) [@antoteo](https://discuss.elastic.co/u/antoteo)\
**Post date:** [November 9, 2020, 9:08am UTC](https://discuss.elastic.co/t/index-management/254724/1 "2020-11-09T09:08:49Z")

</div>

Hello i have a case that people will upload their dataset excels csvs and other with python. their data will have different types, i want them to be able to search from all data. I know create one index in every upload but thousand will be created. In the Index patterns i use \* and i combine them all.  
Is it better to have one large index or thousands of smalls???

I will create i cluster in production.  
2 vms with CPU: 16 CPU cores  
Memory: 32GB RAM  
Storage : 1 TB SSD with minimum 3k dedicated IOPS

will be enough?  
thanks

---

<div class="post-metadata">

**Author:** ![ylasri](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ylasri/32/86120_2.png) [@ylasri](https://discuss.elastic.co/u/ylasri)\
**Post date:** [November 9, 2020, 9:14am UTC](https://discuss.elastic.co/t/index-management/254724/2 "2020-11-09T09:14:15Z")

</div>

All depend on how much data will be loaded on a daily basic and the retention period !  
Also it's not recommanded that shard exceed 50Gb of size

---

<div class="post-metadata">

**Author:** ![antoteo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/antoteo/32/88393_2.png) [@antoteo](https://discuss.elastic.co/u/antoteo)\
**Post date:** [November 9, 2020, 9:23am UTC](https://discuss.elastic.co/t/index-management/254724/3 "2020-11-09T09:23:59Z")

</div>

lets suppose that each day 5 small excels will be uploaded so 5 indeces every day will be created but i will not erase them until a few years.

If i use one index per file with upload-date\*  
i will have thousands of indexs  
if i use one big it will become more than 50G

what is the best approach?

another issue is the types, types will become too much. Should i tell them to upload files with same column's which means same types?

thanks

---

<div class="post-metadata">

**Author:** ![ylasri](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ylasri/32/86120_2.png) [@ylasri](https://discuss.elastic.co/u/ylasri)\
**Post date:** [November 9, 2020, 9:35am UTC](https://discuss.elastic.co/t/index-management/254724/4 "2020-11-09T09:35:46Z")

</div>

There is [ILM](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-lifecycle-management.html) to take care of the index size, you can ask to rollover automtically when 50Gb exceeded  
As much as possible if you can normalize types that would be better

---

<div class="post-metadata">

**Author:** ![antoteo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/antoteo/32/88393_2.png) [@antoteo](https://discuss.elastic.co/u/antoteo)\
**Post date:** [November 9, 2020, 12:54pm UTC](https://discuss.elastic.co/t/index-management/254724/5 "2020-11-09T12:54:21Z")

</div>

thank you for your quick response

i read about ilm but this means i have to define template for index in advance but i dont know the fields for my indexed documents these would be generated dynamic when people add documents.  
do i get it well?

---

<div class="post-metadata">

**Author:** ![ylasri](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ylasri/32/86120_2.png) [@ylasri](https://discuss.elastic.co/u/ylasri)\
**Post date:** [November 9, 2020, 12:57pm UTC](https://discuss.elastic.co/t/index-management/254724/6 "2020-11-09T12:57:34Z")

</div>

The template mapping may not be restrictive on new fields  
you can define some fields, some dynamic mapping ... etc, or you can have a template without mapping just to force ILM policy

---

<div class="post-metadata">

**Author:** ![antoteo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/antoteo/32/88393_2.png) [@antoteo](https://discuss.elastic.co/u/antoteo)\
**Post date:** [December 1, 2020, 6:47am UTC](https://discuss.elastic.co/t/index-management/254724/7 "2020-12-01T06:47:38Z")

</div>

Of course i will suggest that there will be a normalization on fields so they use the same name in same fields and not create more. Are there going to be performance issue on search if we reach more than 1000. can i use \_source field and put them all there besides some id fields? is this gone help? My index pattern says i already have 108 fields but i am still in the begging.

---

<div class="post-metadata">

**Author:** ![antoteo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/antoteo/32/88393_2.png) [@antoteo](https://discuss.elastic.co/u/antoteo)\
**Post date:** [December 2, 2020, 12:00pm UTC](https://discuss.elastic.co/t/index-management/254724/8 "2020-12-02T12:00:38Z")

</div>

I learned about flattened fields, but to map my dynamically created fields in the flattened field?

thanks

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 30, 2020, 12:00pm UTC](https://discuss.elastic.co/t/index-management/254724/9 "2020-12-30T12:00:44Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
