# Large number of indexes for multi-tenant product

**URL:** <https://discuss.elastic.co/t/large-number-of-indexes-for-multi-tenant-product/22699>\
**Category:** Elasticsearch\
**Created:** [March 17, 2015, 2:11am UTC](https://discuss.elastic.co/t/large-number-of-indexes-for-multi-tenant-product/22699 "2015-03-17T02:11:01Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Richard\_Blaylock](https://avatars.discourse-cdn.com/v4/letter/r/a183cd/32.png) [@Richard\_Blaylock](https://discuss.elastic.co/u/Richard_Blaylock)\
**Post date:** [March 17, 2015, 2:11am UTC](https://discuss.elastic.co/t/large-number-of-indexes-for-multi-tenant-product/22699/1 "2015-03-17T02:11:01Z")

</div>

Hi all,

We have a multi-tenant product and are leaning towards dynamically creating  
(and deleting) various indexes relevant to a tenant at runtime: as a tenant  
is created, so are that tenant's indexes. When a tenant is deleted so are  
that tenant's indexes. Each index is specific to that tenant and could  
vary in size, but we do not expect any given index to ever be larger than a  
single disk (e.g. 80 GB).

Due to index shard issues (static, too many shards per index = a hit on  
performance (more map/reduce work to do), etc.), and due to the nature of  
our application, we are currently opting for a single-shard-per-index model

- each index will have one and only one shard. We will have replicas for  
fault tolerance.

On the surface, this appears to be an ideal design choice for multi-tenant  
applications: for any given index, one and only one shard will be 'hit' -  
no need to search across multiple shards, ever. It also reduces contention  
because indexes are always tenant-specific: as an index becomes larger, any  
slowness due to the large index _only_ impacts the corresponding tenant  
(customer), whereas the alternative - using one index across tenants - one  
tenant's use/load could negatively impact other tenants' query performance.

So for multi-tenancy, this single-shard-per-index model sounds ideal for  
our use case - the _only_ issue here is that the number of indexes  
increases dramatically as the number of tenants (customers) increases.  
Consider a system with 20,000 tenants, each having (potentially) hundreds  
or thousands, or even 10s of thousands of indexes, resulting in millions of  
indexes overall. This is manageable from our product's perspective, but  
what impact would this have on ElasticSearch, if any?

Are there practical limits? IIUC, there is a Lucene index (file) per shard,  
so if there are hundreds of thousands or millions of Lucene indexes/files -  
other than disk space and file descriptor count per ES node, are there any  
other limits? Does performance degrade as the number of  
single-shard-indexes increases? Or is there no problem at all?

Thanks,  
Richard

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 17, 2015, 9:11pm UTC](https://discuss.elastic.co/t/large-number-of-indexes-for-multi-tenant-product/22699/2 "2015-03-17T21:11:30Z")

</div>

There are practical limits, based on your dataset, node sizing, version etc.

You'd be better off segregating indices by a higher level definition (eg  
customer number, 1-999, 1000-1999 etc), using routing and then aliases on  
top. This way you conceptually get the same layout as a single index per  
customer, but it gives you the option to split larger customers out to  
their own index and without wasting resources on small use customers.

On 16 March 2015 at 19:11, Richard Blaylock [richard@stormpath.com](mailto:richard@stormpath.com) wrote:

> Hi all,
> 
> We have a multi-tenant product and are leaning towards dynamically  
> creating (and deleting) various indexes relevant to a tenant at runtime: as  
> a tenant is created, so are that tenant's indexes. When a tenant is  
> deleted so are that tenant's indexes. Each index is specific to that  
> tenant and could vary in size, but we do not expect any given index to ever  
> be larger than a single disk (e.g. 80 GB).
> 
> Due to index shard issues (static, too many shards per index = a hit on  
> performance (more map/reduce work to do), etc.), and due to the nature of  
> our application, we are currently opting for a single-shard-per-index model
> 
> - each index will have one and only one shard. We will have replicas for  
> fault tolerance.
> 
> On the surface, this appears to be an ideal design choice for multi-tenant  
> applications: for any given index, one and only one shard will be 'hit' -  
> no need to search across multiple shards, ever. It also reduces contention  
> because indexes are always tenant-specific: as an index becomes larger, any  
> slowness due to the large index _only_ impacts the corresponding tenant  
> (customer), whereas the alternative - using one index across tenants - one  
> tenant's use/load could negatively impact other tenants' query performance.
> 
> So for multi-tenancy, this single-shard-per-index model sounds ideal for  
> our use case - the _only_ issue here is that the number of indexes  
> increases dramatically as the number of tenants (customers) increases.  
> Consider a system with 20,000 tenants, each having (potentially) hundreds  
> or thousands, or even 10s of thousands of indexes, resulting in millions of  
> indexes overall. This is manageable from our product's perspective, but  
> what impact would this have on Elasticsearch, if any?
> 
> Are there practical limits? IIUC, there is a Lucene index (file) per  
> shard, so if there are hundreds of thousands or millions of Lucene  
> indexes/files - other than disk space and file descriptor count per ES  
> node, are there any other limits? Does performance degrade as the number  
> of single-shard-indexes increases? Or is there no problem at all?
> 
> Thanks,  
> Richard
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAEYi1X\_UA2XX8M8bDCCS%2Bx4p9Ta5-nk1vj45pLh9JDSePY0AGQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAEYi1X_UA2XX8M8bDCCS%2Bx4p9Ta5-nk1vj45pLh9JDSePY0AGQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 17, 2015, 10:35pm UTC](https://discuss.elastic.co/t/large-number-of-indexes-for-multi-tenant-product/22699/3 "2015-03-17T22:35:46Z")

</div>

This is a super timely blog from the Found crew -

> **[Discovering the Need for an Indexing Strategy in Multi-Tenant Applications](https://www.elastic.co/blog/found-multi-tenancy)**
>
> There are many buzzwords that may be applied to Elasticsearch. Multi-tenancy is one. Getting started with a multi-tenant use case can be deceptively easy - there are some pitfalls that will require a little careful design.

On 17 March 2015 at 14:11, Mark Walkom [markwalkom@gmail.com](mailto:markwalkom@gmail.com) wrote:

> There are practical limits, based on your dataset, node sizing, version  
> etc.
> 
> You'd be better off segregating indices by a higher level definition (eg  
> customer number, 1-999, 1000-1999 etc), using routing and then aliases on  
> top. This way you conceptually get the same layout as a single index per  
> customer, but it gives you the option to split larger customers out to  
> their own index and without wasting resources on small use customers.
> 
> On 16 March 2015 at 19:11, Richard Blaylock [richard@stormpath.com](mailto:richard@stormpath.com) wrote:
> 
> > Hi all,
> > 
> > We have a multi-tenant product and are leaning towards dynamically  
> > creating (and deleting) various indexes relevant to a tenant at runtime: as  
> > a tenant is created, so are that tenant's indexes. When a tenant is  
> > deleted so are that tenant's indexes. Each index is specific to that  
> > tenant and could vary in size, but we do not expect any given index to ever  
> > be larger than a single disk (e.g. 80 GB).
> > 
> > Due to index shard issues (static, too many shards per index = a hit on  
> > performance (more map/reduce work to do), etc.), and due to the nature of  
> > our application, we are currently opting for a single-shard-per-index model
> > 
> > - each index will have one and only one shard. We will have replicas for  
> > fault tolerance.
> > 
> > On the surface, this appears to be an ideal design choice for  
> > multi-tenant applications: for any given index, one and only one shard will  
> > be 'hit' - no need to search across multiple shards, ever. It also reduces  
> > contention because indexes are always tenant-specific: as an index becomes  
> > larger, any slowness due to the large index _only_ impacts the  
> > corresponding tenant (customer), whereas the alternative - using one index  
> > across tenants - one tenant's use/load could negatively impact other  
> > tenants' query performance.
> > 
> > So for multi-tenancy, this single-shard-per-index model sounds ideal for  
> > our use case - the _only_ issue here is that the number of indexes  
> > increases dramatically as the number of tenants (customers) increases.  
> > Consider a system with 20,000 tenants, each having (potentially) hundreds  
> > or thousands, or even 10s of thousands of indexes, resulting in millions of  
> > indexes overall. This is manageable from our product's perspective, but  
> > what impact would this have on Elasticsearch, if any?
> > 
> > Are there practical limits? IIUC, there is a Lucene index (file) per  
> > shard, so if there are hundreds of thousands or millions of Lucene  
> > indexes/files - other than disk space and file descriptor count per ES  
> > node, are there any other limits? Does performance degrade as the number  
> > of single-shard-indexes increases? Or is there no problem at all?
> > 
> > Thanks,  
> > Richard
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/607f62c1-5854-43e0-9d25-3f748aca44a4%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAEYi1X99jYR7a%2BYuf3o-C\_bxE5OvxybTAKr2rQL4HEEDqS0R6Q%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAEYi1X99jYR7a%2BYuf3o-C_bxE5OvxybTAKr2rQL4HEEDqS0R6Q%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Isaias\_Barroso](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/isaias_barroso/32/3980_2.png) [@Isaias\_Barroso](https://discuss.elastic.co/u/Isaias_Barroso)\
**Post date:** [August 2, 2015, 2:40pm UTC](https://discuss.elastic.co/t/large-number-of-indexes-for-multi-tenant-product/22699/4 "2015-08-02T14:40:14Z")

</div>

Hi @Richard_Blaylock, what approach have you used to solve your problem? @warkolm reference to Found blog post is very nice. I'm facing the same problem. Another interesting point to separate the tenants is to isolate Tenants from eventual index corruption, initially my solution was using a TenantId as a discriminator, I have an another factor that can increase the Indexes numbers, my application can index in multiple languages and the recommendation is to separate the index by language.

Best regards

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:57pm UTC](https://discuss.elastic.co/t/large-number-of-indexes-for-multi-tenant-product/22699/5 "2017-07-05T23:57:51Z")

</div>


