# What is best indexing strategy for multitenant data?

**URL:** <https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353>\
**Category:** Elasticsearch\
**Created:** [May 5, 2011, 6:32pm UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353 "2011-05-05T18:32:44Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ellery\_Crane\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ellery_crane_2/32/3369_2.png) [@Ellery\_Crane\_2](https://discuss.elastic.co/u/Ellery_Crane_2)\
**Post date:** [May 5, 2011, 6:32pm UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/1 "2011-05-05T18:32:44Z")

</div>

I'm attempting to integrate elasticsearch into a multitenant web  
application. I have data segmented into tens of thousands of  
'tenants', and then further subdivided by user within a tenant. I'd  
like to make it so that my users can readily access any data within  
their tenant, with optional visibility rules allowing finer grained  
sharing (for instance, sharing certain types of data with other users  
in the tenant, while retaining exclusive access to other types).

Towards this goal, I'm trying to figure out the best way of indexing  
my documents within ES. My initial impulse was to create an index for  
each tenant, but some cursory research indicated this was a Bad Idea.  
Maintaining tens of thousands of indexes while adding more every time  
a new tenant is created is almost certainly untenable. I'm stuck,  
therefore, trying to decide what criteria to use when creating  
indexes. I have a few ideas, mostly centering around heuristic data  
such as geographic location, number of active users and so forth, but  
nothing jumps out as the obviously best course of action. Though,  
regardless of how many indexes I'm running and how I'm determining  
which data to index in each, it seems like routing documents based on  
the tenant id would be ideal for my needs. Can anyone offer some  
advice on what kind of indexing strategy to employ for this type of  
use case?

Some additional information that might be relevant:

- Each tenant/user has the same types of data to index, but there may  
be differences in how each type is mapped. That is, a type might have  
some fields for one user, and others for another, and may need to be  
tokenized/analyzed differently for both. This seems to indicate that  
establishing different indexes based on different type mappings may be  
the way to go, but I doubt there are enough such differences to  
warrant more than a handful of different indexes. Is there any  
performance hit associated with putting vast amounts of data into a  
small number of indexes, assuming a per-tenant id routing strategy?

- Almost all queries will need to be filtered by tenant, by user, or  
by some combination of visibility rules. That said, some users need  
the ability to query across all tenants, but the performance of such  
queries need not be as high.

- I'm using MongoDB as my data store, and see a fairly obvious one-to-  
one mapping of Mongo Collection to ES document type.This suggests that  
using types as a way of dividing data within an index by tenant might  
not work, since I will likely need to use the types for collection  
mapping.

Any advice on this issue is much appreciated.

---

<div class="post-metadata">

**Author:** ![Drew\_H](https://avatars.discourse-cdn.com/v4/letter/d/e79b87/32.png) [@Drew\_H](https://discuss.elastic.co/u/Drew_H)\
**Post date:** [May 5, 2011, 11:39pm UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/2 "2011-05-05T23:39:47Z")

</div>

Hi Ellery,  
I'm new to ES so I'm afraid I don't have answers for you, but I am  
curious what led you to the conclusion that creating an index per  
tenant was a bad idea?

Thanks,  
Drew

On May 5, 2:32 pm, Ellery Crane [seid...@gmail.com](mailto:seid...@gmail.com) wrote:

> I'm attempting to integrate elasticsearch into a multitenant web  
> application. I have data segmented into tens of thousands of  
> 'tenants', and then further subdivided by user within a tenant. I'd  
> like to make it so that my users can readily access any data within  
> their tenant, with optional visibility rules allowing finer grained  
> sharing (for instance, sharing certain types of data with other users  
> in the tenant, while retaining exclusive access to other types).
> 
> Towards this goal, I'm trying to figure out the best way of indexing  
> my documents within ES. My initial impulse was to create an index for  
> each tenant, but some cursory research indicated this was a Bad Idea.  
> Maintaining tens of thousands of indexes while adding more every time  
> a new tenant is created is almost certainly untenable. I'm stuck,  
> therefore, trying to decide what criteria to use when creating  
> indexes. I have a few ideas, mostly centering around heuristic data  
> such as geographic location, number of active users and so forth, but  
> nothing jumps out as the obviously best course of action. Though,  
> regardless of how many indexes I'm running and how I'm determining  
> which data to index in each, it seems like routing documents based on  
> the tenant id would be ideal for my needs. Can anyone offer some  
> advice on what kind of indexing strategy to employ for this type of  
> use case?
> 
> Some additional information that might be relevant:
> 
> - Each tenant/user has the same types of data to index, but there may  
> be differences in how each type is mapped. That is, a type might have  
> some fields for one user, and others for another, and may need to be  
> tokenized/analyzed differently for both. This seems to indicate that  
> establishing different indexes based on different type mappings may be  
> the way to go, but I doubt there are enough such differences to  
> warrant more than a handful of different indexes. Is there any  
> performance hit associated with putting vast amounts of data into a  
> small number of indexes, assuming a per-tenant id routing strategy?
> 
> - Almost all queries will need to be filtered by tenant, by user, or  
> by some combination of visibility rules. That said, some users need  
> the ability to query across all tenants, but the performance of such  
> queries need not be as high.
> 
> - I'm using MongoDB as my data store, and see a fairly obvious one-to-  
> one mapping of Mongo Collection to ES document type.This suggests that  
> using types as a way of dividing data within an index by tenant might  
> not work, since I will likely need to use the types for collection  
> mapping.
> 
> Any advice on this issue is much appreciated.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 6, 2011, 10:00am UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/3 "2011-05-06T10:00:34Z")

</div>

Hi Ellery

> My initial impulse was to create an index for  
> each tenant, but some cursory research indicated this was a Bad Idea.

Yes, I'd agree with that. Each index comes with overhead, so 10,20,100  
maybe more indices would be fine. 10,000 wouldn't.

> Maintaining tens of thousands of indexes while adding more every time  
> a new tenant is created is almost certainly untenable. I'm stuck,  
> therefore, trying to decide what criteria to use when creating  
> indexes. I have a few ideas, mostly centering around heuristic data  
> such as geographic location, number of active users and so forth, but  
> nothing jumps out as the obviously best course of action. Though,  
> regardless of how many indexes I'm running and how I'm determining  
> which data to index in each, it seems like routing documents based on  
> the tenant id would be ideal for my needs. Can anyone offer some  
> advice on what kind of indexing strategy to employ for this type of  
> use case?

I'd say that you should just give each doc that belongs to a particular  
tenant a tenant ID, then you can filter the results based on that. And  
I agree with your idea of using the tenant ID for routing.

> - Each tenant/user has the same types of data to index, but there may  
> be differences in how each type is mapped. That is, a type might have  
> some fields for one user, and others for another, and may need to be  
> tokenized/analyzed differently for both. This seems to indicate that  
> establishing different indexes based on different type mappings may be  
> the way to go, but I doubt there are enough such differences to  
> warrant more than a handful of different indexes.

A few options here. It may be possible to use a single type for all of  
your tenants. For instance:

- if one tenant has fields foo and bar, and another has bar and baz,  
you can store docs from both tenants in the same type, just adding  
the relevant fields

- you mention different analysis. how would this be different? If  
it is a question of language, then you might be able to make  
this work by using the \_analyzer field:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/analyzer-field.html)

- if the mappings are so different that you don't want to combine them  
into one type, then you could use different types within the  
same index, eg user\_v1, user\_v2

> Is there any  
> performance hit associated with putting vast amounts of data into a  
> small number of indexes, assuming a per-tenant id routing strategy?

No, this is something that Elasticsearch is good at. First you can  
play around with the number of shards that you set when you create the  
index. Second, the number of replicas, which can be updated dynamically.  
Third, you can look at using aliases to combine multiple indices (for  
read purposes, to write, you will need to write to one index, or an  
alias that points to only one index).

> - Almost all queries will need to be filtered by tenant, by user, or  
> by some combination of visibility rules. That said, some users need  
> the ability to query across all tenants, but the performance of such  
> queries need not be as high.

Querying across indices is easy, and querying the same index filtering  
by one or many tenant and user IDs is also easy and fast.

hth

clint

---

<div class="post-metadata">

**Author:** ![Alexandre\_Heimburger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexandre_heimburger/32/2941_2.png) [@Alexandre\_Heimburger](https://discuss.elastic.co/u/Alexandre_Heimburger)\
**Post date:** [May 6, 2011, 10:04am UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/4 "2011-05-06T10:04:04Z")

</div>

hi

I also work for a multi-tenant product and we have built one index per  
tenant.

Then we have indices for users, spaces, content etc..

On Fri, May 6, 2011 at 1:39 AM, Drew H [hite.drew@gmail.com](mailto:hite.drew@gmail.com) wrote:

> Hi Ellery,  
> I'm new to ES so I'm afraid I don't have answers for you, but I am  
> curious what led you to the conclusion that creating an index per  
> tenant was a bad idea?
> 
> Thanks,  
> Drew
> 
> On May 5, 2:32 pm, Ellery Crane [seid...@gmail.com](mailto:seid...@gmail.com) wrote:
> 
> > I'm attempting to integrate elasticsearch into a multitenant web  
> > application. I have data segmented into tens of thousands of  
> > 'tenants', and then further subdivided by user within a tenant. I'd  
> > like to make it so that my users can readily access any data within  
> > their tenant, with optional visibility rules allowing finer grained  
> > sharing (for instance, sharing certain types of data with other users  
> > in the tenant, while retaining exclusive access to other types).
> > 
> > Towards this goal, I'm trying to figure out the best way of indexing  
> > my documents within ES. My initial impulse was to create an index for  
> > each tenant, but some cursory research indicated this was a Bad Idea.  
> > Maintaining tens of thousands of indexes while adding more every time  
> > a new tenant is created is almost certainly untenable. I'm stuck,  
> > therefore, trying to decide what criteria to use when creating  
> > indexes. I have a few ideas, mostly centering around heuristic data  
> > such as geographic location, number of active users and so forth, but  
> > nothing jumps out as the obviously best course of action. Though,  
> > regardless of how many indexes I'm running and how I'm determining  
> > which data to index in each, it seems like routing documents based on  
> > the tenant id would be ideal for my needs. Can anyone offer some  
> > advice on what kind of indexing strategy to employ for this type of  
> > use case?
> > 
> > Some additional information that might be relevant:
> > 
> > - Each tenant/user has the same types of data to index, but there may  
> > be differences in how each type is mapped. That is, a type might have  
> > some fields for one user, and others for another, and may need to be  
> > tokenized/analyzed differently for both. This seems to indicate that  
> > establishing different indexes based on different type mappings may be  
> > the way to go, but I doubt there are enough such differences to  
> > warrant more than a handful of different indexes. Is there any  
> > performance hit associated with putting vast amounts of data into a  
> > small number of indexes, assuming a per-tenant id routing strategy?
> > 
> > - Almost all queries will need to be filtered by tenant, by user, or  
> > by some combination of visibility rules. That said, some users need  
> > the ability to query across all tenants, but the performance of such  
> > queries need not be as high.
> > 
> > - I'm using MongoDB as my data store, and see a fairly obvious one-to-  
> > one mapping of Mongo Collection to ES document type.This suggests that  
> > using types as a way of dividing data within an index by tenant might  
> > not work, since I will likely need to use the types for collection  
> > mapping.
> > 
> > Any advice on this issue is much appreciated.

--  
Alexandre Heimburger  
R&D Manager  
blueKiwi Software  
tel : +33687880997  
email : [ahb@bluekiwi-software.com](mailto:ahb@bluekiwi-software.com)  
adress : 93 rue Vieille du Temple, 75003 Paris

What is blueKiwi? blueKiwi - the first Enterprise Social Software Suite in  
the world building professional networks on conversations and relationships

- helps large organizations increase their productivity, foster innovations  
and boost people satisfaction.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 6, 2011, 10:13am UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/5 "2011-05-06T10:13:22Z")

</div>

Hiya

> I also work for a multi-tenant product and we have built one index per  
> tenant.

Using one index per tenant is often the right solution, but not if you  
have 10,000 tenants.

Each shard in each index is a separate Lucene instance, which has some  
overhead. On top of that, each node in the the cluster needs to  
maintain information about all indices and shards that exist in the  
cluster, which has its own overhead.

clint

---

<div class="post-metadata">

**Author:** ![Ellery\_Crane\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ellery_crane_2/32/3369_2.png) [@Ellery\_Crane\_2](https://discuss.elastic.co/u/Ellery_Crane_2)\
**Post date:** [May 6, 2011, 2:43pm UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/6 "2011-05-06T14:43:52Z")

</div>

On May 6, 6:00 am, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) wrote:

> A few options here. It may be possible to use a single type for all of  
> your tenants. For instance:
> 
> - if one tenant has fields foo and bar, and another has bar and baz,  
> you can store docs from both tenants in the same type, just adding  
> the relevant fields
> 
> - you mention different analysis. how would this be different? If  
> it is a question of language, then you might be able to make  
> this work by using the \_analyzer field:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/analyzer-field.html)
> 
> - if the mappings are so different that you don't want to combine them  
> into one type, then you could use different types within the  
> same index, eg user\_v1, user\_v2

Thanks for the ideas- I shall read up on each of them!

I also realize I might not have given a suitable example when I was  
discussing the document type requirements- my apologies. A clearer  
picture of my needs is something like this: within my application, I  
have different types of data, such as "User", "Location", "BlogPost",  
and so on. Every tenant would need to have documents of those types,  
but the structure of the data might be different from tenant to  
tenant. For instance, my "Location" documents might have Street, City,  
Zipcode, County and State for tenants in the United States, but  
Prefecture, Municipality, City, District, City Block, House Number,  
and Postal Code for tenants in Japan. This is a simple example- the  
structure could vary quite a bit more than that, including deeply  
nested fields, parent/child relationships, and so on. In addition to  
the different fields, I may also want to tokenize/analyze/etc the  
documents differently from tenant to tenant, depending on the needs of  
the users in the tenant. Given that, for instance, a US tenant would  
never want to see Location documents with a "Prefecture" field, or a  
Spanish tenant might want to analyze fields in 'BlogPost' documents  
differently than an English tenant, it seemed that multiple indexes  
with different type mappings in each was the way to go. However, I  
shall investigate the options that you presented- as I mentioned, I am  
still very much an elasticsearch newbie 🙂 Thank you again for your  
help!

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 6, 2011, 4:43pm UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/7 "2011-05-06T16:43:02Z")

</div>

Hi Ellery

> Given that, for instance, a US tenant would  
> never want to see Location documents with a "Prefecture" field,

If you don't set one, they won't see it. The mapping knows how to  
handle the different field types, but it won't "auto-create" them in  
your doc if they are missing.

> Spanish tenant might want to analyze fields in 'BlogPost' documents  
> differently than an English tenant,

Sure - using the \_analyze field I mentioned, you might be able to  
achieve what you want here.

clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:06am UTC](https://discuss.elastic.co/t/what-is-best-indexing-strategy-for-multitenant-data/4353/8 "2017-07-06T04:06:47Z")

</div>


