# IDF per customer, many customers per index - best practices

**URL:** <https://discuss.elastic.co/t/idf-per-customer-many-customers-per-index-best-practices/17831>\
**Category:** Elasticsearch\
**Created:** [May 30, 2014, 10:48am UTC](https://discuss.elastic.co/t/idf-per-customer-many-customers-per-index-best-practices/17831 "2014-05-30T10:48:50Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Igor\_Kupczynski](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_kupczynski/32/1157_2.png) [@Igor\_Kupczynski](https://discuss.elastic.co/u/Igor_Kupczynski)\
**Post date:** [May 30, 2014, 10:48am UTC](https://discuss.elastic.co/t/idf-per-customer-many-customers-per-index-best-practices/17831/1 "2014-05-30T10:48:50Z")

</div>

Dear Elasticsearch Community,

There are many sources over the internet which recommend putting many  
customers into one index. One example is the Shay Banon's talk given at  
Berlin Buzzwords [1]. This approach has many advantages and the alternative

- one customer per index seems like a huge over-provisioning. By using  
aliases (with the "filter" clause) its trivial to create a virtual  
namespace per customer.

There is one thing the worries me a bit tough. As per the documentation [2]

Inverse document frequency

> How often does each term appear in the index? The more often, the _less_ relevant.  
> Terms that appear in many documents have a lower _weight_ than more  
> uncommon terms.

It seems that IDF will be calculated over the entire index. It makes sense,  
because this is calculated at the index time and not at the query time. Is  
this is a problem in the field? Do you know what can be the impact of other  
customers' documents over a single customer doing the search? Do you have  
any advices on optimizing the queries for such a use case? Any best  
practices?

To sum up, I'm a bit concerned with putting many customers on a single  
index, because the search ranking may be affected; but the alternative -  
index per customer is not feasible because of the huge number of customer.  
Do you have any hints here?

Thanks,  
Igor Kupczyński

[1] [https://speakerdeck.com/kimchy/elasticsearch-big-data-search-analytics](https://speakerdeck.com/kimchy/elasticsearch-big-data-search-analytics)  
[2] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/guide/current/relevance-intro.html)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [May 30, 2014, 12:48pm UTC](https://discuss.elastic.co/t/idf-per-customer-many-customers-per-index-best-practices/17831/2 "2014-05-30T12:48:16Z")

</div>

IDF is calculated per shard, and only in DFS search types, it is calculated  
over all nodes in an initial scatter phase.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

If you are concerned about IDF in a single multi-user index per aliased  
user index, you should consider to index as many docs as possible into the  
multi-user index. The more global docs the better. This will flatten out  
skewed IDF.

Another option is to route a customer to a single shard, this will avoid  
DFS search types at all to get global IDF, but does not scale for large  
number of docs per user.

If you have customers with very small indexes, and they can evaluate  
relevance scores, they can count IDF and may notice IDF is  
misleading/wrong. In that case, to hide this skew effect, you could group  
your users into users with classes of almost equal amount of docs (a "small  
doc number" customers index, a "medium doc number" customers index, and  
a "big doc number" customers index for example) . Also, you could try to  
classify customers into users with same kind of docs (if possible at all).

If you want proficient customers to take advanced control of their  
distributed scoring you would have to create an index per user and offer  
DFS search types to them.

Jörg

On Fri, May 30, 2014 at 12:48 PM, Igor Kupczyński [puszczyk@gmail.com](mailto:puszczyk@gmail.com)  
wrote:

> Dear Elasticsearch Community,
> 
> There are many sources over the internet which recommend putting many  
> customers into one index. One example is the Shay Banon's talk given at  
> Berlin Buzzwords [1]. This approach has many advantages and the alternative
> 
> - one customer per index seems like a huge over-provisioning. By using  
> aliases (with the "filter" clause) its trivial to create a virtual  
> namespace per customer.
> 
> There is one thing the worries me a bit tough. As per the documentation [2]
> 
> Inverse document frequency
> 
> > How often does each term appear in the index? The more often, the _less_ relevant.  
> > Terms that appear in many documents have a lower _weight_ than more  
> > uncommon terms.
> 
> It seems that IDF will be calculated over the entire index. It makes  
> sense, because this is calculated at the index time and not at the query  
> time. Is this is a problem in the field? Do you know what can be the impact  
> of other customers' documents over a single customer doing the search? Do  
> you have any advices on optimizing the queries for such a use case? Any  
> best practices?
> 
> To sum up, I'm a bit concerned with putting many customers on a single  
> index, because the search ranking may be affected; but the alternative -  
> index per customer is not feasible because of the huge number of customer.  
> Do you have any hints here?
> 
> Thanks,  
> Igor Kupczyński
> 
> [1] [https://speakerdeck.com/kimchy/elasticsearch-big-data-search-analytics](https://speakerdeck.com/kimchy/elasticsearch-big-data-search-analytics)  
> [2]  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/guide/current/relevance-intro.html)
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGdsVO9OZEMHZsimOJqC\_1\_\_0NorH1FgmvFa0VQZQoddg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGdsVO9OZEMHZsimOJqC_1__0NorH1FgmvFa0VQZQoddg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Igor\_Kupczynski](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_kupczynski/32/1157_2.png) [@Igor\_Kupczynski](https://discuss.elastic.co/u/Igor_Kupczynski)\
**Post date:** [May 30, 2014, 8:58pm UTC](https://discuss.elastic.co/t/idf-per-customer-many-customers-per-index-best-practices/17831/3 "2014-05-30T20:58:29Z")

</div>

Hi Jörg,

Thanks for your quick answer. I was not aware of this IDF calculation per  
shard in regular queries, but it makes sense - one more scatter-gather  
phase is required for the global stats. I'll probably end up with putting  
many (if possible similar) customers on a single index to make "avarage"  
the IDF. I do not want to go with a "customer" per index approach because,  
as you mentioned, it does not scale.

Cheers,  
Igor

On Friday, 30 May 2014 14:48:31 UTC+2, Jörg Prante wrote:

> IDF is calculated per shard, and only in DFS search types, it is  
> calculated over all nodes in an initial scatter phase.
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/guide/current/_search_options.html#_literal_search_type_literal)
> 
> If you are concerned about IDF in a single multi-user index per aliased  
> user index, you should consider to index as many docs as possible into the  
> multi-user index. The more global docs the better. This will flatten out  
> skewed IDF.
> 
> Another option is to route a customer to a single shard, this will avoid  
> DFS search types at all to get global IDF, but does not scale for large  
> number of docs per user.
> 
> If you have customers with very small indexes, and they can evaluate  
> relevance scores, they can count IDF and may notice IDF is  
> misleading/wrong. In that case, to hide this skew effect, you could group  
> your users into users with classes of almost equal amount of docs (a "small  
> doc number" customers index, a "medium doc number" customers index, and  
> a "big doc number" customers index for example) . Also, you could try to  
> classify customers into users with same kind of docs (if possible at all).
> 
> If you want proficient customers to take advanced control of their  
> distributed scoring you would have to create an index per user and offer  
> DFS search types to them.
> 
> Jörg
> 
> On Fri, May 30, 2014 at 12:48 PM, Igor Kupczyński \<[pusz...@gmail.com](mailto:pusz...@gmail.com)  
> \<javascript:\>\> wrote:
> 
> > Dear Elasticsearch Community,
> > 
> > There are many sources over the internet which recommend putting many  
> > customers into one index. One example is the Shay Banon's talk given at  
> > Berlin Buzzwords [1]. This approach has many advantages and the alternative
> > 
> > - one customer per index seems like a huge over-provisioning. By using  
> > aliases (with the "filter" clause) its trivial to create a virtual  
> > namespace per customer.
> > 
> > There is one thing the worries me a bit tough. As per the documentation  
> > [2]
> > 
> > Inverse document frequency
> > 
> > > How often does each term appear in the index? The more often, the _less_  
> > > relevant. Terms that appear in many documents have a lower _weight_ than  
> > > more uncommon terms.
> > 
> > It seems that IDF will be calculated over the entire index. It makes  
> > sense, because this is calculated at the index time and not at the query  
> > time. Is this is a problem in the field? Do you know what can be the impact  
> > of other customers' documents over a single customer doing the search? Do  
> > you have any advices on optimizing the queries for such a use case? Any  
> > best practices?
> > 
> > To sum up, I'm a bit concerned with putting many customers on a single  
> > index, because the search ranking may be affected; but the alternative -  
> > index per customer is not feasible because of the huge number of customer.  
> > Do you have any hints here?
> > 
> > Thanks,  
> > Igor Kupczyński
> > 
> > [1]  
> > [https://speakerdeck.com/kimchy/elasticsearch-big-data-search-analytics](https://speakerdeck.com/kimchy/elasticsearch-big-data-search-analytics)  
> > [2]  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/guide/current/relevance-intro.html)
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/f85af449-842d-4d11-b854-db4fcd6705f3%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4321ef70-877b-4810-b198-5a3cd0d2a4b9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4321ef70-877b-4810-b198-5a3cd0d2a4b9%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:25am UTC](https://discuss.elastic.co/t/idf-per-customer-many-customers-per-index-best-practices/17831/4 "2017-07-06T01:25:38Z")

</div>


