# Index type effective utilization

**URL:** <https://discuss.elastic.co/t/index-type-effective-utilization/58706>\
**Category:** Elasticsearch\
**Created:** [August 23, 2016, 3:45pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706 "2016-08-23T15:45:31Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![pranay.pramod](https://avatars.discourse-cdn.com/v4/letter/p/74df32/32.png) [@pranay.pramod](https://discuss.elastic.co/u/pranay.pramod)\
**Post date:** [August 23, 2016, 3:45pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/1 "2016-08-23T15:45:31Z")

</div>

I am trying to understand and effectively use the index type  
available in elasticsearch.  
However, I am still not clear how \_type meta field is different from any  
regular field of an index in terms of storage/implementation. I do  
understand avoiding\_type\_gotchas

For example, if I have 1 million records (say posts) and each post  
has a creation\_date. How will things play out if one of my index types  
is creation\_date itself (leading to ~ 1 million types)? I don't think it  
affects the way Lucene stores documents, does it?  
In what way my elasticsearch query performance be affected if I use  
creation\_date as index type against a namesake type say 'post'?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [August 23, 2016, 4:05pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/2 "2016-08-23T16:05:08Z")

</div>

While elasticsearch is scalable in many dimensions there is one where it is limited. This is the metadata about your indices which includes the various indices, doc types and fields they contain.  
These "mappings" exist in memory and are updated and shared around all nodes with every change. For this reason it does not make sense to endlessly grow the list of indices, types (and therefore fields) that exist in this cluster state. A type-per-document-creation-date registers a million on the one-to-ten scale of bad design decisions 🙂

---

<div class="post-metadata">

**Author:** ![pranay.pramod](https://avatars.discourse-cdn.com/v4/letter/p/74df32/32.png) [@pranay.pramod](https://discuss.elastic.co/u/pranay.pramod)\
**Post date:** [August 23, 2016, 5:10pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/3 "2016-08-23T17:10:22Z")

</div>

Thanks Mark. That was helpful. A colleague of mine proposed that idea and despite my intuition on bad design, I didn't have any documentation on index\_type to prove it.

---

<div class="post-metadata">

**Author:** ![fai](https://avatars.discourse-cdn.com/v4/letter/f/87869e/32.png) [@fai](https://discuss.elastic.co/u/fai)\
**Post date:** [August 23, 2016, 6:17pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/4 "2016-08-23T18:17:16Z")

</div>

Hi Mark, Thanks for sharing some insights on the internal design of elasticsearch. Let me clarify that we'll never come close to 1 million types. The retention policy is to keep 10-year data, which accounts for 365 x 10 types at most. To our understanding, types are designed for us to keep the documents in corresponding partitions. During query time, elasticsearch will only go through the documents in the type (date, in our specific case) range specified in the query. Since we have no other fields containing unchanging or fairly distributed (i.e. 90% on one, 10 on the others) values, date, as a primary filter on our report app, seems to be our only option to partition the data. Alternatively, we can have all our data in one default type. Which one do you think is more likely to be the bottleneck? Having 3650 types in the metadata sitting in memory or not leveraging the type feature to partition the data.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [August 23, 2016, 6:25pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/5 "2016-08-23T18:25:11Z")

</div>

> [@fai](#):
>
> To our understanding, types are designed for us to keep the documents in corresponding partitions

No, types in the same index share the same physical Lucene files. Behind the scenes elasticsearch applies a filter to the Lucene docs to only return the ones with your chosen type. This filter is no different to one you could implement with a custom field e.g. by defining a filtered alias [1]. The one difference with the many-types approach is that you would pollute your cluster state with thousands of near-identical mappings.

[1] [Index Aliases | Elasticsearch Guide [2.3] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/2.3/indices-aliases.html#filtered)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 23, 2016, 6:39pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/6 "2016-08-23T18:39:43Z")

</div>

It sounds to me like you are trying to use types to solve a problem [time-based indices](https://www.elastic.co/guide/en/elasticsearch/guide/current/time-based.html) are generally used for.

---

<div class="post-metadata">

**Author:** ![fai](https://avatars.discourse-cdn.com/v4/letter/f/87869e/32.png) [@fai](https://discuss.elastic.co/u/fai)\
**Post date:** [August 23, 2016, 7:32pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/7 "2016-08-23T19:32:33Z")

</div>

ok, if i understand the time-based-index feature correctly, I'll make several api calls (each with a specific time-based index specified) to cover the date range specified by the user. Is that correct? Since we can't really predict what date range the user would pick, we won't be able to come up with alias to group indices together. And do you imply that index is the only way we can get the data partitioning effect?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 23, 2016, 7:40pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/8 "2016-08-23T19:40:34Z")

</div>

You can specify multiple indices in a single call, so multiple requests is generally not needed. Kibana can efficiently determine which indices that may hold data for the specific time window selected through a call to the [field stats API](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-field-stats.html), and is then able to target only the indices required when it creates the request. Before the field stats API was available it instead used the naming convention of the indices to limit the indices that it needed to query. You also have the option to query indices using an index pattern, e.g. a common prefix, although this naturally will hit all indices. Indices with no data in the interval should however return quite quickly.

Another benefit with time based indices is that you can adapt the number of shards each index has over time and that way adapt to increasing or decreasing daily volumes. It also makes it very easy and efficient to manage retention period as entire indices can be deleted.

---

<div class="post-metadata">

**Author:** ![fai](https://avatars.discourse-cdn.com/v4/letter/f/87869e/32.png) [@fai](https://discuss.elastic.co/u/fai)\
**Post date:** [August 23, 2016, 9:10pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/9 "2016-08-23T21:10:09Z")

</div>

Great! This is very helpful. Thank you so much for putting that all together.

---

<div class="post-metadata">

**Author:** ![pranay.pramod](https://avatars.discourse-cdn.com/v4/letter/p/74df32/32.png) [@pranay.pramod](https://discuss.elastic.co/u/pranay.pramod)\
**Post date:** [August 24, 2016, 2:17am UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/10 "2016-08-24T02:17:01Z")

</div>

Hi Christian,

Good to learn about time based indices. I am still somewhat unsure if that's best fit for our needs.  
Our dashboard needs to display date\_histogram by default for a year. That means if we create an index per each day, we are querying 365 indices every time the dashboard is accessed (happens to be the landing page of our app). I am curious to understand how aggregation (date\_histogram) will perform over those many indices. Are date\_histograms designed to perform better with a single or handful of indices?

Many thanks for your inputs!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 24, 2016, 5:17am UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/11 "2016-08-24T05:17:14Z")

</div>

When using time based indices, each index does not necessarily have to correspond to exactly a day. The time period covered by a time based index and the number of shards used often depends of the data volumes being indexed. As you have a very long retention period, it may make more sense to use monthly indices than daily. This does however depend on the amount of data you are indexing each month. Having large numbers of very small indices is inefficient, both with respect to querying and resource utilisation, so you want to make sure that your average shard size typically is between a few GB and a few tens of GB in size.

Kibana is used with date histograms against time based indices all the time, so they work just fine with time based indices.

---

<div class="post-metadata">

**Author:** ![fai](https://avatars.discourse-cdn.com/v4/letter/f/87869e/32.png) [@fai](https://discuss.elastic.co/u/fai)\
**Post date:** [August 25, 2016, 12:55am UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/12 "2016-08-25T00:55:29Z")

</div>

We're considering the alternative of putting everything in one index. To better understand the trade-off (other than what you've mentioned above), we would like to confirm that all the documents in that one index will always be scanned through regardless of what kind of search queries is being performed. In other words, there's no other out-of-the-box mechanism available to make the search effort limited to a subset of the documents.

---

<div class="post-metadata">

**Author:** ![fai](https://avatars.discourse-cdn.com/v4/letter/f/87869e/32.png) [@fai](https://discuss.elastic.co/u/fai)\
**Post date:** [August 25, 2016, 3:01am UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/13 "2016-08-25T03:01:02Z")

</div>

Is that correct? Please confirm.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 28, 2016, 7:49am UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/14 "2016-08-28T07:49:16Z")

</div>

One of the problems with using a single index is that you can not change the number of shards once the index has been created. This means that you ideally need to know your data volumes up front in order to not end up with too many small shards or too few very large shards. Each query/aggregation is executed across all shards in parallel, but the processing of each shard is single-threaded. The query performance therefore depend on the size as well as number of shards.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:24pm UTC](https://discuss.elastic.co/t/index-type-effective-utilization/58706/15 "2017-07-05T22:24:38Z")

</div>


