# Dealing with large index collection strategy?

**URL:** <https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258>\
**Category:** Elasticsearch\
**Created:** [July 24, 2015, 4:57pm UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258 "2015-07-24T16:57:51Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![eeveaud](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eeveaud/32/3601_2.png) [@eeveaud](https://discuss.elastic.co/u/eeveaud)\
**Post date:** [July 24, 2015, 4:57pm UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258/1 "2015-07-24T16:57:51Z")

</div>

Hi,

we would like to deal with one year of various source of logs. that range from 80^6 events per day to 1000 per day with kibana on top to have reporting and dashbord of activity from those different sources.

we plan to have different sources of input so at the end we will have to deal with a collection of 365\*5 index  
all index collection will be standardized as much as possible in terms of idexed fields

as a POC we tried to index 6 month of logs from only one source and hited "memory heap" and "too many files open" and encountered some latency in the kibana search. 😉

as far as I understand ES + kibana is most used for short time analysis not realy for long lasting log analysis.

does ES is suitable for this kind of task and will it support thhis kind of scaling

and what will be the best architecture we can eploy to cover this kind of task.

best regards

Eric

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 25, 2015, 12:40am UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258/2 "2015-07-25T00:40:24Z")

</div>

You can do this, you just have to scale across a lot of nodes as you mention.  
How many is up to your use case though.

Make sure you implement doc values as much as possible, that will help.

---

<div class="post-metadata">

**Author:** ![eeveaud](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eeveaud/32/3601_2.png) [@eeveaud](https://discuss.elastic.co/u/eeveaud)\
**Post date:** [July 25, 2015, 7:44am UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258/3 "2015-07-25T07:44:15Z")

</div>

currently I am testing it on my desktop machine (16Gb ram, 8cpu) and I have set up a cluster with 4 nodes  
this is the initial setup for the toy study and to go further we will scale up over various VM in order to set up a more robust cluster.  
I try to push the toy case as far as possible and stress it to check the robustness.

> [@warkolm](#):
>
> Make sure you implement doc values as much as possible, that will help.

sorry can you emphasis what you mean by this ?

Eric

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 25, 2015, 8:15am UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258/4 "2015-07-25T08:15:23Z")

</div>

How many shards do you have in the cluster and what is the average shard size? Each shard in Elasticsearch is a separate Lucene index and carries with it a certain amount of memory and file descriptor overhead. For logging use cases a reasonable shard size is often from a few GB to tens of GBs, although we generally recommend keeping it below 50GB as very large shards can have a negative impact on recovery. If your shards are quite small, it may make sense to have applications/streams share indices (assuming mappings allow this), reduce the number of shards for the indices or even go from daily indices to weekly or monthly. One of the benefits of using time-based indices is that you can change the number of shards for an index for the next period if volumes change.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 26, 2015, 4:28am UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258/5 "2015-07-26T04:28:06Z")

</div>

Doc values = [https://www.elastic.co/guide/en/elasticsearch/guide/current/doc-values.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/doc-values.html)

---

<div class="post-metadata">

**Author:** ![eeveaud](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eeveaud/32/3601_2.png) [@eeveaud](https://discuss.elastic.co/u/eeveaud)\
**Post date:** [July 29, 2015, 2:25pm UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258/6 "2015-07-29T14:25:28Z")

</div>

thanks

curently running tests with  
expanded number of shard, doc-values and weekly index  
I will let you know.

Eric

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:58pm UTC](https://discuss.elastic.co/t/dealing-with-large-index-collection-strategy/26258/7 "2017-07-05T23:58:19Z")

</div>


