# Need advice on building a new production ELK cluster

**URL:** <https://discuss.elastic.co/t/need-advice-on-building-a-new-production-elk-cluster/354344>\
**Category:** Elasticsearch\
**Created:** [February 28, 2024, 12:50pm UTC](https://discuss.elastic.co/t/need-advice-on-building-a-new-production-elk-cluster/354344 "2024-02-28T12:50:59Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![anon85145925](https://avatars.discourse-cdn.com/v4/letter/a/85f322/32.png) [@anon85145925](https://discuss.elastic.co/u/anon85145925)\
**Post date:** [February 28, 2024, 12:50pm UTC](https://discuss.elastic.co/t/need-advice-on-building-a-new-production-elk-cluster/354344/1 "2024-02-28T12:50:59Z")

</div>

Hello,

We are in the process of building a new ELK cluster, we are looking to incorporate HA and data resiliency in a DR setup (across 2 sites).  
We will be using a SAN in each of our 2 sites and will be looking to be ingesting anywhere between 500-700GB and retain the data in the cluster for about 1 year.

We are struggling to determine if we should be using containers or VMs.  
Another thing we are trying to understand is if we can make the storage independent from the elasticsearch nodes - ex. if a node fails the data is preserved on the SAN.

In terms of DR, we have also been exploring the different options of using CCR, doing logshipping to the secondary site or simply leveraging the snapshot/restore feature from elasticsearch.

Few of our biggest goals here are: 1) Make it as easy as possible to scale up; 2) Make it as hard as possible to lose data; 3) Make it as easy as possible to restore in case of a DR scenario.

We wanted to see if the community can give us some advise/help us determine which of the above paths we should/shouldn't take.

Thanks!

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 28, 2024, 2:20pm UTC](https://discuss.elastic.co/t/need-advice-on-building-a-new-production-elk-cluster/354344/3 "2024-02-28T14:20:49Z")

</div>

> [@anon85145925](#):
>
> maybe worth clarifying the ingestion rates I mentioned 500-700GB are per day and retained for a year, so in reality the underlying storage should be able to gold 700GBx365days.

Do you need the data to be searchable for 1 year or could you use data tier with different kinds of retention, like 15 days for hot data, 60 days for warm data and everything else in snapshots?

Also, to have some resiliency you need at least 1 replica, so 700GB/day with 1 replica will be 1.4 TB/day, to store this for a full year you will need more than 500 TB, which can be really expensive, even more if you want to have this on-premises and have the same infrastructure in multiple sites.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 28, 2024, 9:55am UTC](https://discuss.elastic.co/t/need-advice-on-building-a-new-production-elk-cluster/354344/5 "2024-03-28T09:55:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
