# Guidance on setting up elasticsearch to handle terrabytes of logs

**URL:** <https://discuss.elastic.co/t/guidance-on-setting-up-elasticsearch-to-handle-terrabytes-of-logs/205579>\
**Category:** Elasticsearch\
**Created:** [October 29, 2019, 3:40am UTC](https://discuss.elastic.co/t/guidance-on-setting-up-elasticsearch-to-handle-terrabytes-of-logs/205579 "2019-10-29T03:40:35Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![deepak\_deore](https://avatars.discourse-cdn.com/v4/letter/d/9fc29f/32.png) [@deepak\_deore](https://discuss.elastic.co/u/deepak_deore)\
**Post date:** [October 29, 2019, 3:40am UTC](https://discuss.elastic.co/t/guidance-on-setting-up-elasticsearch-to-handle-terrabytes-of-logs/205579/1 "2019-10-29T03:40:35Z")

</div>

I need guidance on setting up an ES cluster which can consume ~1TB of logs per day and storing them for 30days.

Here is what I am thinking to do by procuring Basic license:

1. Elastisearch - 3 master nodes (x-pack and tls enabled) running on kubernetes (aws EKS), I have no idea about how much heap so will start with 10G heap and will increase gradually as needed. 1 replica of each index.
2. ES storage - gp2 EBS for starting, its max size is 16TB, I may need to increase the master nodes to accommodate more data if disks are getting full.
3. fluentd daemonset configured to send logs from k8 to ES (this is straight forward)
4. Kibana running on k8 with 2G heap to start with.

Can anyone guide me if I need to do the things differently here?

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [October 29, 2019, 3:52am UTC](https://discuss.elastic.co/t/guidance-on-setting-up-elasticsearch-to-handle-terrabytes-of-logs/205579/2 "2019-10-29T03:52:26Z")

</div>

Might be worth taking a look at [this blog](https://www.elastic.co/blog/sizing-hot-warm-architectures-for-logging-and-metrics-in-the-elasticsearch-service-on-elastic-cloud)

Recently we have changed Warm ratio to 160:1

Also might take a look at [this](https://www.elastic.co/blog/implementing-hot-warm-cold-in-elasticsearch-with-index-lifecycle-management)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 29, 2019, 6:12am UTC](https://discuss.elastic.co/t/guidance-on-setting-up-elasticsearch-to-handle-terrabytes-of-logs/205579/3 "2019-10-29T06:12:19Z")

</div>

I would also recommend you look at these webinars:

> **[Quantitative Cluster Sizing](https://www.elastic.co/webinars/elasticsearch-sizing-and-capacity-planning)**
>
> This webinar covers the capacity planning frameworks, methodologies, and best practices used by the solutions architects at Elastic. You will learn how to estimate the architecture requirements for typical Elasticsearch use cases.

> **[Optimizing Storage Efficiency in Elasticsearch](https://www.elastic.co/webinars/optimizing-storage-efficiency-in-elasticsearch)**
>
> This video covers the different types of nodes we use in Hot/Warm/Cold architectures and discuss their characteristics and the factors that determine how much data each node type can hold and how you go about optimizing for this.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 26, 2019, 6:22am UTC](https://discuss.elastic.co/t/guidance-on-setting-up-elasticsearch-to-handle-terrabytes-of-logs/205579/4 "2019-11-26T06:22:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
