# Memory/cpu ratio to disk size

**URL:** <https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752>\
**Category:** Elasticsearch\
**Created:** [September 23, 2019, 8:27pm UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752 "2019-09-23T20:27:17Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![yuecong](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuecong/32/49061_2.png) [@yuecong](https://discuss.elastic.co/u/yuecong)\
**Post date:** [September 23, 2019, 8:27pm UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/1 "2019-09-23T20:27:17Z")

</div>

Could I get some comments on concerns/ insights on the following resource( cpu/ memory and disk size) configuration for one of my Elasticsearch cluster?

Data volume:

- throughput: 18K docs/ second ( very continous load)
- size: 720Gb per day.

index setting:  
replica: 1  
shard: 18 shards

node configurs:

3 cordinating nodes

- for each node: 8 cpus, 32 GB memory, 16GB java heap.

3 master nodes:

- for each node: 1 cpu, 8GB memory, 4Gb java heap, 50GB ssd disk

6 data nodes:

- for each node: 20 cpus, 100 GB memory, 32GB java heap, 10TB data disks

Thanks!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 24, 2019, 4:40am UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/2 "2019-09-24T04:40:41Z")

</div>

I have a few questions:

- How large and complex are your documents?

- What is your retention period?

- How will you query the data? How frequently? What are the query latency requirements?

- Are you using the latest version?

- What type of storage will your data nodes use?

---

<div class="post-metadata">

**Author:** ![yuecong](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuecong/32/49061_2.png) [@yuecong](https://discuss.elastic.co/u/yuecong)\
**Post date:** [September 24, 2019, 5:01am UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/3 "2019-09-24T05:01:49Z")

</div>

Thanks!

1, each document is a piece of log, like lo4j logs and nginx logs. The size for each document is from 1000 bytes to 2000 bytes.  
2, I am setting the retention period as 30 days  
3, We are using kibana to query the logs. The query latency requirement is not that strict. like less than 1 minute for a complicated query, but several seconds for normal queries. Besides, I have a job to periodically query the last doc to calculate some latency between the timestamp in the doc and the time I am indexing the doc and some \_cat api to get the current state of the cluster per 30 seconds.  
4, we are using 7.1 version. btw, I think upgrading from 7.1 to 7.x should not be as hard as upgrading from 6.x to 7.x, right?  
5, we are using ssd type of EBS. (e.g. io1 for AWS)

---

<div class="post-metadata">

**Author:** ![yuecong](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuecong/32/49061_2.png) [@yuecong](https://discuss.elastic.co/u/yuecong)\
**Post date:** [September 25, 2019, 11:00pm UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/4 "2019-09-25T23:00:34Z")

</div>

@Christian_Dahlqvist could you help give some insights when you have time. Thanks

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 26, 2019, 4:29am UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/5 "2019-09-26T04:29:18Z")

</div>

I would recommend watching the following videos:

> **[Quantitative Cluster Sizing](https://www.elastic.co/webinars/elasticsearch-sizing-and-capacity-planning)**
>
> This webinar covers the capacity planning frameworks, methodologies, and best practices used by the solutions architects at Elastic. You will learn how to estimate the architecture requirements for typical Elasticsearch use cases.

> **[Optimizing Storage Efficiency in Elasticsearch](https://www.elastic.co/webinars/optimizing-storage-efficiency-in-elasticsearch)**
>
> This video covers the different types of nodes we use in Hot/Warm/Cold architectures and discuss their characteristics and the factors that determine how much data each node type can hold and how you go about optimizing for this.

If we make the simplified assumption that your data will take up the same size on disk as the raw size and that you will have a replica for high availability you will generate 1.44TB indices per day. that will be around 7TB of data per node. As the nodes will be handling a lot of indexing as well as querying I would not be surprised to see some heap pressure before you reach that volume. I would therefore suspect you might need a larger cluster in terms of data nodes, but the only way to know for sure is to test.

---

<div class="post-metadata">

**Author:** ![yuecong](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuecong/32/49061_2.png) [@yuecong](https://discuss.elastic.co/u/yuecong)\
**Post date:** [September 26, 2019, 4:57am UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/6 "2019-09-26T04:57:19Z")

</div>

Thanks so much for the guidance.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 26, 2019, 4:58am UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/7 "2019-09-26T04:58:26Z")

</div>

Also make sure you read [this blog post about sharding practices](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 24, 2019, 4:58am UTC](https://discuss.elastic.co/t/memory-cpu-ratio-to-disk-size/200752/8 "2019-10-24T04:58:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
