# Large Capacity Sizing: Single Cluster vs Multiple Clusters, Index Sizing, Memory Problems

**URL:** <https://discuss.elastic.co/t/large-capacity-sizing-single-cluster-vs-multiple-clusters-index-sizing-memory-problems/160812>\
**Category:** Elasticsearch\
**Created:** [December 13, 2018, 8:48pm UTC](https://discuss.elastic.co/t/large-capacity-sizing-single-cluster-vs-multiple-clusters-index-sizing-memory-problems/160812 "2018-12-13T20:48:45Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![thehybridtech](https://avatars.discourse-cdn.com/v4/letter/t/e79b87/32.png) [@thehybridtech](https://discuss.elastic.co/u/thehybridtech)\
**Post date:** [December 13, 2018, 8:48pm UTC](https://discuss.elastic.co/t/large-capacity-sizing-single-cluster-vs-multiple-clusters-index-sizing-memory-problems/160812/1 "2018-12-13T20:48:45Z")

</div>

Hiya all...

Some quick scaling questions to get best practices. We deal with a very large pipeline of events every second and doing another round of evaluation now that we have moved to ES6. I will share our current setup. Love to get thoughts and recommendations.

Average Event Size: 1.8kB  
Typical Event Per Second: 300k EPS (projected to 1 million EPS by end of next year)  
3 Racks of 12 servers split at rack into separate clusters.  
Each Server:

- 56 CPUs

- 768 GB Memory

- 12 spinning disks split into 6 RAID0

- 6 ES Data instances

- CPU pinning giving each Instance 9 cpus and 2 cpus dedicated to system operations

- Utilizing the Elasticsearch Docker image on Centos 7

- Support of 30 days worth of data at 100k EPS per rack

Each Cluster and Index Level:  
5 Dedicated Coordinator Nodes  
5 Dedicated Master Nodes  
72 Data Nodes

Each Index is time bucketed at 3 hours with 24 shards / 1 replica (Shards are 30 - 50gb)  
Each cluster runs around 25k Shards and 150 Billion Documents

Questions:

- Multiple Clusters or one giant cluster??
  - I have found that crossing around 100 data nodes there were some very interesting performance problems and I could achieve better ingest on multiple clusters instead of a single large clusters on ES5. Any improvements on ES6 or future ES7 to rethink this question?

- Small amount of large indexes with lots of shards or lots of smaller indexes with less shards? (Same amount of shards between the two forms)
  - I utilize search aliases and bucketing end user searches to limit index hits. Larger indexes will mean search aliases will mean less value.

- Recommendations on Memory heap problems. We are running in to memory problems due to a mix of fielddata and segments with the current design. This is limiting length of data we can search. This is also why we keep adding instances to the servers because I cannot get more hardware but need memory.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 13, 2018, 9:33pm UTC](https://discuss.elastic.co/t/large-capacity-sizing-single-cluster-vs-multiple-clusters-index-sizing-memory-problems/160812/2 "2018-12-13T21:33:35Z")

</div>

The fact that you have spinning disks across the board may be limiting your cluster. Have you looked at disk I/O and iowait? In order to minimise heap usage and be able to handle and query larger data volumes, you way want to make sure you follow the guidelines outlined in [this webinar](https://www.elastic.co/webinars/optimizing-storage-efficiency-in-elasticsearch).

25k shards of an average size of 40GB across 72 nodes gives about 13.5TB per node. Is that what you have?

---

<div class="post-metadata">

**Author:** ![thehybridtech](https://avatars.discourse-cdn.com/v4/letter/t/e79b87/32.png) [@thehybridtech](https://discuss.elastic.co/u/thehybridtech)\
**Post date:** [December 13, 2018, 9:42pm UTC](https://discuss.elastic.co/t/large-capacity-sizing-single-cluster-vs-multiple-clusters-index-sizing-memory-problems/160812/3 "2018-12-13T21:42:31Z")

</div>

Thank you @Christian_Dahlqvist for responding.

I have to live with current hardware limitations. As to disk IO though, I am not seeing any specific problems. We are typically around 2% IOWAIT with the rare spikes up to 5%.

Each ES data instance has access to 13.8TB so really close. 😃

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 10, 2019, 9:42pm UTC](https://discuss.elastic.co/t/large-capacity-sizing-single-cluster-vs-multiple-clusters-index-sizing-memory-problems/160812/4 "2019-01-10T21:42:36Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
