# Recommended disk configuration for reasonable shard reallocation time

**URL:** <https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741>\
**Category:** Elasticsearch\
**Created:** [January 16, 2017, 1:29pm UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741 "2017-01-16T13:29:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![YuWatanabe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuwatanabe/32/13259_2.png) [@YuWatanabe](https://discuss.elastic.co/u/YuWatanabe)\
**Post date:** [January 16, 2017, 1:29pm UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/1 "2017-01-16T13:29:41Z")

</div>

I am estimating hardware configuration for customer environment but hesitating how should I configure the disk.

My reqirements are

2500 evt/sec (Maximum)  
58 TB of logs ( holding for 1year )  
5 shards, 1 replica for each index.

I am considering to split into 3 nodes which ends up with about 20 TB / node.  
Considering the size of storage, I want to use 4TB 7200 RPM NL-SAS with RAID 1 + 0.

However, I assume that 20 TB of NL-SAS RAID 1 + 0 will require long recovery time  
approximately ,

1491 hours (100 IOPS per disk , 5 disks per node )

May I ask for advice about how should I configure my node and disk in this case? I would appreciate any advice.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 16, 2017, 2:36pm UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/2 "2017-01-16T14:36:48Z")

</div>

20TB of data on a single node is in my experience too much. It is however hard to determine exactly how much data a single node can handle. I would recommend watching the following [video on cluster sizing](https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing) to get an idea how to best determine how much data a node can handle through benchmarking.

---

<div class="post-metadata">

**Author:** ![YuWatanabe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuwatanabe/32/13259_2.png) [@YuWatanabe](https://discuss.elastic.co/u/YuWatanabe)\
**Post date:** [January 16, 2017, 3:11pm UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/3 "2017-01-16T15:11:13Z")

</div>

@Christian_Dahlqvist

Thank you for the reply. Would you please describe the part **too much**?

Do people generally keep the volume tight , perhaps \<10TB, and increase the number of nodes when dealing with large volume of data?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 16, 2017, 3:29pm UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/4 "2017-01-16T15:29:02Z")

</div>

I typically see less than 10 TB per node for logging use cases, but am not aware of any hard limit, as this depends a lot on the use case. The time it takes to move data around in case of node failure is part of this, but query patterns and performance also play a part. It is typically recommended to keep shard sizes below 50 GB, as larger shards can cause problems for recovery unless you have very good network bandwidth. Each shard also comes with some overhead, which means that a node can not handle an infinite number of shards.

---

<div class="post-metadata">

**Author:** ![YuWatanabe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuwatanabe/32/13259_2.png) [@YuWatanabe](https://discuss.elastic.co/u/YuWatanabe)\
**Post date:** [January 18, 2017, 12:01am UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/5 "2017-01-18T00:01:02Z")

</div>

Thank you for clear experience.

To be honest this deployment will be our first large volume deployment.  
Would mixing SSD and Spiining disk using hot-warm architecture recommended solutin to deal with large volume data?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 18, 2017, 6:13am UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/6 "2017-01-18T06:13:52Z")

</div>

As you indexing load seems reasonable low, especially when spread out across the nodes in the cluster, I don't think a hot/warm archhitecture will buy you much. You need to optimize your mappings in order to optimize the size indices take up on disk and use the best\_compression codec. Then you need to benchmark to see how many nodes you need.

---

<div class="post-metadata">

**Author:** ![YuWatanabe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuwatanabe/32/13259_2.png) [@YuWatanabe](https://discuss.elastic.co/u/YuWatanabe)\
**Post date:** [January 20, 2017, 7:40am UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/7 "2017-01-20T07:40:54Z")

</div>

Ok. Thank you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 17, 2017, 7:41am UTC](https://discuss.elastic.co/t/recommended-disk-configuration-for-reasonable-shard-reallocation-time/71741/8 "2017-02-17T07:41:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
