# Using huge NVMe disks with elasticsearch

**URL:** <https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228>\
**Category:** Elasticsearch\
**Created:** [June 17, 2021, 7:59am UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228 "2021-06-17T07:59:05Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Timur\_Makarchuk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timur_makarchuk/32/88755_2.png) [@Timur\_Makarchuk](https://discuss.elastic.co/u/Timur_Makarchuk)\
**Post date:** [June 17, 2021, 7:59am UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/1 "2021-06-17T07:59:05Z")

</div>

Hello everyone.

We're currently in process of choosing new hardware for our elasticsearch cluster.

Our current cluster consists of 32 nodes and holds 42TB of data across 2000 indices.

We're choosing hardware from the specs our provider has.  
One of options we're considering is the one that has 256GB RAM and 4x2TB NVMe SSD.  
We're planning to join those in RAID0 which would get us 8TB NVMe SSD per node.  
My questing is is it maybe a bit too much since some of our shards are pretty small and we may cross `20 shards or fewer per GB of heap memory` boundary.

And seeing is this node has too much RAM as well (although it can be used as cache) we were considering splitting those nodes into 4 LXC containers 64GB RAM and 2TB each. Which option would be preferable from elasticsearch perspective?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 17, 2021, 8:18am UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/2 "2021-06-17T08:18:41Z")

</div>

> [@Timur\_Makarchuk](#):
>
> Which option would be preferable from elasticsearch perspective?

Bare metal. But honestly, containerising things makes much more logical sense.

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [June 17, 2021, 12:58pm UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/3 "2021-06-17T12:58:56Z")

</div>

- why use raid0. you should use single disk /data01, /data02, /data03 etc... and elasticsearch will manage them. if you loose one disk you are only loosing 25% of shard on that node. if you use 8TB with Raid0 then one dead disk and you have 100% shard lost for that node.
- 256 RAM might be overkill and on that sense your logic is write to split it.

I am also in process of setting up same amount of NVME but 98gig ram 20 node cluster.

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [June 17, 2021, 1:01pm UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/4 "2021-06-17T13:01:09Z")

</div>

Performomance vs other benefit

[https://discuss.elastic.co/t/large-cluster-on-vm-vs-bare-metal/276017/2](https://discuss.elastic.co/t/large-cluster-on-vm-vs-bare-metal/276017/2)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 17, 2021, 1:05pm UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/5 "2021-06-17T13:05:36Z")

</div>

I believe multiple data paths is getting deprecated so raid0 is likely a better option. Recall seeing a discussion about that around here somewhere…

---

<div class="post-metadata">

**Author:** ![Timur\_Makarchuk](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timur_makarchuk/32/88755_2.png) [@Timur\_Makarchuk](https://discuss.elastic.co/u/Timur_Makarchuk)\
**Post date:** [June 17, 2021, 1:23pm UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/6 "2021-06-17T13:23:48Z")

</div>

Hi! Thank you for your reply.

@elasticforme Your questing mentions VMs which implies performance overhead much larger then one of LXC (which is not VMs, but Containers), so I'm not sure if reply to your question is applicable here.

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [June 17, 2021, 1:58pm UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/7 "2021-06-17T13:58:16Z")

</div>

whole, I didn't see that anywhere and I do have all my system with multiple data path.

Timur, yes vm/container not same but close because using same resource on same hardware

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 19, 2021, 8:28am UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/8 "2021-06-19T08:28:54Z")

</div>

> [@elasticforme](#):
>
> I didn't see that anywhere and I do have all my system with multiple data path.

Christian is right, they are deprecated as of 7.13 and 8.0 will require a node per data path instead:

> <https://github.com/elastic/elasticsearch/issues/71205>
>
> Multiple Data Paths (MDP) is a pseudo-software-RAID-0 feature within Elasticsear…ch allowing multiple paths to be specified in the path.data setting (which usually point to different disks). Although it has been used in the past as a simple way to run a multi-disk setups, it has long been a source of user complaints due to confusing or unintuitive behavior. Additionally, the implementation is complex, and not well-tested nor maintained, with practically no benefit over spanning the data path filesystem across multiple drives and/or running one node for each data path.
> 
> We have long advised against using MDP, and are now ready to deprecate and remove it. This is a meta-issue to track that work.
> 
> \- \[x\] Deprecate MDP in 7.13
> \- \[x\] Document migration path #71871
> \- \[\] ~Remove documentation from 8.0~
> \- \[\] ~Block MDP in 8.0~
> \- \[\] ~Remove MDP from 8.0~

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [June 20, 2021, 1:03am UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/9 "2021-06-20T01:03:44Z")

</div>

IMHO this is wrong move.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 18, 2021, 1:03am UTC](https://discuss.elastic.co/t/using-huge-nvme-disks-with-elasticsearch/276228/10 "2021-07-18T01:03:46Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
