# ES Cluster on ZFS with PCIe3.0 SSD/SATA SSD

**URL:** <https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911>\
**Category:** Elasticsearch\
**Created:** [February 16, 2016, 4:50pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911 "2016-02-16T16:50:14Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [February 16, 2016, 4:50pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/1 "2016-02-16T16:50:14Z")

</div>

Hi,

Just wondering if anyone has configured ES on a ZFS Pool of PCIe x3.0 SSDs and SATA3 SSDs yet and if how they found performance?

If so did you do anything special with the PCIe SSD and the L2ARC?

Thanks,  
BigPete

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 16, 2016, 6:27pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/2 "2016-02-16T18:27:06Z")

</div>

Are you running Linux?

---

<div class="post-metadata">

**Author:** ![german23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/german23/32/11052_2.png) [@german23](https://discuss.elastic.co/u/german23)\
**Post date:** [February 17, 2016, 7:54am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/3 "2016-02-17T07:54:38Z")

</div>

I cant say anything about PCI SSDs, but we use a RAID 5 of SATA-SSDs and switched from BTRFS to ZFS about 1 year ago.

The performance is much better with ZFS now and it runs very smoothly, with no problems since the migration.

Luckily our server have plenty of RAM, so we didnt configured L2ARC, as the ARC and the speed of the SSDs are more than enough for ES (ES is rather CPU , than disk bound)

We also enable the built-in ZFS compression with the fast LZ4 compress algorithm - the CPU overhead is rather low and its saving a huge amount of disk space.

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [February 17, 2016, 8:18am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/4 "2016-02-17T08:18:20Z")

</div>

> [@jprante](#):
>
> Are you running Linux?

Yep, my test box is running Debian 8

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [February 17, 2016, 8:20am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/5 "2016-02-17T08:20:46Z")

</div>

Good to know, I figured this could be a useful way to get the 'Hot ' (all SSD) cluster and 'Warm' (all HDD) a nice speed boost with little cost, even if I added 2x SSDs to the 'Warm' as L2ARC+LZ4 and left the 'Hot' as just standard without compression?

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 17, 2016, 11:12am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/6 "2016-02-17T11:12:23Z")

</div>

While ZFS is surely an advanced file system, if you want performance with SSD, I recommend XFS and hardware RAID. HW RAID controller are much faster than ZFS pools.

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [February 17, 2016, 11:53am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/7 "2016-02-17T11:53:53Z")

</div>

If budget is a serious constraint where nice array controllers, enterprise level disks etc aren't achievable would you say that ZFS would be a better use of disk resource than just going Ext3?

---

<div class="post-metadata">

**Author:** ![ttys0](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ttys0/32/462_2.png) [@ttys0](https://discuss.elastic.co/u/ttys0)\
**Post date:** [February 17, 2016, 4:13pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/8 "2016-02-17T16:13:05Z")

</div>

I'm just working on preliminary setup and config, but so far putting Elasticsearch on ZFS has been as performant as XFS on CentOS 7 systems. Personally, I find the other features of ZFS to be the benefits that push me over to using ZFS on all but the OS disk.

Now, there is currently one **HUGE** caveat to this. If you are going to put Elasticsearch on ZFS using the current ZoL release (0.6.5.4), **MAKE SURE** you create the ZFS filesystem with the _xattr=sa_ option. Without this, there's a very good chance that the ZFS filesystem will not correctly free up deleted blocks.

The GitHub Issue is here : [https://github.com/zfsonlinux/zfs/issues/1548](https://github.com/zfsonlinux/zfs/issues/1548)

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [February 17, 2016, 4:46pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/9 "2016-02-17T16:46:55Z")

</div>

Thats a good shout thanks, it wasnt something I was aware of!

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [February 17, 2016, 5:20pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/10 "2016-02-17T17:20:25Z")

</div>

Just so you know I've been following along with this. I don't have anything to add other than "cool" and "good luck". Personally I'd love to be able to use ZFS and I salute you for doing so. Just keep in mind that you are one of the few folks who do so getting help might be difficult.

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 17, 2016, 6:05pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/11 "2016-02-17T18:05:07Z")

</div>

While you're at it, here is the best summary I know of what has to be planned before using ZFS, a little old, so some facts are outdated, but most points are still valid:

> **[ZFS: Read Me 1st](http://nex7.blogspot.de/2013/03/readme1st.html)**
>
> Things Nobody Told You About ZFS Yes, it's back. You may also notice it is now hosted on my Blogger page - just don't have time to deal w...

For what it's worth, I have never been able to set up ZFS to match the performance of XFS under Linux.

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [February 18, 2016, 8:23am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/12 "2016-02-18T08:23:07Z")

</div>

Great information thanks everyone, I'll maybe add XFS into my bench-marking to ensure I've covered all bases. The Hot tier is likely to be working with 8,035,200,000 docs (4.3TB) a month so need to try and optimize every possible piece.

---

<div class="post-metadata">

**Author:** ![rusty](https://avatars.discourse-cdn.com/v4/letter/r/f17d59/32.png) [@rusty](https://discuss.elastic.co/u/rusty)\
**Post date:** [February 18, 2016, 10:30am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/13 "2016-02-18T10:30:30Z")

</div>

Hi! As I can remember ZFS on Linux by default is very memory hungry and in most cases should be limited by tuning [ARC min/max memory](https://wiki.gentoo.org/wiki/ZFS#ARC). Another point that Marvel does not understand how to get disk usage from ZFS, but it can be fixed with recent versions (don't know).

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [August 15, 2016, 7:47am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/14 "2016-08-15T07:47:46Z")

</div>

I figured I'd post a follow up on this. I've got the hardware end and have been benchmarking with Bonnie++

2x Xeon 2.1Ghz E5-2620 v4  
Supermicro X10DAI  
128GB DDR4 2133mhz ECC  
2x Samsung 512GB 950 Pro (via PCIe Gen3) (for cache drives)  
1x LSI Megaraid 93622-8i  
8x Samsung 1TB 850 Pro in a Raid5 hardware array.

I'm running VMWare ESXi6 to separate the guest OS from the underlying hardware so was able to spin up a Windows VM and run CrystalDiskMark as a comparison.

8x 850 Drives in Raid 5

- CrystalDiskMark Results
- Seq Q32T1 - Read: 4137MB/s - Write: 2932MB/s
- 4K Q32T1 - Read: 492MB/s - Write: 155MB/s
- Seq - Read: 2544MB/s - Write: 2357MB/s
- 4K - Read: 30MB/s - Write: 65MB/s

2x 950 NVMe Drives in Raid 0

- CrystalDiskMark Results
- Seq Q32T1 - Read: 3219MB/s - Write: 3065MB/s
- 4K Q32T1 - Read: 447MB/s - Write: 436MB/s
- Seq - Read: 2574MB/s - Write: 2511MB/s
- 4K - Read: 41MB/s - Write: 96MB/s

Which when I compare a zfs pool (of the raid5 volume) + 2 nvme cache drives the performance isn't great!

Output Block; 298MB/sec  
Output Rewrite: 166MB/sec  
Input Block: 1564MB/sec

Does anyone have any suggestions on a better configuration or disk config? I'm going to try EXT4 and XFS as a comparison now to rule out an Debian/Bonnie issues.

Thanks,  
BigPete

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [August 16, 2016, 7:30pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/15 "2016-08-16T19:30:44Z")

</div>

Don't use RAID5 if you are after write speed.

ZFS uses advanced features (checksums for everything, deduplication, compression) which are not targeted for maximum speed.

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [August 22, 2016, 7:19am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/16 "2016-08-22T07:19:04Z")

</div>

I've checked that controller and hoped to see JBOD but don't think it supports it, it was also to help reduce the number of drives needed to be passed through to ESXi (and then the guest VM). I'm not that worried about write speed for ES as it'll be continuous slow small writes where as the reads would be where the performance is needed.

---

<div class="post-metadata">

**Author:** ![BigPete](https://avatars.discourse-cdn.com/v4/letter/b/d07c76/32.png) [@BigPete](https://discuss.elastic.co/u/BigPete)\
**Post date:** [October 13, 2016, 7:51am UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/17 "2016-10-13T07:51:51Z")

</div>

Just as an update.

I've tried ZFS/XFS/bcache and cant get any of the benchmarks to match the Windows Crystal Disk Mark results... really at a loss as to why performance is so different between the two OS's. (Unless its just the way Bonnie++ vs Crystal are displaying the results?)

Anyone any thoughts?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:12pm UTC](https://discuss.elastic.co/t/es-cluster-on-zfs-with-pcie3-0-ssd-sata-ssd/41911/18 "2017-07-05T22:12:44Z")

</div>


