# Raid 0 SSD?

**URL:** <https://discuss.elastic.co/t/raid-0-ssd/46504>\
**Category:** Elasticsearch\
**Created:** [April 6, 2016, 9:36am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504 "2016-04-06T09:36:17Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 6, 2016, 9:36am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/1 "2016-04-06T09:36:17Z")

</div>

I was considering setting up servers, with the data disk being raid 0.. since we have a replica on another server.. I figured that would a good way to save a lot of money (SSD's is by far the most expensive part of a new cluster setup).

I figured I'd use 1TB disks - and have approx 10 pcs. per server.. that ofcourse means that the risk of failure is ~X10.. since only of the 10 disks needs to fail, for the entire servers datastore to fail.

Does anyone have any experience with doing that.. or is it just a stupid idea? 🙂

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 6, 2016, 9:50am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/2 "2016-04-06T09:50:14Z")

</div>

Why not use multiple `path.data` entries and let ES "stripe" it?

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 6, 2016, 10:19am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/3 "2016-04-06T10:19:23Z")

</div>

interesting.. so I should simply just present EACH SSD disk.. on seperate mount points.. and then let ES handle it.. meaning I'd only loose the shards on one disk.

Only issue with that - is if one shard ever grew above ~960GB (the size of one SSD).. but that should be very unlikely.. since I'm using daily indices and at minimum 4 shards per index.

Can elasticsearch handle distribution disk usage over multiple paths like that? that would be pretty sweet.

Anyone using something like that?

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 6, 2016, 11:05am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/4 "2016-04-06T11:05:43Z")

</div>

Someone tried to ask the same here - with no real answer.. [Using RAID 0 vs multiple data paths after commit #10461](https://discuss.elastic.co/t/using-raid-0-vs-multiple-data-paths-after-commit-10461/855/4)

it seems no one is actually using the multiple paths approach.. even though with \*TB cluster sizes - thats a lot of money saved if its more stable than raid 0 (which has 10X the risk - with 10 disk backing it- of failure).

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 6, 2016, 11:11am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/5 "2016-04-06T11:11:25Z")

</div>

@warkholm - it seems ES no longer stripes shards.. avoiding striping causing shard failures because all shards are spread over many paths.. [https://github.com/elastic/elasticsearch/issues/9498](https://github.com/elastic/elasticsearch/issues/9498) - so from 2.0+ it should be good to have multiple paths.. and you'll only loose parts of your data, when a disk fails (instead of an entire raid 0)

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 6, 2016, 11:14am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/6 "2016-04-06T11:14:15Z")

</div>

how write performance is - compared to shards.. it should be much worse.. depending on how well ES shares writes over shards.. and with multiple paths PER server.. it would probably make sense to have more shards per index ?

6 servers with 10 paths = 60 disks.. and if we want to share the write load as well as possible.. we'd need a lot more shards than the usual 6.. so write performance will suffer, compared to raid 0.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [April 6, 2016, 2:40pm UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/7 "2016-04-06T14:40:46Z")

</div>

Personally I prefer raid 0 to multiple data paths, even with 2.0's fixes. I see it as a performance vs safety tradeoff. Usually I'm fine with the safety that comes from sharding. But I agree that it is a nice tradeoff to be able to make. I wouldn't for example, raid 0 four disks together. It is just too much bother.

If you have a hot/warm setup where you have historical data in Elasticsearch you could use multiple data.paths to spinning disks on nodes with a dozen disks for the warm nodes and raid 0 to two or three SSDs for your hot nodes.

I think one shard getting bigger than a whole SSD isn't a good argument. You have other problems if you let a shard get that big like recovery time. I'd shoot for shards an order of magnitude smaller than your SSDs.

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [April 6, 2016, 8:48pm UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/8 "2016-04-06T20:48:10Z")

</div>

Why multiple paths, addressing each SSD as single disk?

If you use RAID 0 on hardware controller (not the OS based crap), you can multiply the speed of each disk. E.g. 8 disks on RAID 0 give 8x speed (assuming the HW controller can cope with that transport capacity, e.g. a 12Gb/s SAS controller)

In my setup, indexing speed is crucial. RAID0 pays you back every EUR you invest in SSD.

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 6, 2016, 9:02pm UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/9 "2016-04-06T21:02:29Z")

</div>

I have one server - with 24 pcs. of 240GB disk.. in raid 10. I've had 2 disks fail now - within a few months. with just 10x times - the MTBF is rather high - and one disk - and the entire raid 0 is out.. that would be avoided if you used each disk seperately. But then you don't spread write load over many SSD's as you also correctly note.. It would be nice to know what people have good experience with 🙂

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 7, 2016, 6:35am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/10 "2016-04-07T06:35:51Z")

</div>

> [@nik9000](#):
>
> If you have a hot/warm setup where you have historical data in Elasticsearch you could use multiple data.paths to spinning disks on nodes with a dozen disks for the warm nodes and raid 0 to two or three SSDs for your hot nodes.

How do I make elasticsearch migrate shards away from the hot area every night? If thats possible, that could be a good solution - to split up into raid0 + LongTerm storage (spinning disks + perhaps som SSD cache)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 7, 2016, 6:38am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/11 "2016-04-07T06:38:29Z")

</div>

Take a read of [https://www.elastic.co/blog/hot-warm-architecture](https://www.elastic.co/blog/hot-warm-architecture)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 7, 2016, 6:38am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/12 "2016-04-07T06:38:54Z")

</div>

> [@KlavsKlavsen](#):
>
> split up into raid0 + LongTerm storage

The only way to do that on a single node is to run multiple ES instances.

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 7, 2016, 6:50am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/13 "2016-04-07T06:50:50Z")

</div>

so what nik9000 suggested actually isn't possible ?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 7, 2016, 6:51am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/14 "2016-04-07T06:51:32Z")

</div>

You can do that, you just need the different storage attached to different nodes.

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 7, 2016, 7:08am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/15 "2016-04-07T07:08:27Z")

</div>

hmm. can I tell ES to move index from one host to another.. or do I have to dump and restore (+some routing to ensure it doesn't end up on the same node as before :)?

---

<div class="post-metadata">

**Author:** ![anhlqn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anhlqn/32/5454_2.png) [@anhlqn](https://discuss.elastic.co/u/anhlqn)\
**Post date:** [April 7, 2016, 3:34pm UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/16 "2016-04-07T15:34:14Z")

</div>

Yes, you can by setting node attributes and route indexes to certain nodes  
[https://www.elastic.co/guide/en/elasticsearch/reference/2.3/shard-allocation-filtering.html](https://www.elastic.co/guide/en/elasticsearch/reference/2.3/shard-allocation-filtering.html)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 7, 2016, 8:29pm UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/17 "2016-04-07T20:29:44Z")

</div>

Read that blog post I posted.

---

<div class="post-metadata">

**Author:** ![KlavsKlavsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/klavsklavsen/32/74145_2.png) [@KlavsKlavsen](https://discuss.elastic.co/u/KlavsKlavsen)\
**Post date:** [April 8, 2016, 6:23am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/18 "2016-04-08T06:23:00Z")

</div>

Link? googleing "Mark Warkom" blog gives no relevant hits

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 8, 2016, 6:40am UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/19 "2016-04-08T06:40:42Z")

</div>

I think he meant the link he posted here: [Raid 0 SSD?](https://discuss.elastic.co/t/raid-0-ssd/46504/11?u=dadoonet)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:01pm UTC](https://discuss.elastic.co/t/raid-0-ssd/46504/20 "2017-07-05T23:01:16Z")

</div>


