# Which RAID config for 2TB \* 12disks

**URL:** <https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245>\
**Category:** Elasticsearch\
**Created:** [August 7, 2018, 6:53am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245 "2018-08-07T06:53:33Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 7, 2018, 6:53am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/1 "2018-08-07T06:53:33Z")

</div>

Hello,  
There are 6 servers, and each have 64GB ram and 12 spinning disks (2 TB per disk. so 24TB in total).

When I give 32GB RAM to ES, It seems that maximum storable index size for a node is around 4TB.  
So, If I set RAID 0, 20TB disks will be unused.  
And I can not increase replica due to the ES heap limit.

How about,  
Make 6 `Raid1` array? I mean (raid1 = 2disk) \* 6 ==\> 6 path, 12TB available disk space.  
Then one raid1 give to OS, `/data0`  
Five raid1 will be configured for `path.data=/data1,/data2,/data3,/data4,/data5`

How does it seems?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 7, 2018, 7:26am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/2 "2018-08-07T07:26:14Z")

</div>

> [@jihun](#):
>
> When I give 32GB RAM to ES, It seems that maximum storable index size for a node is around 4TB.

This will generally depend on your data and use-case as well as how well you optimize heap usage.

- What type of data are you indexing?
- Have you gone through and [optimised your mappings](https://www.elastic.co/guide/en/elasticsearch/reference/6.3/tune-for-disk-usage.html)?
- What is the average shard size in your cluster? Having lots of small shards and indices can be inefficient and [drive up heap usage](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster).
- Are you using any coordinating-only nodes?

---

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 7, 2018, 7:50am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/3 "2018-08-07T07:50:35Z")

</div>

@Christian_Dahlqvist

OK,  
I store user click stream logs to ES, a index per a day, with 5 shard, no replica.  
In a day, 2 billion logs indexed and it's index size is about 750GB.

I thinks, this click stream logs are not mission critical, so replica required not importantly.

Actually there are 6 hot-nodes, 2 warm-nodes, 3 master dedicated nodes

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/8/a/8a7ab2fa3ea07cec8518967321f58530afc28f46.png)

What I described in the first comment is about NEW warm-nodes.  
because I want to increase retention date for warm-node data.

Below is all about the current warm-node.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/5/f5e12af4fba523e6a9a9eebd2778599fa9d7dbf4.png)

## Kibana monitoring

### single node

![image](https://us1.discourse-cdn.com/elastic/original/3X/6/0/60e150a12c24a4ddec76bedbbe48a84815bc9d54.png)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/8/5/8550e9e58d366d7863aa791c730659be281d069d.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/a/6ad21f44475d0dc321258884f46fa39c2ddb0cc1.png)

### single index

![image](https://us1.discourse-cdn.com/elastic/original/3X/9/9/99919068a297dd79243e07cbab0e0afeecb4cd57.png)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/c/7cd3b289408eb3a168eeac6440f434bed06ea6e5.png)  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/3/63777627a3651f27ac604738ff46776cbdb8590a.png)

## Here is mapping of index.

```auto
{
  "jpl_raw_20180724": {
    "mappings": {
      "_default_": {
        "dynamic": "false",
        "_all": {
          "enabled": false
        },
        "_source": {
          "excludes": [
            "ac_hash",
            "event_hash"
          ]
        },
        "properties": {
          "action_id": {
            "type": "keyword"
          },
          "app_ver": {
            "type": "keyword",
            "ignore_above": 30
          },
          "classifier": {
            "type": "keyword"
          },
          "client_ip": {
            "type": "keyword"
          },
          "country": {
            "type": "keyword"
          },
          "deliver_delay_time": {
            "type": "long"
          },
          "device_id": {
            "type": "keyword",
            "ignore_above": 200
          },
          "event_time": {
            "type": "date"
          },
          "ingest_host": {
            "type": "keyword",
            "ignore_above": 50
          },
          "ingest_time": {
            "type": "date",
            "format": "epoch_millis"
          },
          "language": {
            "type": "keyword",
            "index": false
          },
          "os_name": {
            "type": "keyword",
            "ignore_above": 30
          },
          "os_ver": {
            "type": "keyword",
            "ignore_above": 30
          },
          "p0value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p1value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p2value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p3value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p4value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p5value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p6value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p7value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p8value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p9value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "product": {
            "type": "keyword"
          },
          "scene_id": {
            "type": "keyword",
            "ignore_above": 50
          },
          "service_id": {
            "type": "keyword"
          },
          "user_key": {
            "type": "keyword"
          }
        }
      },
      "default": {
        "dynamic": "false",
        "_all": {
          "enabled": false
        },
        "_source": {
          "excludes": [
            "ac_hash",
            "event_hash"
          ]
        },
        "properties": {
          "action_id": {
            "type": "keyword"
          },
          "app_ver": {
            "type": "keyword",
            "ignore_above": 30
          },
          "classifier": {
            "type": "keyword"
          },
          "client_ip": {
            "type": "keyword"
          },
          "country": {
            "type": "keyword"
          },
          "deliver_delay_time": {
            "type": "long"
          },
          "device_id": {
            "type": "keyword",
            "ignore_above": 200
          },
          "event_time": {
            "type": "date"
          },
          "ingest_host": {
            "type": "keyword",
            "ignore_above": 50
          },
          "ingest_time": {
            "type": "date",
            "format": "epoch_millis"
          },
          "language": {
            "type": "keyword",
            "index": false
          },
          "os_name": {
            "type": "keyword",
            "ignore_above": 30
          },
          "os_ver": {
            "type": "keyword",
            "ignore_above": 30
          },
          "p0value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p1value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p2value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p3value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p4value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p5value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p6value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p7value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p8value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "p9value": {
            "type": "keyword",
            "ignore_above": 300
          },
          "product": {
            "type": "keyword"
          },
          "scene_id": {
            "type": "keyword",
            "ignore_above": 50
          },
          "service_id": {
            "type": "keyword"
          },
          "user_key": {
            "type": "keyword"
          }
        }
      }
    }
  }
}

```

## Here is a single index stats

> <https://gist.github.com/jeesim2/5d7f4fae102923dcefd2a3f7576db0c7>

## Here is a single node info

> <https://gist.github.com/jeesim2/76cb7ac5dc3429ffd8c499e787bdabf6>

Additionally,  
when I force merged an index there were no remarkable heap reduce(just 5%?), as I remember..

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 7, 2018, 8:22am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/4 "2018-08-07T08:22:52Z")

</div>

Mappings look fine and properly optimised, so no problem there. You do however have quite large shards and the terms heap usage is high. I would recommend trying to reduce the average shard size to closer to 50GB to see if this makes a difference, e.g. by increasing the number of primary shards to between 15 and 18.

---

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 7, 2018, 8:51am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/5 "2018-08-07T08:51:13Z")

</div>

I'll try to do that.  
BTW, having more shard with smaller size can reduce the total heap usage?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 7, 2018, 8:56am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/6 "2018-08-07T08:56:47Z")

</div>

It may, which is why I am asking you to test it and then look at the index stats.

---

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 7, 2018, 11:29pm UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/7 "2018-08-07T23:29:00Z")

</div>

Now reindexing ongoing with 15 shards.  
It might takes 6 hours...

It seems that even though 15 shards reduces heap usage by 2-30% compared to 5 shards, I still cannot have replica. If 50% reduced I can have 1 replica.

Anyway I will check out the 15 shards index's stats.

Raid configuration, what I asked first, seems a little independent from the this 15 shard test. Because disk space quietly large. How do you think about set raid up regardless of heap optimizing if multiple raid1 make sense.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 8, 2018, 5:21am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/8 "2018-08-08T05:21:10Z")

</div>

> [@jihun](#):
>
> Raid configuration, what I asked first, seems a little independent from the this 15 shard test. Because disk space quietly large. How do you think about set raid up regardless of heap optimizing if multiple raid1 make sense.

Multiple RAID1 maks sense, and if you manage to reduce heap usage you will be able to use a larger portion of that storage.

---

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 8, 2018, 5:56am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/9 "2018-08-08T05:56:49Z")

</div>

Of course!  
Thanks for your help! : D

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [August 8, 2018, 6:01am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/10 "2018-08-08T06:01:33Z")

</div>

If a single daily index is nearly 750GB, I would consider writing hourly indices rather than daily. That will reduce them to a far more manageable ~30GB/index.

---

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 9, 2018, 12:22am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/11 "2018-08-09T00:22:35Z")

</div>

> If a single daily index is nearly 750GB, I would consider writing hourly indices rather than daily. That will reduce them to a far more manageable ~30GB/index.

It seems a good idea!

---

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 9, 2018, 12:31am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/12 "2018-08-09T00:31:09Z")

</div>

@Christian_Dahlqvist  
15 shards index does not shows an notable heap reduce.

## 15 shards

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/a/b/ab863d8ff2b252fac2da03af538e7a0a5c05705e.png)

## 5 shards

![image](https://us1.discourse-cdn.com/elastic/original/3X/1/9/19db1bf362746d3c3f4db430de3a2191ee52c157.png)

I did not reindexed but applied 15 shard to a new day's index.  
document size of both index are almost same.

one different is that

- 20180803 index's(15shards) segment count is around 650
- but 20180808 index's(5shards) segment count is around 290

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 9, 2018, 6:24am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/13 "2018-08-09T06:24:10Z")

</div>

It is a shame it did help, but I have to admit it was a long shot. As the mappings look good I am not sure I have any other suggestions apart from tweaking the circuit breaker thresholds a bit, but this is unlikely to give any massive improvement and could cause instability if pushed too far.

There is one thing I forgot to ask earlier: Do you allow Elasticsearch to automatically assign document IDs or do you set them yourself? If you set them yourself, what do the IDs look like and how are they generated?

---

<div class="post-metadata">

**Author:** ![jihun](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@jihun](https://discuss.elastic.co/u/jihun)\
**Post date:** [August 9, 2018, 6:31am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/14 "2018-08-09T06:31:51Z")

</div>

@Christian_Dahlqvist  
Could you tell me the technical background of 'more shards to reduce terms heap'? for next time when I meet similar situation and need heap optimizing. : )

> Do you allow Elasticsearch to automatically assign document IDs or do you set them yourself?

Document \_id automatically generated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 9, 2018, 7:19am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/15 "2018-08-09T07:19:58Z")

</div>

> [@jihun](#):
>
> Could you tell me the technical background of 'more shards to reduce terms heap'? for next time when I meet similar situation and need heap optimizing.

This was based on something I saw in a test of large shards (small sample size), but it seems there is no such direct correlation to shard size.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 6, 2018, 7:20am UTC](https://discuss.elastic.co/t/which-raid-config-for-2tb-12disks/143245/16 "2018-09-06T07:20:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
