# Threshold selection - How to define it?

**URL:** <https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433>\
**Category:** Elasticsearch\
**Created:** [July 1, 2020, 8:13am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433 "2020-07-01T08:13:55Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Thomas74](https://avatars.discourse-cdn.com/v4/letter/t/3bc359/32.png) [@Thomas74](https://discuss.elastic.co/u/Thomas74)\
**Post date:** [July 1, 2020, 8:13am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/1 "2020-07-01T08:13:55Z")

</div>

Hi,

I have 4 data nodes with 25 Tb on each for a total of 100Tb of storage.

I want to change the cluster threshold because, by default, the option

`cluster.routing.allocation.disk.watermark.flood_stage` is at 95%.

That will mean I lost 1 250 Gb of data per node and so 5Tb in total 😕

I would like to change it but I don't know the best practice to have an optimized value.

I have 2 idea :

1. Based on my bigest index size (100Gb on primary index)

```auto
cluster.routing.allocation.disk.watermark.low : 200gb 
cluster.routing.allocation.disk.watermark.high : 150gb
cluster.routing.allocation.disk.watermark.flood_stage: 120gb 

```

Total losted space : 120x4=480gb

1. Based on the fact that my machines use LVM and risk of a complete crash is quite limited.

```auto
cluster.routing.allocation.disk.watermark.low : 50gb 
cluster.routing.allocation.disk.watermark.high : 20gb
cluster.routing.allocation.disk.watermark.flood_stage: 15gb 

```

Total losted space : 15x4=60Gb

Is my logic correct ?

Best regards,

Thomas R.

---

<div class="post-metadata">

**Author:** ![defalt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/defalt/32/71379_2.png) [@defalt](https://discuss.elastic.co/u/defalt)\
**Post date:** [July 1, 2020, 8:29am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/2 "2020-07-01T08:29:00Z")

</div>

Do you use SSD's? Because if not I don't know why you want to risk data loss just because of 5TB. 5TB are only 5% of the price you payed for all that storage. Just invest another 150$ and you are fine. But thats just my opinion. Maybe there is a good reason for your question.

---

<div class="post-metadata">

**Author:** ![Thomas74](https://avatars.discourse-cdn.com/v4/letter/t/3bc359/32.png) [@Thomas74](https://discuss.elastic.co/u/Thomas74)\
**Post date:** [July 1, 2020, 8:43am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/3 "2020-07-01T08:43:20Z")

</div>

We will reach our limit in few weeks and as I don't have more storage for now, I 'm trying to save some space with this.  
Also, I will need to buy another server to increase my storage size. That implie hardware, licences, etc.. extra cost.  
It's a shame to lose 1Tb per node I think but maybe I'm not right.

---

<div class="post-metadata">

**Author:** ![defalt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/defalt/32/71379_2.png) [@defalt](https://discuss.elastic.co/u/defalt)\
**Post date:** [July 1, 2020, 8:52am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/4 "2020-07-01T08:52:54Z")

</div>

If so your settings should be okay but bear in mind the possible risk that comes with it. Also if you have only a few weeks left these **4,4% of extra space** will give you only a **couple of days more**. So your investment has to be done either way. Default elasticsearch limitations are there for a reason, of course you can change them to your liking. The only problem that I see is that if you are at 99,4% full disks you will have no time to plan your further steps. If you have already planed what you will do when your disks reach that limit go for it^^.

---

<div class="post-metadata">

**Author:** ![Thomas74](https://avatars.discourse-cdn.com/v4/letter/t/3bc359/32.png) [@Thomas74](https://discuss.elastic.co/u/Thomas74)\
**Post date:** [July 1, 2020, 9:26am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/5 "2020-07-01T09:26:18Z")

</div>

> [@defalt](#):
>
> Also if you have only a few weeks left these **4,4% of extra space** will give you only a **couple of days more**. So your investment has to be done either way.

Ya I totally agree with you. We will have to invest etiher way.

> [@defalt](#):
>
> Default elasticsearch limitations are there for a reason

Yes I think it also but I don't really understand the reason. For cluster/shard sizing, the common response is "it depends".

Why not the same answer also for this subject ?

Like taking in count :

- Number of document/amount of data ingest per day
- Size of the biggest index

If you lose 5Gb on 100Gb storage it doesn't matter but when you work with petabytes of data it start to be annoying.

There's for sure a reason for this but I don't really understand the logic behind ^^'

> [@defalt](#):
>
> if you are at 99,4% full disks you will have no time to plan your further steps

About further steps, instead of deleting indexes, what kind of steps you think that we can do ? I think about these actions but I'm not sure :

- Reindex
- Remove replicas
- force\_merge
- Snapshot and delete

Maybe I'm going too further in my reflexions 🙂

---

<div class="post-metadata">

**Author:** ![defalt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/defalt/32/71379_2.png) [@defalt](https://discuss.elastic.co/u/defalt)\
**Post date:** [July 1, 2020, 9:37am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/6 "2020-07-01T09:37:30Z")

</div>

> [@Thomas74](#):
>
> it depends

is the standard answer of every elastic team member 😉. I thing they train it on their first they when they join the company. But I think its just hard to create general rules which work for people using 1GB of data and some others using petabytes and supercomputers (who knows what elastic is used for). I think the logic is that if you run into these limitations that you still have time to take steps against it. Maybe if a company doesn't check their cluster state and the suddenly realise its close to full they can still increase this size and order new servers.

> [@Thomas74](#):
>
> - Reindex
> - Remove replicas
> - force\_merge
> - Snapshot and delete

sounds good to me. Removing replicas is the first thing you should try. Maybe force merges will help but they will lead to longer query times and I dont think it helps that much. Snapshots are an idea. But you will need a lot of storage for them.

---

<div class="post-metadata">

**Author:** ![Thomas74](https://avatars.discourse-cdn.com/v4/letter/t/3bc359/32.png) [@Thomas74](https://discuss.elastic.co/u/Thomas74)\
**Post date:** [July 1, 2020, 10:16am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/7 "2020-07-01T10:16:10Z")

</div>

> [@defalt](#):
>
> is the standard answer of every elastic team member 😉

Haha that's completly that 😃

> [@defalt](#):
>
> I think the logic is that if you run into these limitations that you still have time to take steps against it.

Yes maybe the best thing is to know your own "SLA" and configure it in function of your actions possibility.

I have a better idea on how I will work on it. Thanks 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 29, 2020, 10:16am UTC](https://discuss.elastic.co/t/threshold-selection-how-to-define-it/239433/8 "2020-07-29T10:16:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
