# Elastic cluster destroyes SSD's?

**URL:** <https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985>\
**Category:** Elasticsearch\
**Created:** [July 6, 2022, 7:07am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985 "2022-07-06T07:07:45Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![azeiner](https://avatars.discourse-cdn.com/v4/letter/a/e19b73/32.png) [@azeiner](https://discuss.elastic.co/u/azeiner)\
**Post date:** [July 6, 2022, 7:07am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/1 "2022-07-06T07:07:45Z")

</div>

I'm using Elastic 7.17.3 in a 3 Node Cluster on Windows - we have Redis as a ingest and cache database and running logstash to transfer daten from redis to the cluster - so far so good.

Now i'm encoutering issues like destroyed SSD drives every now and then ... we've had the cluster no running for almost 2 years, then one SSD got corrupted and was done - we replaced it, another one went down 2 month later ... and then the climax - two nodes went down within a day ... so far so not soo good ...

we've analyzed the harddisks and got the following data (read with crystaldiskinfo):

HOST writes (sum): 342 PB  
NAND writes (sum): 1997 PB

This disk is readable although it can not be used any longer since we can not format or delete data on this SSD.

This behavior is pretty much the same on all other SSD's which corrupted over time.

So my question is, what are we doing wrong with this cluster? Can Elastic somehow be setup to not make that huge amout of writes or is this normal for a cluster of this size? Has anybody else encountered such an issue or similar?

our data (6 indices) is approx. 800-900 GB all in all so not really that much ...

Kind regards  
Andy

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [July 10, 2022, 10:40pm UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/2 "2022-07-10T22:40:53Z")

</div>

Are you constantly changing the data in Elasticsearch? Or is it write once, and never change over those 2 years?

---

<div class="post-metadata">

**Author:** ![azeiner](https://avatars.discourse-cdn.com/v4/letter/a/e19b73/32.png) [@azeiner](https://discuss.elastic.co/u/azeiner)\
**Post date:** [July 11, 2022, 6:29am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/3 "2022-07-11T06:29:04Z")

</div>

we are constantly importing data with logstash - not too much but its more or less constant 365. ist this a problem or how should we overcome this issue - we've had a single node cluster where we duplicating the data and there it is no problem at all - maybe the rebalancing and replicating of the data makes thing different and also leads to this huge amount of data written? i don't know to be honest ...

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2022, 9:27am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/4 "2022-07-11T09:27:38Z")

</div>

Are you importing immutable data or are you also updating/deleting existing data?

---

<div class="post-metadata">

**Author:** ![azeiner](https://avatars.discourse-cdn.com/v4/letter/a/e19b73/32.png) [@azeiner](https://discuss.elastic.co/u/azeiner)\
**Post date:** [July 11, 2022, 11:49am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/5 "2022-07-11T11:49:11Z")

</div>

we are only importing data - we've a redis cache where logstash imports all keys to Elasticsearch for further processing - for now no data is deleted. could it be that we've to reconfigure logstash? we are not doing any conversions or so so - just transfer from redis to elastic.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2022, 12:24pm UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/6 "2022-07-11T12:24:18Z")

</div>

Can you show the configuration of the Elasticsearch output(s) in Logstash?

Can you show the output of the [cat indices API](https://www.elastic.co/guide/en/elasticsearch/reference/8.2/cat-indices.html)?

---

<div class="post-metadata">

**Author:** ![azeiner](https://avatars.discourse-cdn.com/v4/letter/a/e19b73/32.png) [@azeiner](https://discuss.elastic.co/u/azeiner)\
**Post date:** [July 11, 2022, 12:41pm UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/7 "2022-07-11T12:41:17Z")

</div>

```auto

input {
	redis{

		host => "10.11.1.132"

		data_type => "list"

		key => "modbusdata"

		threads => 5

		type => "modbus"

	}
	

	redis{

		host => "10.11.1.132"

		data_type => "list"

		key => "ethercatdata"

		threads => 5

		type => "ethercat"

	}
	
	redis{

		host => "10.11.1.132"

		data_type => "list"

		key => "criodata"

		threads => 5

		type => "crio"

	}
	
	redis{

		host => "10.11.1.132"

		data_type => "list"

		key => "batchdata"

		threads => 5

		type => "batch"

	}
	
	redis{

		host => "10.11.1.132"

		data_type => "list"

		key => "parallel"

		threads => 5

		type => "parallel"

	}
	
	
}

filter {
  date {
    match => ["timestamp" , "yyyy-MM-dd HH:mm:ss.SSS"]
	target => "@timestamp"
  }

}

output {

	stdout { codec => rubydebug }

	if [type] == "modbus" {
		elasticsearch {

			hosts => ["10.11.1.132"]

			document_type => "modbus"

			index => "modbusdata"

		}
		elasticsearch {

			hosts => ["10.11.3.99"]

			index => "modbusdata"

		}		
	
	
	}

	
	if [type] == "ethercat" {
		elasticsearch {

			hosts => ["10.11.1.132"]

			document_type => "ethercat"

			index => "ethercatdata"

		}
		elasticsearch {

			hosts => ["10.11.3.99"]

			index => "ethercatdata"

		}
	
	}
	
	if [type] == "crio" {
		elasticsearch {

			hosts => ["10.11.1.132"]

			document_type => "modbus"

			index => "criodata"

		}
		elasticsearch {

			hosts => ["10.11.3.99"]

			index => "criodata"

		}		
	
	
	}
	
	if [type] == "batch" {
		elasticsearch {

			hosts => ["10.11.1.132"]

			document_type => "batch"

			index => "batchdata"

		}
		
		elasticsearch {

			hosts => ["10.11.3.99"]

			index => "batchdata"

		}
	}
	
	if [type] == "parallel" {
		elasticsearch {

			hosts => ["10.11.1.132"]

			document_type => "parallel"

			index => "parallel"

		}
		
		elasticsearch {

			hosts => ["10.11.3.99"]

			index => "parallel"
		}
		}
		}

```

green open .geoip\_databases AwH3KKrtRxScQSH4yMPHcA 1 0 40 40 37.7mb 37.7mb  
green open .kibana\_task\_manager\_7.17.3\_001 dN\_O0elCSYSJxllgWZSvgA 1 0 17 452473 128.9mb 128.9mb  
green open .apm-custom-link akvx0XUDQdWFL0HpYeJXZw 1 0 0 0 226b 226b  
yellow open parallel \_SrV8OE3RlynylAIGLymWw 1 1 15978223 0 2gb 2gb  
green open .apm-agent-configuration XktJLBlrSNGOge8Su9\_H7A 1 0 0 0 226b 226b  
yellow open ethercatdata ZdNsBIFuSnmyJ4vCjojbcA 1 1 3228353 0 854.8mb 854.8mb  
green open .async-search wwlChbkZTQecHyphNDeFGw 1 0 0 0 252b 252b  
yellow open batchdata TljGnOnmQd-5ZSPyKzJ7LA 1 1 15978239 0 1.6gb 1.6gb  
yellow open modbusdata 7tWIQap8RXyWtI0jl9Xdgw 1 1 3226520 0 915.3mb 915.3mb  
yellow open criodata VneVTPKLSh64KJ2zJ6cNIQ 1 1 3228355 0 1.5gb 1.5gb  
green open .kibana\_7.17.3\_001 3GoRaHAWSTCtVX8MUAaiRg 1 0 21 6 2.4mb 2.4mb

There is only one Node (in a single Node Cluster) available, the others are down because of the mentioned problems - one Node would be operable but sind it's the last Node in a 3 Node Cluster i didn't manage to get it up and running again - haven't tried the Node Tools yet!

Thank you Christian!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2022, 12:51pm UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/8 "2022-07-11T12:51:33Z")

</div>

It does not look like you are updating or deleting any data and the indices are quite small so I am not sure what would result in so much writes. Let's see if anyone else have any ideas or suggestions.

---

<div class="post-metadata">

**Author:** ![BenB196](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benb196/32/83401_2.png) [@BenB196](https://discuss.elastic.co/u/BenB196)\
**Post date:** [July 11, 2022, 7:17pm UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/9 "2022-07-11T19:17:18Z")

</div>

I think a good next step is seeing what is actually using the disk here, and confirm that Elasticsearch is really the issue. You can use the [Metricbeat System Diskio module](https://www.elastic.co/guide/en/beats/metricbeat/current/metricbeat-metricset-system-diskio.html), to track the disk usage. I don't think it directly shows what processes are using the disk, but should at least give you an idea. Might need to look at some other tools to see at the process level what processes are using the disk.

**Note:** If you can, I'd recommend sending this monitoring data to a secondary Elasticsearch cluster on different nodes to:

1. Not affect the performance of your current cluster
2. Not adversely affect the diskio metrics by collecting data then writing data to the same disks, thus further increasing the disk usage.

Edit: To further contextualize this question, I have a cluster which processes ~1TB/day across 6 hot/content nodes. Each node has its own backing SSD, and has been running ~1.25 years, and each SSD only has ~1.4PB written to it. So, Elasticsearch at your scale, writing 340PB is kind of insane. (Note: I run on Linux/Kubernetes, so the infrastructure isn't the same, but unless there is some bug on Windows, I doubt Elasticsearch would be the issue here.

---

<div class="post-metadata">

**Author:** ![azeiner](https://avatars.discourse-cdn.com/v4/letter/a/e19b73/32.png) [@azeiner](https://discuss.elastic.co/u/azeiner)\
**Post date:** [July 12, 2022, 5:59am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/10 "2022-07-12T05:59:26Z")

</div>

Thank you Ben, i totally agree - i'm using Elastic for a while now (mostly on smaller Clusters) and never had such an issue - but this one drives me insane. It might be, and this is something we've to check further, that the batch of SSD's is some kind of problematic since it's the third SSD we've got with such an issue. Metricbeat is a good idea, i'll set it up as you said since we are going to split from one cluster wit 3 nodes to 3 clusters with one node ... i think this should solve the purpose and wonder why we haven't done it before!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 12, 2022, 6:04am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/11 "2022-07-12T06:04:26Z")

</div>

I am not very familiar with Windows internals but it may be worthwhile checking whether you have any swap configured as that potentially could cause a lot of disk I/O.

---

<div class="post-metadata">

**Author:** ![azeiner](https://avatars.discourse-cdn.com/v4/letter/a/e19b73/32.png) [@azeiner](https://discuss.elastic.co/u/azeiner)\
**Post date:** [July 12, 2022, 6:07am UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/12 "2022-07-12T06:07:25Z")

</div>

I'll take a look at this, very good idea, since this is the only Cluster on Windows we're running there has to be some I/O problem - i don't think it's logstash or elastic itself!

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 22, 2022, 12:11pm UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/13 "2022-07-22T12:11:49Z")

</div>

Another user reports the same kind of thing here:

> [@Elastic causes very high disk usage with minimal data](https://discuss.elastic.co/t/elastic-causes-very-high-disk-usage-with-minimal-data/310384):
>
> Hello, as my other post explained our 3 Node cluster fried 2 SSDs in 4 hours. We are currently investigating what really happened. Our System: 3 Nodes, Elastic 7.17.3, Kibana 7.17.3, Every node has a 2TB SSD and 32GB of Ram, Node 1 Receives all the Data and moves it to the rest. The Drives are ONLY used by Elastic, nothing else is stored on them. Elastic has 16GB of ram for itself, no other tasks are running on the machines. The machines are physical, not virtual. We now have the cluster bac…

No idea what's going on here tho, Elasticsearch isn't doing anything to directly cause this level of write traffic.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 19, 2022, 12:12pm UTC](https://discuss.elastic.co/t/elastic-cluster-destroyes-ssds/308985/14 "2022-08-19T12:12:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
