# IO / Disc intensive load testing

**URL:** <https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722>\
**Category:** Elasticsearch\
**Created:** [September 7, 2020, 7:43am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722 "2020-09-07T07:43:30Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Oleg\_Ruchovets](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleg_ruchovets/32/60137_2.png) [@Oleg\_Ruchovets](https://discuss.elastic.co/u/Oleg_Ruchovets)\
**Post date:** [September 7, 2020, 7:43am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/1 "2020-09-07T07:43:30Z")

</div>

Hello,  
My task is to make an elastic search disk/io "hard life" :-). To not reinventing the wheel - is there a tool for simulation high disk load on elastic?  
what type of query is the most disk/io intensive? I saw many articles and the best practices on how to tune performance but my task is really to kill disk/io for elastic and I didn't find any useful open information on the web.

Thank you in advance.

---

<div class="post-metadata">

**Author:** ![whatgeorgemade](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/whatgeorgemade/32/103246_2.png) [@whatgeorgemade](https://discuss.elastic.co/u/whatgeorgemade)\
**Post date:** [September 7, 2020, 8:14am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/2 "2020-09-07T08:14:44Z")

</div>

Welcome!

You could use [Rally](https://esrally.readthedocs.io/en/stable/) for this. It's a purpose-built benchmarking tool for Elasticsearch.

There are several specimen datasets available and you can configure the load to put on the cluster.

---

<div class="post-metadata">

**Author:** ![Oleg\_Ruchovets](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleg_ruchovets/32/60137_2.png) [@Oleg\_Ruchovets](https://discuss.elastic.co/u/Oleg_Ruchovets)\
**Post date:** [September 7, 2020, 8:38am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/3 "2020-09-07T08:38:31Z")

</div>

Thank you for the answer @whatgeorgemade. I saw Rally. I probably wrong but what I understood Rally is mostly used or benchmark of different elastic search versions aka comparing the performance of version A against version B.

In my case, I am trying to cause disk high load testing. For example - a very complicated query requires a full index scan and running or heavy write insertions. I don't know what kind of operation/queries this could be. For example, relation databases "hates" nested joins. I can hang DB by this kind of query.  
In the case of elastic what kind of query/insertions can cause disk saturation.

Does Rally have such capabilities?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 7, 2020, 8:53am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/4 "2020-09-07T08:53:43Z")

</div>

Heavy indexing is probably the easiest way to saturate disk I/O assuming you have sufficient networking throughput and CPU for that to not be a bottleneck.

---

<div class="post-metadata">

**Author:** ![whatgeorgemade](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/whatgeorgemade/32/103246_2.png) [@whatgeorgemade](https://discuss.elastic.co/u/whatgeorgemade)\
**Post date:** [September 7, 2020, 8:54am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/5 "2020-09-07T08:54:32Z")

</div>

You can use Rally to benchmark your own cluster. If you have data already, you can even benchmark writing those documents, and running your own queries.

If you do have your own data, the best way to start is by creating a new track based on one of your existing indices. Have a look [here](https://esrally.readthedocs.io/en/stable/adding_tracks.html) to see how that's done.

If you don't have your own data, you can look through the sample datasets and use one that is close to the type of document you're going to be indexing.

---

<div class="post-metadata">

**Author:** ![Oleg\_Ruchovets](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleg_ruchovets/32/60137_2.png) [@Oleg\_Ruchovets](https://discuss.elastic.co/u/Oleg_Ruchovets)\
**Post date:** [September 7, 2020, 9:09am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/6 "2020-09-07T09:09:34Z")

</div>

Ok, what kind of documents it should be. Is it a massive index of simple documents or I should be kind of Wikipedia. By default elastic is indexing by every JSON field , right? if I will simulate load with 100 - 1000 field each document, does the number of fields effect on indexing effort and as a result performance?

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 7, 2020, 9:17am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/7 "2020-09-07T09:17:29Z")

</div>

I am not sure what the relationship between document size and disk I/O is so suspect you will need to test. I do not think you need very large documents to saturate I/O so you can probably use one of the existing Rally datasets.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 5, 2020, 9:17am UTC](https://discuss.elastic.co/t/io-disc-intensive-load-testing/247722/8 "2020-10-05T09:17:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
