# Esrally to test resiliency of the server

**URL:** <https://discuss.elastic.co/t/esrally-to-test-resiliency-of-the-server/260150>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [January 5, 2021, 2:32am UTC](https://discuss.elastic.co/t/esrally-to-test-resiliency-of-the-server/260150 "2021-01-05T02:32:33Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ven1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ven1/32/81355_2.png) [@ven1](https://discuss.elastic.co/u/ven1)\
**Post date:** [January 5, 2021, 2:32am UTC](https://discuss.elastic.co/t/esrally-to-test-resiliency-of-the-server/260150/1 "2021-01-05T02:32:33Z")

</div>

I like to test the resiliency of the elasticsearch cluster when there is a server failure (1 Elasticsearch Node is down).

I am following the below process

- esrally rally as loadgenerator and to deploy a 4 node cluster.  
\*` "number_of_replicas=1"`:
- Using event data track with challenge "`--challenge=index-logs-fixed-daily-volume` "
- Shutdown one of the elasticsearch node I get "`asyncio.exceptions.CancelledError"` error and challenge does not complete.

Since there are 2 replicas my expectation was elasticsearch workload will continue with warning but that is not happening.

Any pointer on how to achieve this test with rally ?

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [January 21, 2021, 7:22am UTC](https://discuss.elastic.co/t/esrally-to-test-resiliency-of-the-server/260150/2 "2021-01-21T07:22:44Z")

</div>

Hi,

I fear Rally is maybe not the right tool for this job as its main purpose is benchmarking and one of the core assumptions is that the system is in a steady state so we can perform reproducible measurements.

What you're after seems more like a QA test ensuring stability of the cluster after it has lost a node. One of the features of the Python client is to use [sniffing](https://elasticsearch-py.readthedocs.io/en/master/index.html#sniffing) to learn about the current cluster topology so it will know about disappearing or new cluster nodes. As a corollary of requiring a system in steady state Rally does not enable cluster sniffing. We also do not retry failed requests but only record that a failure has happened as retries will skew measurement results as well.

It would probably make sense that you instead write a small test harness e.g. using the Elasticsearch Python client and enable sniffing and maybe also implement retry handing in case of failed requests. You could still leverage the data sets that we offer with Rally though in order to generate some load. At the end of your test run you'd likely also want to assert that all documents have been ingested successfully.

Hope that helps.

Daniel

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 18, 2021, 7:22am UTC](https://discuss.elastic.co/t/esrally-to-test-resiliency-of-the-server/260150/3 "2021-02-18T07:22:50Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
