# Chaos engineering in Elasticsearch

**URL:** <https://discuss.elastic.co/t/chaos-engineering-in-elasticsearch/158293>\
**Category:** Elasticsearch\
**Created:** [November 27, 2018, 8:21am UTC](https://discuss.elastic.co/t/chaos-engineering-in-elasticsearch/158293 "2018-11-27T08:21:48Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![sumanthkumarc](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumanthkumarc/32/39338_2.png) [@sumanthkumarc](https://discuss.elastic.co/u/sumanthkumarc)\
**Post date:** [November 27, 2018, 8:21am UTC](https://discuss.elastic.co/t/chaos-engineering-in-elasticsearch/158293/1 "2018-11-27T08:21:48Z")

</div>

How can i break Elasticsearch in controlled manner, so that i can test its resiliency in my environment. I do understand that we could tamper with the external factors like memory , network, cpu etc, but i was just wondering if something is built into/along with elasticsearch as a plugin for development, which help us in testing this. If nothing of this sort exists, how can i create one such thing? possibly in a form i could contribute back to community so that others can re-use it??

Any pointers are appreciated.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 27, 2018, 10:35am UTC](https://discuss.elastic.co/t/chaos-engineering-in-elasticsearch/158293/2 "2018-11-27T10:35:22Z")

</div>

The class `org.elasticsearch.discovery.AbstractDisruptionTestCase` is the base class for a number of test suites that build a cluster of Elasticsearch nodes (running in a single JVM) and test its behaviour under various network disruptions.

There are other tests in the test suite that apply other kinds of disruption too. A good starting point is to look at the implementations of `org.elasticsearch.test.disruption.ServiceDisruptionScheme` and see where they are used.

We have also been running some tests, particularly benchmarks using [Rally](https://esrally.readthedocs.io/en/stable/), in low-memory environments, and improving how Elasticsearch deals with memory pressure as a result.

I have also done some fruitful experimentation with a cluster running on a single machine in Docker containers, applying network disruptions using `iptables` rules.

Mike McCandless also [wrote up some of the testing he did](http://blog.mikemccandless.com/2014/04/testing-lucenes-index-durability-after.html) on Lucene's safety in the event of power loss.

I hope these are useful pointers.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 25, 2018, 10:41am UTC](https://discuss.elastic.co/t/chaos-engineering-in-elasticsearch/158293/3 "2018-12-25T10:41:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
