# ES alerting mechanism for failure scenarios, high latency situations

**URL:** <https://discuss.elastic.co/t/es-alerting-mechanism-for-failure-scenarios-high-latency-situations/16206>\
**Category:** Elasticsearch\
**Created:** [March 6, 2014, 6:24pm UTC](https://discuss.elastic.co/t/es-alerting-mechanism-for-failure-scenarios-high-latency-situations/16206 "2014-03-06T18:24:33Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![T\_Vinod\_Gupta](https://avatars.discourse-cdn.com/v4/letter/t/fbc32d/32.png) [@T\_Vinod\_Gupta](https://discuss.elastic.co/u/T_Vinod_Gupta)\
**Post date:** [March 6, 2014, 6:24pm UTC](https://discuss.elastic.co/t/es-alerting-mechanism-for-failure-scenarios-high-latency-situations/16206/1 "2014-03-06T18:24:33Z")

</div>

is there a plugin or api support for monitoring ES key metrics and alerting  
the dev ops about situations when some node in a cluster fails or there is  
a spike in latency due to whatever reason?

what are the best practices here and what do people usually do?

thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAHau4yv9L%2B5zXtDQcNKmK-b\_30Q2MdrTtPjHUWsDYKEgFX8hnQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAHau4yv9L%2B5zXtDQcNKmK-b_30Q2MdrTtPjHUWsDYKEgFX8hnQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [March 7, 2014, 5:44am UTC](https://discuss.elastic.co/t/es-alerting-mechanism-for-failure-scenarios-high-latency-situations/16206/2 "2014-03-07T05:44:07Z")

</div>

Hi,

We use our own SPM for Elasticsearch. It has classic threshold-based  
alerts as well as alerts based on automatic anomaly detection -

> **[Metrics & Log Alerts: Anomalies, Thresholds, Heartbeats](https://sematext.com/alerts)**
>
> Utilize different types of Alerts or Anomaly Detection and connect them to Slack, PagerDuty, and other ChatOps tools, email, mobile push notifications.

. It's a SaaS, not a plugin, but maybe it would work for you.

## Otis

Performance Monitoring \* Log Analytics \* Search Analytics  
Solr & Elasticsearch Support \* [http://sematext.com/](http://sematext.com/)

On Thursday, March 6, 2014 1:24:33 PM UTC-5, T Vinod Gupta wrote:

> is there a plugin or api support for monitoring ES key metrics and  
> alerting the dev ops about situations when some node in a cluster fails or  
> there is a spike in latency due to whatever reason?
> 
> what are the best practices here and what do people usually do?
> 
> thanks

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/41c27e2b-5031-44f2-9d8d-4130d451446e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/41c27e2b-5031-44f2-9d8d-4130d451446e%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [March 7, 2014, 10:45pm UTC](https://discuss.elastic.co/t/es-alerting-mechanism-for-failure-scenarios-high-latency-situations/16206/3 "2014-03-07T22:45:14Z")

</div>

This is a very good point, I'm thinking about this for years.

Node failures should be easy to monitor by OS services. But latency spikes  
are totally different.

It is a very, very hard job to measure anomalies in latency correctly. Just  
consider the skews of wrong programming, or of the hostile environments  
JVMs do run in (clocks, OSes, VMs, ...) If anomalies are detected wrongly,  
no or false alerts are emitted, and all of the effort would lead to  
annoyance or frustration.

Lately I read about Gil Tene's LatencyUtils

> **[GitHub - LatencyUtils/LatencyUtils: Utilities for latency measurement and...](https://github.com/LatencyUtils/LatencyUtils)**
>
> Utilities for latency measurement and reporting. Contribute to LatencyUtils/LatencyUtils development by creating an account on GitHub.

[https://groups.google.com/forum/#!topic/mechanical-sympathy/oZSv5QnpAYs](https://groups.google.com/forum/#!topic/mechanical-sympathy/oZSv5QnpAYs)

which I find a promising tool to measure anomalies in histograms.

Some of this might be possible to get implemented by an ES plugin, but I  
haven't tried LatencyUtils yet, and how it can be connected to ES metrics  
is still open to me.

Jörg

On Thu, Mar 6, 2014 at 7:24 PM, T Vinod Gupta [tvinod@readypulse.com](mailto:tvinod@readypulse.com) wrote:

> is there a plugin or api support for monitoring ES key metrics and  
> alerting the dev ops about situations when some node in a cluster fails or  
> there is a spike in latency due to whatever reason?
> 
> what are the best practices here and what do people usually do?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGXNqJkF5uL2oCKmBsHYqQJxFdxUrW%2BF0maVSJupOGupQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGXNqJkF5uL2oCKmBsHYqQJxFdxUrW%2BF0maVSJupOGupQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![T\_Vinod\_Gupta](https://avatars.discourse-cdn.com/v4/letter/t/fbc32d/32.png) [@T\_Vinod\_Gupta](https://discuss.elastic.co/u/T_Vinod_Gupta)\
**Post date:** [March 7, 2014, 10:49pm UTC](https://discuss.elastic.co/t/es-alerting-mechanism-for-failure-scenarios-high-latency-situations/16206/4 "2014-03-07T22:49:39Z")

</div>

i was playing around with marvel on a test machine and it is clear that a  
lot of thought, effort and time has gone into building it. it is super. but  
what will really take it to the next level is alerts - you can configure  
certain kinds of events to trigger an alert. and then have rules around  
latency spikes. i agree that full automation can lead to false triggers and  
annoyance. but if you make it like aws cloudwatch triggers where you say  
that if the cluster/node is in a certain state (e.g. search latency \> 1s  
for a period of 10 min), then trigger.

thanks

On Fri, Mar 7, 2014 at 2:45 PM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<[joergprante@gmail.com](mailto:joergprante@gmail.com)

> wrote:

> This is a very good point, I'm thinking about this for years.
> 
> Node failures should be easy to monitor by OS services. But latency spikes  
> are totally different.
> 
> It is a very, very hard job to measure anomalies in latency correctly.  
> Just consider the skews of wrong programming, or of the hostile  
> environments JVMs do run in (clocks, OSes, VMs, ...) If anomalies are  
> detected wrongly, no or false alerts are emitted, and all of the effort  
> would lead to annoyance or frustration.
> 
> Lately I read about Gil Tene's LatencyUtils
> 
> [GitHub - LatencyUtils/LatencyUtils: Utilities for latency measurement and reporting](https://github.com/LatencyUtils/LatencyUtils)
> 
> [Redirecting to Google Groups](https://groups.google.com/forum/#!topic/mechanical-sympathy/oZSv5QnpAYs)
> 
> which I find a promising tool to measure anomalies in histograms.
> 
> Some of this might be possible to get implemented by an ES plugin, but I  
> haven't tried LatencyUtils yet, and how it can be connected to ES metrics  
> is still open to me.
> 
> Jörg
> 
> On Thu, Mar 6, 2014 at 7:24 PM, T Vinod Gupta [tvinod@readypulse.com](mailto:tvinod@readypulse.com)wrote:
> 
> > is there a plugin or api support for monitoring ES key metrics and  
> > alerting the dev ops about situations when some node in a cluster fails or  
> > there is a spike in latency due to whatever reason?
> > 
> > what are the best practices here and what do people usually do?
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGXNqJkF5uL2oCKmBsHYqQJxFdxUrW%2BF0maVSJupOGupQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGXNqJkF5uL2oCKmBsHYqQJxFdxUrW%2BF0maVSJupOGupQ%40mail.gmail.com)[https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGXNqJkF5uL2oCKmBsHYqQJxFdxUrW%2BF0maVSJupOGupQ%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGXNqJkF5uL2oCKmBsHYqQJxFdxUrW%2BF0maVSJupOGupQ%40mail.gmail.com?utm_medium=email&utm_source=footer)  
> > .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAHau4yvm0HKXK%2Bhuvejq%2B0WT4TrWEJYMTnCnYsSWWaipq828ag%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAHau4yvm0HKXK%2Bhuvejq%2B0WT4TrWEJYMTnCnYsSWWaipq828ag%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:44am UTC](https://discuss.elastic.co/t/es-alerting-mechanism-for-failure-scenarios-high-latency-situations/16206/5 "2017-07-06T01:44:45Z")

</div>


