# Advanced Watcher - If Failed % is greater than some defined threshold value

**URL:** https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727
**Category:** Kibana
**Tags:** elastic-stack-alerting
**Created:** [March 25, 2022, 8:31pm UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727 "2022-03-25T20:31:13Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![Freddy.Raj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/freddy.raj/32/77781_2.png) [@Freddy.Raj](https://discuss.elastic.co/u/Freddy.Raj)
#### Post date: [March 25, 2022, 8:31pm UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/1 "2022-03-25T20:31:13Z")

</div>

Hi All,  
I need some help to achieve below monitoring condition to enable watcher alert.

1. We have a domain named "XYZ" as a value to one of the term field - "main.domain"
2. Under this domain "XYZ", we see logs for several recipients domain - [abc.com](http://abc.com), [def.com](http://def.com), [ghi.com](http://ghi.com) etc.,
3. There are 3 types of message status - Sent, Sending and Failed

At any given time, messages will be either in Sent, Sending or Failed status for the respective recipient domains of the main domain named "XYZ".

Monitoring condition: at any time, if "Failed" message type for any recipient domain ([abc.com](http://abc.com), [def.com](http://def.com), [ghi.com](http://ghi.com) etc.,) crosses greater than 5% then an alert need to be triggered.

Currently, we have filtered the main domain named "XYZ" for 5 last minutes and pulled 2 set of aggregations.  
a. for all recipient domains along with all 3 message status types  
b. for all recipient domains and only with Failed message status type

Now need to figure out a way where in, i can set up alert for below example.  
Lets say, in last 5 mins, for the main domain named "XYZ" and for recipient domain "[abc.com](http://abc.com)", if failed messages go greater than 5% then an alert need to be triggered.

Any suggestions or advise would be of great help. Thank you all in advance.

---

<div class="post-metadata">

### Author: ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)
#### Post date: [March 28, 2022, 7:49am UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/2 "2022-03-28T07:49:20Z")

</div>

Hey,

it would help what you already tried in order to get an overview. To me this sounds, as if you need to compare two values returned by an aggregation response within your domain. This can be done using a script condition.

The most important thing here is probably not writing the watch, but the correct query that returns all the data required to do the parsing in the condition, you should focus on that first, before putting anything in a watch.

Hope that helps as a start.

--Alex

---

<div class="post-metadata">

### Author: ![Freddy.Raj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/freddy.raj/32/77781_2.png) [@Freddy.Raj](https://discuss.elastic.co/u/Freddy.Raj)
#### Post date: [March 28, 2022, 10:37am UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/3 "2022-03-28T10:37:08Z")

</div>

Thank you Alex. @spinscale

I was working on creating scripts that would generate desired results. Didn't put anything into watcher section 🙂 . However, below are my queries tried so far.

1. _ **Removed indent to save space here** _  
This gives me result of total messages in all status broken with each message type and recipient domain for the main domain.  
And if use "value\_count" instead of "terms" for message type aggs then i get total value of all 3 message types under each recipient domain.

{"aggs":{"RecipientDomain":{"terms":{"field":"recipientDomain.keyword"},"aggs":{"MessageTypes":{"terms":{"field":"messageType.keyword"}}}}},"size":0,"\_source":{"excludes":},  
"script\_fields":{},"docvalue\_fields":[{"field":"@timestamp","format":"date\_time"}],"query":{"bool":{"must":[{"match\_all":{}},{"match\_phrase":{"sendingDomain":  
{"query":"[main.domain.com](http://main.domain.com)"}}},{"range":{"@timestamp":{"format":"strict\_date\_optional\_time","gte":"2022-03-24T15:49:42.890Z","lte":"2022-03-24T15:55:42.891Z"}}}]}}}

My result for this is:  
"aggregations" : {  
"recipient\_domain" : {  
"doc\_count\_error\_upper\_bound" : 12,  
"sum\_other\_doc\_count" : 576,  
"buckets" : [  
{  
"key" : "[ABC.com](http://ABC.com)",  
"doc\_count" : 3479,  
"Delivered\_Message" : {  
"buckets" : {  
"messageType:Delivered" : {  
"doc\_count" : 3420  
}  
}  
},  
"Transient\_Message" : {  
"buckets" : {  
"messageType:Transient" : {  
"doc\_count" : 32  
}  
}  
},  
"Failed\_Message" : {  
"buckets" : {  
"messageType:Failed" : {  
"doc\_count" : 27  
}  
}  
}  
},

How can i add script condition to calculate % of failures?

1. _ **Removed indent to save space here** _  
This case, i get results again on total messages under each domain and under each status.

{"size":0,"query":{"bool":{"must":[{"match\_phrase":{"sendingDomain":{"query":"[main.domain.com](http://main.domain.com)"}}},{"range":{"@timestamp":{"gte":"2022-03-25T19:44:09.034Z",  
"lte":"2022-03-25T19:50:09.034Z"}}}],"must\_not":}},"aggs":{"recipient\_domain":{"terms":{"field":"recipientDomain.keyword","min\_doc\_count":20,"size":10,"order":{"\_count":"desc"}},  
"aggs":  
{"Failed\_Message":{"filters":{"filters":{"messageType:Failed":{"bool":{"must":,"filter":[{"bool":{"should":[{"match":{"messageType":"Failed"}}],"minimum\_should\_match":1}}],"should":,"must\_not":}}}}},  
"Delivered\_Message":{"filters":{"filters":{"messageType:Delivered":{"bool":{"must":,"filter":[{"bool":{"should":[{"match":{"messageType":"Delivered"}}],"minimum\_should\_match":1}}],"should":,"must\_not":}}}}},  
"Transient\_Message":{"filters":{"filters":{"messageType:Transient":{"bool":{"must":,"filter":[{"bool":{"should":[{"match":{"messageType":"Transient"}}],"minimum\_should\_match":1}}],"should":,"must\_not":}}}}},

To this when i add script condition to compute % calculation, i get same results as above but nothing for below condition. As in, i get results for total records in each domain and each status but computing % script part does not show any reference at all. What am i missing here?  
I test these in Kibana DevTools.

"ComputePercentage":{"bucket\_script":{"buckets\_path":{"TD":"Delivered\_Message.doc\_count","TT":"Transient\_Message.doc\_count","TF":"Failed\_Message.doc\_count"},  
"script":"(params.TF / (params.TD + params.TT + params.TF)) \* 100"}}}}}}

My result for this is:  
"aggregations" : {  
"RD" : {  
"doc\_count\_error\_upper\_bound" : 15,  
"sum\_other\_doc\_count" : 882,  
"buckets" : [  
{  
"key" : "[ABC.com](http://ABC.com)",  
"doc\_count" : 3651,  
"MT" : {  
"doc\_count\_error\_upper\_bound" : 0,  
"sum\_other\_doc\_count" : 0,  
"buckets" : [  
{  
"key" : "Delivered",  
"doc\_count" : 3573  
},  
{  
"key" : "Transient",  
"doc\_count" : 51  
},  
{  
"key" : "Failed",  
"doc\_count" : 27  
}  
]  
}  
},

How can i add script condition to calculate % of failures?

---

<div class="post-metadata">

### Author: ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)
#### Post date: [March 29, 2022, 8:16am UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/4 "2022-03-29T08:16:09Z")

</div>

Hey,

so basically you have to loop through `ctx.payload.aggregations.RD.buckets` and for each element you have to first extract the bucket `doc_count` with `key == Delivered` in `ctx.payload.aggregations.RD.buckets[0]` and the same for Transient/Failed. When you have both doc\_counts you can divide them.

I think you should find some help in the watcher examples, even though they are already a bit older. See [examples/Alerting/Sample Watches at master · elastic/examples · GitHub](https://github.com/elastic/examples/tree/master/Alerting/Sample%20Watches)

--Alex

---

<div class="post-metadata">

### Author: ![Freddy.Raj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/freddy.raj/32/77781_2.png) [@Freddy.Raj](https://discuss.elastic.co/u/Freddy.Raj)
#### Post date: [March 29, 2022, 12:47pm UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/5 "2022-03-29T12:47:11Z")

</div>

Thank you for your prompt response Alex @spinscale  
Was able to use the ctx.payload aggregations and set up watcher alert too.  
I know receive alert when any domain crosses threshold of 5% failure from overall messages.

However, if there are multiple domain failures at the same time, how can i show all failures within same email alert. Rather than sending 1 alert each for each failed domains?

---

<div class="post-metadata">

### Author: ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)
#### Post date: [March 30, 2022, 7:51am UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/6 "2022-03-30T07:51:06Z")

</div>

If your search query contains the data for all failures, then you would explicitely add a [foreach](https://www.elastic.co/guide/en/elasticsearch/reference/8.1/action-foreach.html) part to your action, to run this for each failure. If you do not do that you should end up with a single alert? Or do you have an own watch for each customer?

---

<div class="post-metadata">

### Author: ![Freddy.Raj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/freddy.raj/32/77781_2.png) [@Freddy.Raj](https://discuss.elastic.co/u/Freddy.Raj)
#### Post date: [March 30, 2022, 8:18am UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/7 "2022-03-30T08:18:26Z")

</div>

This will be just 1 watcher alert for all recipient domains failing at a given time.  
Now we see multiple alerts triggering at a time if there are more than 1 recipient domain failing with 5% rate.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [April 27, 2022, 8:18am UTC](https://discuss.elastic.co/t/advanced-watcher-if-failed-is-greater-than-some-defined-threshold-value/300727/8 "2022-04-27T08:18:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
