# Machine Learning module is triggering alerts when there is no anomaly

**URL:** <https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [April 29, 2019, 6:23pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974 "2019-04-29T18:23:22Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [April 29, 2019, 6:23pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/1 "2019-04-29T18:23:22Z")

</div>

We are using the Machine Learning module to identify an unusual rate of firewall logs by host and by severity (e.g., warning, informational, error).

We keep receiving Watch alerts that look like this:

 ![ML_Watcher_Alert_-josh_ablett_adeliarisk_com-_Adelia_Risk_Mail](https://us1.discourse-cdn.com/elastic/original/3X/b/7/b7f58fd81861b78cb679ab0adbe29a4857885c18.png)

But when we click on the link the alert, there's nothing there:

 ![sonicwall-anomalies-by-host-and-severity_-_Kibana](https://us1.discourse-cdn.com/elastic/original/3X/7/5/753689c9674c68f24dbc545cbe31e75921f7dab3.png)

When we look at the timeframe in Single Metric Viewer, there are clearly no anomalies (the email alert was triggered on 3/28/19 at 7pm eastern).

 ![sonicwall-anomalies-by-host-and-severity_-_Kibana2](https://us1.discourse-cdn.com/elastic/original/3X/c/1/c14658a0dfd8daf6361952c20703c98619be07d7.png)

Also, our Watch is configured to only trigger on anomaly scores over 90, so it should only email when something highly unusual happens:

 ![Kibana](https://us1.discourse-cdn.com/elastic/original/3X/6/5/6596c1693f3e7517e6881af4d09f937a298ac928.png)

Anyone have any idea why a Watcher would trigger an alert when there's nothing there? Anomaly detection is kind of the whole point of our use of Elastic as a SIEM, and if it can't be trusted, we need to look for other solutions.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [April 29, 2019, 11:10pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/2 "2019-04-29T23:10:23Z")

</div>

Seems similar to: [ML alerts triggering on interim result](https://discuss.elastic.co/t/ml-alerts-triggering-on-interim-result/158408)

See that thread for the workaround and the related bug report at: [https://github.com/elastic/ml-cpp/issues/324](https://github.com/elastic/ml-cpp/issues/324)

The above was ultimately fixed in v6.6.2

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [April 30, 2019, 1:50pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/3 "2019-04-30T13:50:52Z")

</div>

Thanks, Rich.

I'm on v6.6.2. Here's a screenshot:

 ![dec576__rsyslog-_Elastic_Cloud](https://us1.discourse-cdn.com/elastic/original/3X/7/9/79bdf66fa44cd52adf8d7feaa603fd8ca713bae8.png)

Can you suggest anything else we can try?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [April 30, 2019, 2:20pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/4 "2019-04-30T14:20:56Z")

</div>

Hmm...the behavior you describe is so similar to the issue pointed out - perhaps I'm not correct in which version the fix was back-ported.

Does the problem go away if you modify the Watch to ignore interim results? Just add a:

```auto
                  { "term" : { "is_interim" : "false"}}

```

to the query to `.ml-anomalies-*` that the Watch is making.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 2, 2019, 3:23pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/5 "2019-05-02T15:23:25Z")

</div>

Update?

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [May 2, 2019, 3:37pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/6 "2019-05-02T15:37:12Z")

</div>

Hi Rich - sorry for the delayed response.

By "add it to the query", should it be nested in a filter statement? Or should it be at the same level as the ""bool" or "filter" that are auto-generated by the system when it creates a Watch?

Also, how would I be able to tell if this works? Is there a way to replay a Watch against historical data, since this email triggered was back in March?

Thanks!

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 3, 2019, 1:01am UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/7 "2019-05-03T01:01:40Z")

</div>

Hi Josh - well, it would be nested in the filter part of the statement. But, your subsequent question about re-running the Watch against historical data has got me reconsidering the nature of the problem you're reporting.

So, help me understand how often this is happening? I originally thought it was happening routinely (as in every bucket\_span) and that my suggested workaround would alleviate that.

So, if that's not the case, then how often are you seeing this? Is every alert email showing a "ghost" anomaly that you cannot link to or do some of the alerts link to "real" ones?

Also, it might be handy if you create a kibana index pattern for `.ml-anomalies-*`. Once you do that, you can use the Discover tab to search for the anomaly records. For example:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/f/ff5eed3000b6ae81513c90dd48803b35c3b2b5fe.png)

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [May 3, 2019, 6:25pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/8 "2019-05-03T18:25:49Z")

</div>

Hi Rich,

It's definitely not happening every bucket\_span (which is set to 60 minutes).

It's also not on a regular cycle. We get email alerts every few days, and most of the time when I click on the link in the Watcher, what I see in the Anomaly Explorer view doesn't have any connection to what was in the email.

For example, here's the most recent email from April 27 (six days ago at this point):  
 ![ML_Watcher_Alert_-josh_ablett_adeliarisk_com-_Adelia_Risk_Mail1](https://us1.discourse-cdn.com/elastic/original/3X/c/2/c2254f279e0a49add828e9e8c2cd065d7ab73ebe.png)

But when clicking on the link, I don't see any anomaly in that time window even close to an anomaly score of 94. There were some anomalies (highlighted by red arrows), but none that should have reached the level of severity to trigger the Watch.

 ![sonicwall-anomalies-by-host-and-severity_-_Kibana1](https://us1.discourse-cdn.com/elastic/original/3X/c/2/c249dfcf9e494fea2284f16bcea5c86fda584084.png)

Would love some suggestions on where to go next. We will create the index pattern you suggest.

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [May 3, 2019, 6:35pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/9 "2019-05-03T18:35:39Z")

</div>

The index pattern is making things even more confusing.

Here's the highest anomaly scores:  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/1/6/160ab3e23dba2f369b91a9186d96e37e15a15b4f.png)

And if I look at the Anomaly timeline, it seems to match up:  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/f/5f7c0780722f1fd13a0076df7cbce3e908930e9f.png)

But when I click on that hour in the Anomaly timeline, no anomalies with a score of 93.994 show up on the Anomaly Explorer page. Here is where I would expect to see the anomaly:  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/e/fe54f6bf1c3fe4662563ba43feb2b70e491925f6.png)

But as you can see, there's on an anomaly with a score of 54 showing up for that hour. And there's only one, not the two that we can see in the new index pattern.

Also, every alert with a score higher than 85 has interim set to false --- here's a screenshot.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/7/f787b12ba65f710787e554fd212114c839a45a2a.png)

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 3, 2019, 7:58pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/10 "2019-05-03T19:58:21Z")

</div>

Ok, now this seems like this is a timezone issue. Hmmm....Are you in a timezone that is exactly the number of hours behind UTC as the anomalies that you're finding in `ml-anomalies-*`? What time zone are you in and what timezone is your kibana set to use?

Also when looking in `ml-anomalies-*`, you should look at a field called `initial_anomaly_score` as that is what the score was at the time the result was first written. The `anomaly_score`can be revised at a later time, but `initial_anomaly_score` will never change once written.

(by the way - you probably can remove the `is_interim` filter on your watch search since it seems like you're not plagued by that other bug. The `is_interim` is only briefly set to true when the latest bucket is still collecting data - within the current hour in your case)

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [May 6, 2019, 1:11pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/11 "2019-05-06T13:11:29Z")

</div>

Hi Rich - we're in the eastern time zone (UTC-5). Timezone is set to "browser."

It seems that the data has changed yet again, which is really scary.

Now when I look in `ml-anomalies-*`, there is no value anywhere near the anomaly score of 94 that was in the email from April 27, 2019, or the 93.994 that was in the `ml-anomalies-*` screenshot that I posted here on May 3 based on data from April 27, 2019.

When I looked in `ml-anomalies-*` today (May 6), here are the highest scores for `anomaly_score`, sorted in the descending order:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/4/3/439e8c1e4f46c14e5073af31cbb40c28d12ba84c.png)

And here are the highest scores for `initial_anomaly_score`, sorted in descending order:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/2/2/22d6c9a465d1d317f6cb35ef4775086b3439d2e8.png)

Also - not entirely following your logic about timezones. I agree that this would explain a time of 13:00 in the email alert and a time of 10:00 in `ml-anomalies-*`, but that doesn't seem to explain why there are different numbers showing up in different parts of the user interface.

Any ideas?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 6, 2019, 2:00pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/12 "2019-05-06T14:00:13Z")

</div>

Hmm... I cannot immediately explain the anomaly score discrepancy unless your watch was really returning and alerting upon the `record_score` and not the `anomaly_score`. There are several different kinds of `result_type` documents in `.ml-anomalies-*`. They include:

- `bucket` - which is an aggregation/summary score of the time period of the bucket span
- `record` - which is the detailed score of an individual occurrence of an anomaly inside a bucket
- `influencer` - an entity-centric version of the scoring

I now suspect that the watch is returning something different than the `anomaly_score` that you're comaring against. You could validate this by either:

1. Posting the full code of your watch for me to see or
2. Looking at the watch code yourself and working out what value it is returning and comparing that against a look in `.ml-anomalies-*` for one of the other `result_type` documents

Also, it is possible that the logic in the watch is incorrect with respect to the time zone. It is likely that there is some scripting that is in the watch to buld the "start" and "end" time for the email link that says "Click here to open in Anomaly Explorer". It is possible that the logic used in building that link is pointing you to an incorrect window of time.

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [May 6, 2019, 2:20pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/13 "2019-05-06T14:20:19Z")

</div>

Hi Rich,

Sorry, I threw us a red herring here. We have two similarly named ML job ID's, and I was confusing them in my most recent post. Please disregard my comments about the numbers not agreeing.

So, to recap where we are, email alert came in April 27 at 9:12 am eastern time:

```
Job: sonicwall-anomalies-by-host-and-severity 
Time: 2019-04-27T13:00:00.000Z 
Anomaly score: 94 

```

Which agrees with what I see in `ml-anomalies-*`:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/3/f361f1e72dc520b564c33587a8206fe9b07cd447.png)

And also agrees with the summary chart on the top of the Anomaly Explorer page:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/b/7bc3a3b6b02cdc18e3966c13396041f2c678801f.png)

But that doesn't agree at all with what's at the bottom of the Anomaly Explorer page:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/1/5/159e33504f97ed6494b1488df54d534d35510812.png)

So I think it's really just that last part that's the problem.

Thanks,  
Josh

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 6, 2019, 2:23pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/14 "2019-05-06T14:23:27Z")

</div>

Hi Josh - glad you caught that discrepancy!

As for the last part (the data that is in the table) - those are the `record_score` entries. So, you need to compare what you see on the screen there with entries in `.ml-anomalies-*` that have `result_type:record` and look at `record_score` and `initial_record_score`

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 7, 2019, 5:30pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/15 "2019-05-07T17:30:43Z")

</div>

update?

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [May 7, 2019, 6:09pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/16 "2019-05-07T18:09:43Z")

</div>

Hi Rich - thanks for nagging me on this, missed yesterday's notification.

In `.ml-anomalies-*`, when I filter down for `result_type: "record"` and `job_id: "sonicwall-anomalies-by-host-and-severity`, I see numbers that match with the bottom of the Anomaly Explorer page:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/f/6fe7f3a91eb64796476cb304a13be177f128b0e5.png)

Any tips or documentation to help me to understand, in practical terms, the difference between the "record" score and the "bucket" score? I thought all of the anomaly detection is comparing the count in each bucket, so I'm not sure what the "record" score is supposed to be telling me?

I found this article [https://www.elastic.co/blog/machine-learning-anomaly-scoring-elasticsearch-how-it-works](https://www.elastic.co/blog/machine-learning-anomaly-scoring-elasticsearch-how-it-works), but I'm still a bit confused on what constitutes an anomaly on an individual records when the machine learning job is being trained to look for an unusual rate of entries.

Thanks!

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 8, 2019, 3:39pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/17 "2019-05-08T15:39:53Z")

</div>

The [results API docs](https://www.elastic.co/guide/en/elasticsearch/reference/7.0/ml-apis.html#ml-api-result-endpoint) detail what's contained in the different results types, but the blog you mentioned is also a good source. I have written a [full reference book on Elastic ML](https://www.packtpub.com/big-data-and-business-intelligence/machine-learning-elastic-stack), so if that's interesting you can look into that as well.

To understand the difference between bucket-level anomalies and record-level anomalies, you need to consider the case that for some ML jobs, there are possibly many entities being modeled (if the job is split) and technically there can be multiple "detectors" configured per job. Therefore, any one instance of an anomaly - for either a particular detector and/or a particular entity would result in a results `record`. The `bucket` level score is the aggregation of all anomalies in that bucket. There would be a difference, for example, if 1 out of 100 hosts were unusual in the last bucket versus 99 out of 100 hosts. The latter would obviously command a higher bucket score.

Hope that helps

---

<div class="post-metadata">

**Author:** ![Josh\_A](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/josh_a/32/42035_2.png) [@Josh\_A](https://discuss.elastic.co/u/Josh_A)\
**Post date:** [May 9, 2019, 12:15pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/18 "2019-05-09T12:15:38Z")

</div>

Hi Rich - it helps conceptually, but I'm still struggling a bit with how to use this in practice.

This specific machine learning job is simply looking for an anomalous count of firewall logs, with a partition of "host.keyword" (which is the IP address of the sending firewall). Our goal is to get an alert when any individual host sends either an unusually high or unusually low number of records in any given hour, as compared to other hours. The bucket\_span is set to 60 minutes.

Given that goal, should we be paying attention to the record-level anomalies or the bucket-level anomalies? And also, in this use case, what would make an individual record anomalous, since each record is simply a firewall log entry?

Thanks!

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 9, 2019, 1:27pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/19 "2019-05-09T13:27:54Z")

</div>

Hey Josh,

When you have a job with splits (like you do, splitting on `host.keyword`) AND your goal is to get an alert when any individual host is unusual, then yes - you should be working with `result_type:record` entries in the results index. What this means, of course, is that if there is a widespread problem affecting many (let's say 100) hosts within the same bucket span, then you will get 100 anomaly records, one for each host. Conversely, if you were only querying at the bucket level, you would only get 1 result doc per bucket\_span and you might not know which host(s) were unusual in that bucket, unless they were prominent `influencers`.

To see an example of how a record watch looks, see here: [https://gist.github.com/richcollier/1c2b8161286bdca6c553859f28d3d66d](https://gist.github.com/richcollier/1c2b8161286bdca6c553859f28d3d66d)

In this watch, the results index is queried for any occurrence of any entity (in my case, this was a job split on an `airline` code) with a minimum `record_score` of `90`. (Also note that this watch isn't built for real-time, as it searches over a multi-year history of results).

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [May 9, 2019, 1:31pm UTC](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974/20 "2019-05-09T13:31:12Z")

</div>

One small detail that this example watch omits. Since the default `size` of a query in elasticsearch returns 10 `hits` - if you expect more than 10 results, you should add a `size` parameter to the query. For example:

```auto
...
    "input": {
      "search": {
        "request": {
          "indices": [
            ".ml-anomalies-*"
          ],
          "body": {
            "size" : 1000,
            "query": {
              "bool": {
                "filter": [
...

```

[Next page](https://discuss.elastic.co/t/machine-learning-module-is-triggering-alerts-when-there-is-no-anomaly/178974.md?page=2)
