# Setting document limits for Machine Learning anomalies

**URL:** <https://discuss.elastic.co/t/setting-document-limits-for-machine-learning-anomalies/206213>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [November 1, 2019, 8:14pm UTC](https://discuss.elastic.co/t/setting-document-limits-for-machine-learning-anomalies/206213 "2019-11-01T20:14:46Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![elasticitous](https://avatars.discourse-cdn.com/v4/letter/e/8e7dd6/32.png) [@elasticitous](https://discuss.elastic.co/u/elasticitous)\
**Post date:** [November 1, 2019, 8:14pm UTC](https://discuss.elastic.co/t/setting-document-limits-for-machine-learning-anomalies/206213/1 "2019-11-01T20:14:46Z")

</div>

I've successfully created some Population machine learning jobs but I'm seeing a lot of false positive anomalies generated from members of the population with too small a sample size for those metrics to converge to anything meaningful.

For instance, suppose an average value of 10 over 100 documents is an anomaly. But one document with a value of 10 isn't even though that's the same average. I want a member of the population returned as an anomaly only if its aggregate metric is high/low, but also if its total document count in the bucket span is high enough to care.

Making the bucket span longer isn't an option.

Is there an easy way of limiting the ML job to a minimum document count to "count?"

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [November 4, 2019, 2:56pm UTC](https://discuss.elastic.co/t/setting-document-limits-for-machine-learning-anomalies/206213/2 "2019-11-04T14:56:44Z")

</div>

As soon as you say "if its total document count is high enough to care" - you're defining a rule. That's perfectly fine, but just so you know - you'll have to manually define what "high enough" means.

The [Custom Rules](https://www.elastic.co/guide/en/elastic-stack-overview/current/ml-rules.html) part of the ML job allows you to override the definition of what is considered anomalous. However, it is only for what's being measured in the detector function. So, if you are using the `mean` function, you can control whether or not the anomaly on `mean` is "high enough" or "low enough". Same with the `count` function if you are measuring the event rate as a function of time. But, you cannot control a different aspect that is _not_ the detector function. In other words, you cannot control anomalies on the `mean` depending on the count of documents.

In order to accomplish this, you'd need to put that logic in the alerting. When creating a Watch, you would use a "chain input", which allows more than one search to define the alert. The first search would be to look in `.ml-anomalies-*` for the anomaly in the `mean` (and locate the offending entity as the influencer), then use a second search to determine the number of docs that this entity has in that timeframe. The `condition` in the watch is where you'd put the rule/threshold that you define what would be "high enough" for the doc count. If those two things are met, then the alert can notify you.

An example of a chain input watch is here: [A watch alert example based on two different searches using CHAIN input and Painless script condition](https://discuss.elastic.co/t/a-watch-alert-example-based-on-two-different-searches-using-chain-input-and-painless-script-condition/113829)

Another example is here: [https://gist.github.com/richcollier/7e5603c366b9fcece6f1a8b1b3cf4d3f](https://gist.github.com/richcollier/7e5603c366b9fcece6f1a8b1b3cf4d3f)

---

<div class="post-metadata">

**Author:** ![elasticitous](https://avatars.discourse-cdn.com/v4/letter/e/8e7dd6/32.png) [@elasticitous](https://discuss.elastic.co/u/elasticitous)\
**Post date:** [November 4, 2019, 7:11pm UTC](https://discuss.elastic.co/t/setting-document-limits-for-machine-learning-anomalies/206213/3 "2019-11-04T19:11:16Z")

</div>

> [@richcollier](#):
>
> In other words, you cannot control anomalies on the `mean` depending on the count of documents.

Thanks for confirming this and the idea to use Watcher as a workaround but I feel you should be able to do this in the job config itself.

My datafeed already specifies a "summary\_count\_field\_name" value of "doc\_count" which means the datafeed knows the total number of documents in any given bucket despite what the detector metric happens to be.

I can see why this would get problematic if you wanted to control one metric you are calculating, with another you're not, but every aggregation bucket already has a total doc count returned along with the aggregated value.

There's no way to use this somehow? The only workaround is to have the anomalies (which I know aren't anomalies) reported incorrectly, and then configure Watcher to ignore them?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [November 7, 2019, 4:02pm UTC](https://discuss.elastic.co/t/setting-document-limits-for-machine-learning-anomalies/206213/4 "2019-11-07T16:02:11Z")

</div>

You could make an aggregate value using a [script field](https://www.elastic.co/guide/en/elastic-stack-overview/current/ml-configuring-transform.html) or a [bucket\_script aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-pipeline-bucket-script-aggregation.html) that is the combination of the document count and the metric (i.e. doc count \* metricvalue). Then have that value modeled over time by ML.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 5, 2019, 4:02pm UTC](https://discuss.elastic.co/t/setting-document-limits-for-machine-learning-anomalies/206213/5 "2019-12-05T16:02:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
