# ML Anomaly Job with exclude\_frequent option

**URL:** <https://discuss.elastic.co/t/ml-anomaly-job-with-exclude-frequent-option/348047>\
**Category:** Elasticsearch\
**Created:** [November 27, 2023, 12:28pm UTC](https://discuss.elastic.co/t/ml-anomaly-job-with-exclude-frequent-option/348047 "2023-11-27T12:28:13Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![marmai16](https://avatars.discourse-cdn.com/v4/letter/m/13edae/32.png) [@marmai16](https://discuss.elastic.co/u/marmai16)\
**Post date:** [November 27, 2023, 12:28pm UTC](https://discuss.elastic.co/t/ml-anomaly-job-with-exclude-frequent-option/348047/1 "2023-11-27T12:28:13Z")

</div>

Hello everyone,

i was reading through the docs and became curious

> **[Create anomaly detection jobs API | Elasticsearch Guide \[8.11\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/ml-put-job.html)**

Say i create two detectors.  
One detector is `high_sum(a) over b`  
The other detector is `high_sum(a) by c`.

Now, if i define `exclude_frequent = over` for the first detector, then it will exclude frequently occurring values in `b` from triggering anomalies completely?

Consequently, it would make no sense to define `exclude_frequent = over` with the second detector `high_sum(a) by c` since `over_field` is empty?  
Thus in case of `high_sum(a) by c` we would need to set `exclude_frequent = by` to make it exclude frequent values in `c` from triggering anomalies?

Am i understanding this correctly?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [November 30, 2023, 2:53pm UTC](https://discuss.elastic.co/t/ml-anomaly-job-with-exclude-frequent-option/348047/2 "2023-11-30T14:53:51Z")

</div>

First and foremost, you should probably understand the major differences in using the `over` field versus using a split field like `by` or `partition`. Using the `over` field results in a population analysis (comparing entities against the population) which is much different than the normal temporal analysis (comparing an entity against its own history).

See: [Temporal vs. Population Analysis in Elastic Machine Learning | Elastic Blog](https://www.elastic.co/blog/temporal-vs-population-analysis-in-elastic-machine-learning)

Secondly, the use of `exclude_frequent` predated the creation of [Filters in Custom Rules](https://www.elastic.co/guide/en/machine-learning/current/ml-ad-run-jobs.html#ml-ad-rules), which give you more flexibility and control on what gets excluded.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 28, 2023, 2:54pm UTC](https://discuss.elastic.co/t/ml-anomaly-job-with-exclude-frequent-option/348047/3 "2023-12-28T14:54:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
