# Anomaly Detection Assistance

**URL:** <https://discuss.elastic.co/t/anomaly-detection-assistance/326004>\
**Category:** APM\
**Tags:** elastic-stack-machine-learning, open-telemetry\
**Created:** [February 20, 2023, 7:28pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004 "2023-02-20T19:28:38Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![tr78](https://avatars.discourse-cdn.com/v4/letter/t/6a8cbe/32.png) [@tr78](https://discuss.elastic.co/u/tr78)\
**Post date:** [February 20, 2023, 7:28pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/1 "2023-02-20T19:28:38Z")

</div>

**Kibana version** : 8.6.1 (ECK)

**Elasticsearch version** : 8.6.1 (ECK)

**APM Server version** : Fleet 8.6.1 (ECK)

**APM Agent language and version** :

**Browser version** : Chrome 110.0.5481.104

**Original install method (e.g. download page, yum, deb, from source, etc.) and version**: ECK

**Fresh install or upgraded from other version?** : Fresh

Looking for some general guidance regarding machine learning jobs. I have an ECK cluster with APM accepting OTEL data from a collector.

My goal is, using machine learning, to detect anomalies/outliers (I'm not sure which is the correct route here) in a specific field within an APM index. For example, I'd like to be able to check via API whether the build of a product took longer than it typically does. So a build runs, and upon completion it checks elastic to determine whether that particular build ID was an anomaly.

This all sounds simple (and probably is), but I'm not sure what path I should be taking to do this. The data is getting to Elastic, but I'm not sure how to analyze it. Should I use an anomaly detection job? Should I use a data frame analysis job? Should I not be using either? I have been able to successfully set up ongoing anomaly detection, but from what I can tell, the anomaly detection works off a mean,median, etc, of a specified bucket. So I'm not sure how to go about determining if just one specific data point is an anomaly.

Hoping this all makes sense. I'm new to all this so let me know if this doesn't make sense and I'll try and clarify.

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [February 21, 2023, 1:37pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/2 "2023-02-21T13:37:01Z")

</div>

Anomaly detection is the right approach for time-series data (metrics, logs, etc.). Yes, it uses bucketing, but the user has control over the size of the bucketing (see [bucket\_span](https://www.elastic.co/blog/explaining-the-bucket-span-in-machine-learning-for-elasticsearch)).

Anomaly detection will not tell you if a "single" measurement is anomalous in time, unless that measurement happens to be the only one in the current bucket\_span. Usage of the `max` function, for example, will approximate that since it self-selects the single largest measurement in the bucket\_span.

---

<div class="post-metadata">

**Author:** ![tr78](https://avatars.discourse-cdn.com/v4/letter/t/6a8cbe/32.png) [@tr78](https://discuss.elastic.co/u/tr78)\
**Post date:** [February 21, 2023, 6:35pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/3 "2023-02-21T18:35:52Z")

</div>

@richcollier Thanks - that makes sense about using `max` . To clarify though, even when using max there is no way to correlate that to a specific document, is that correct? So for example I run a build, I want to check if that build's final duration is an anomaly - is that not feasible?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [February 22, 2023, 2:16pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/4 "2023-02-22T14:16:20Z")

</div>

You can always link from the Anomaly Detection results back to Discover to manually inspect the documents in that time bucket:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/a/0/a0a336e546a4cd355c3949e27e92ef343d2e423c.png)

---

<div class="post-metadata">

**Author:** ![tr78](https://avatars.discourse-cdn.com/v4/letter/t/6a8cbe/32.png) [@tr78](https://discuss.elastic.co/u/tr78)\
**Post date:** [February 22, 2023, 2:38pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/5 "2023-02-22T14:38:42Z")

</div>

I do not have that option, it is greyed out. When I hover I get:

`Unable to link to Discover; no data view exists for index 'traces-apm*'`

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [February 22, 2023, 6:48pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/6 "2023-02-22T18:48:54Z")

</div>

Just manually create a [Data View](https://www.elastic.co/guide/en/kibana/master/data-views.html) for that index?

---

<div class="post-metadata">

**Author:** ![tr78](https://avatars.discourse-cdn.com/v4/letter/t/6a8cbe/32.png) [@tr78](https://discuss.elastic.co/u/tr78)\
**Post date:** [February 22, 2023, 8:50pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/7 "2023-02-22T20:50:37Z")

</div>

Got it - thank you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 15, 2023, 4:51pm UTC](https://discuss.elastic.co/t/anomaly-detection-assistance/326004/8 "2023-03-15T16:51:15Z")

</div>

This topic was automatically closed 20 days after the last reply. New replies are no longer allowed.
