# Anomaly detection in Machine learning kibana

**URL:** <https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018>\
**Category:** Kibana\
**Tags:** elastic-stack-machine-learning\
**Created:** [November 24, 2021, 7:34am UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018 "2021-11-24T07:34:49Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![MAHALAKSHMI\_S](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mahalakshmi_s/32/97482_2.png) [@MAHALAKSHMI\_S](https://discuss.elastic.co/u/MAHALAKSHMI_S)\
**Post date:** [November 24, 2021, 7:34am UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/1 "2021-11-24T07:34:49Z")

</div>

I created a "Categorization" based anomaly detection job and explored the job results in Kibana.  
I have used "Payload"(i.e String) field for categorization.  
Here I'm not sure what does "typical" value in anomaly explorer results signifies?

P.S I knew for any numerical feature, "typical" value signifies the median of those values. But not sure in case of string

---

<div class="post-metadata">

**Author:** ![droberts195](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/droberts195/32/17692_2.png) [@droberts195](https://discuss.elastic.co/u/droberts195)\
**Post date:** [November 24, 2021, 10:36am UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/2 "2021-11-24T10:36:26Z")

</div>

With ML categorization jobs you still do an anomaly detection as well as a categorization. Usually this would use a function of the category ID. It's almost always `rare by mlcategory` or `count by mlcategory`, and since you don't know which you've got it must be one of these that's been added by the categorization wizard. You can find out by looking at the job configuration in the ML jobs list.

If it's `count by mlcategory` then your `typical` and `actual` will be how many categories of Payload typically and actually occur per time bucket. If it's `rare by mlcategory` then `typical` will be the probability of seeing that category in a typical bucket.

You can see the category definitions without the anomaly information using the [Get Categories API](https://www.elastic.co/guide/en/elasticsearch/reference/current/ml-get-category.html).

---

<div class="post-metadata">

**Author:** ![MAHALAKSHMI\_S](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mahalakshmi_s/32/97482_2.png) [@MAHALAKSHMI\_S](https://discuss.elastic.co/u/MAHALAKSHMI_S)\
**Post date:** [November 24, 2021, 11:01am UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/3 "2021-11-24T11:01:39Z")

</div>

Thanks for your reply @droberts195 ..

I used **count by mlcategory**. However the **actual** field gives out the count of that payload category in the datafeed, which I felt totally different from what you said in the last post.  
In addition to that **typical** field is in **float** type.

 ![doubt](https://us1.discourse-cdn.com/elastic/original/3X/7/4/749827af29d0b58821c6404d292bea8d2ee57fa3.png)

---

<div class="post-metadata">

**Author:** ![droberts195](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/droberts195/32/17692_2.png) [@droberts195](https://discuss.elastic.co/u/droberts195)\
**Post date:** [November 24, 2021, 11:20am UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/4 "2021-11-24T11:20:11Z")

</div>

What that information is saying is that for category 8 there were 1330 documents on 26th December 2021 that were classified as category 8. On average there are 221.6 documents per bucket in category 8. The reason `typical` is a `float` is because the expected value of a distribution of integers isn't always an integer. For example, the expected value from rolling a standard 6-sided dice is 3.5, but you'll never roll 3.5 on a single roll of the dice.

---

<div class="post-metadata">

**Author:** ![MAHALAKSHMI\_S](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mahalakshmi_s/32/97482_2.png) [@MAHALAKSHMI\_S](https://discuss.elastic.co/u/MAHALAKSHMI_S)\
**Post date:** [November 24, 2021, 1:43pm UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/5 "2021-11-24T13:43:43Z")

</div>

Thanks @droberts195 , for clarifying this.

In case of **population analysis** with bucket span of 1 hour, it would be very helpful if you can explain about the **typical** value here.  
In the screenshot attached, typical value remains the same for all anomalous messages. Why is it so?

 ![doubt_1](https://us1.discourse-cdn.com/elastic/original/3X/e/8/e8c349d5fec07f87292a688727a004ae92b79b5e.png)

---

<div class="post-metadata">

**Author:** ![droberts195](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/droberts195/32/17692_2.png) [@droberts195](https://discuss.elastic.co/u/droberts195)\
**Post date:** [November 24, 2021, 3:03pm UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/6 "2021-11-24T15:03:43Z")

</div>

For a population job anomalies are created for entities that are significantly different to the population within the current time bucket. So the average count for the population in that time bucket is 2.27, and those two entities have much higher counts.

---

<div class="post-metadata">

**Author:** ![MAHALAKSHMI\_S](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mahalakshmi_s/32/97482_2.png) [@MAHALAKSHMI\_S](https://discuss.elastic.co/u/MAHALAKSHMI_S)\
**Post date:** [November 25, 2021, 5:28am UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/7 "2021-11-25T05:28:57Z")

</div>

Is it possible to get all buckets and count of documents in each bucket for category 8?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 23, 2021, 5:29am UTC](https://discuss.elastic.co/t/anomaly-detection-in-machine-learning-kibana/290018/8 "2021-12-23T05:29:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
