# Kibana ML anomalies detection and regression

**URL:** <https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894>\
**Category:** Kibana\
**Tags:** elastic-stack-machine-learning\
**Created:** [July 20, 2020, 1:17pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894 "2020-07-20T13:17:58Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [July 20, 2020, 1:17pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/1 "2020-07-20T13:17:58Z")

</div>

Hi,

I have trend data and i want to identify customers whose turnover climb or fall .  
Witch Kibana ML model can i do to identify them?  
Anomalies detection or regression ?  
I also want to know how to interpret the results of anomalies detection based on poplulation metric and how to interpret the result of a regression?  
For regression per example, i have done an analysis witch predict the turnover. But i dont know what to do with the result?  
How can i interpret that( i have training r2 =0,7 and testing r2=0,4).  
when can i say my model is good?  
How the regression analysis in kibana works? I have used a training percent of 90.  
This normaly means that 90% of data is used for training and 10% for testing.  
is the 10% a part of my data or is this unseen data that kibana will generate?

Thanks for answers.

Cordialy

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 20, 2020, 4:32pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/2 "2020-07-20T16:32:39Z")

</div>

If you have a time-series based trend-line of the rates of customer turnovers, you can certainly use anomaly detection to assess if the current rate is higher/lower than typical.

If you are trying to assess/predict whether or not a specific customer is likely to turnover or not (based upon the values of other fields that could be indicative), then a classification analytics job would be the right approach. See a good example of that here: [https://www.elastic.co/webinars/introduction-to-supervised-machine-learning-in-elastic](https://www.elastic.co/webinars/introduction-to-supervised-machine-learning-in-elastic)

---

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [July 31, 2020, 12:08pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/3 "2020-07-31T12:08:43Z")

</div>

Thanks Richcollier.. So can you explain me please how the typical value is calculated?? I have some values that i don't understand.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 31, 2020, 12:51pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/4 "2020-07-31T12:51:26Z")

</div>

The `actual` value is from the raw data itself (the data ML is analyzing). The `typical` value is the highest probable value from the internal statistical model that ML has constructed for that data set.

Perhaps a [basic introductory video](https://www.youtube.com/watch?v=n6xW6YWYgs0) on ML's anomaly detection would be helpful.

---

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [July 31, 2020, 12:55pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/5 "2020-07-31T12:55:51Z")

</div>

Okay thank you.. But what are exactly the statistical models that are used to calculate typical values?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 31, 2020, 1:10pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/6 "2020-07-31T13:10:03Z")

</div>

We designed them. You can get a sense of the foundation by looking at this whitepaper:

[http://www.ijmlc.org/papers/398-LC018.pdf](http://www.ijmlc.org/papers/398-LC018.pdf)

Or probably easier is to watch this:

[https://www.elastic.co/elasticon/conf/2017/sf/machine-learning-and-statistical-methods-for-time-series-analysis](https://www.elastic.co/elasticon/conf/2017/sf/machine-learning-and-statistical-methods-for-time-series-analysis)

---

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [August 3, 2020, 8:36am UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/7 "2020-08-03T08:36:59Z")

</div>

Hi Richcollier, I have another question..my typical value are like this.see screenshot.  
I think this is not normal..what do you think??

 ![Capture_typical](https://us1.discourse-cdn.com/elastic/original/3X/c/e/ced7925d498e7e673d0aec7b00068e5233af0a3f.png)

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [August 3, 2020, 1:01pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/8 "2020-08-03T13:01:41Z")

</div>

I'm not sure I understand why you think those values look incorrect. Can you provide more context and/or the reason why you think that?

---

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [August 4, 2020, 8:45am UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/9 "2020-08-04T08:45:59Z")

</div>

I think that because typical values are supposed to be the "highest probable values (p-value)" so it should be a value between 0 and 1. is it correct??  
but in my screenshot I have values ​​greater than 1.  
on the other hand the p value is not smaller than 0.05 does that mean that the result is not significant?

Cordialy,

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [August 4, 2020, 11:51am UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/10 "2020-08-04T11:51:02Z")

</div>

The typical value is the highest probable value of the measurement, not the highest probability. As an analogy, the highest probable value of a two dice being rolled is "7"

---

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [August 4, 2020, 12:06pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/11 "2020-08-04T12:06:04Z")

</div>

Okay I understand better.. But i still dont understand my result.  
I have got this result after the analysis

 ![Capture_typical_actual](https://us1.discourse-cdn.com/elastic/original/3X/5/5/55f6f7a5841ee2fffa3848a0491453a911d51b34.png)

---

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [August 4, 2020, 12:07pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/12 "2020-08-04T12:07:39Z")

</div>

As you see there is a big gap beetween the actual and the typycal values..

---

<div class="post-metadata">

**Author:** ![ahadi](https://avatars.discourse-cdn.com/v4/letter/a/ee7513/32.png) [@ahadi](https://discuss.elastic.co/u/ahadi)\
**Post date:** [August 4, 2020, 12:15pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/13 "2020-08-04T12:15:59Z")

</div>

ignore my last posts @richcollier i understood.  
Thank you very much

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 1, 2020, 12:16pm UTC](https://discuss.elastic.co/t/kibana-ml-anomalies-detection-and-regression/241894/14 "2020-09-01T12:16:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
