# Model used in anomaly detection

**URL:** <https://discuss.elastic.co/t/model-used-in-anomaly-detection/300264>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-monitoring\
**Created:** [March 22, 2022, 3:03am UTC](https://discuss.elastic.co/t/model-used-in-anomaly-detection/300264 "2022-03-22T03:03:57Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![zanattabruno](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zanattabruno/32/103328_2.png) [@zanattabruno](https://discuss.elastic.co/u/zanattabruno)\
**Post date:** [March 22, 2022, 3:03am UTC](https://discuss.elastic.co/t/model-used-in-anomaly-detection/300264/1 "2022-03-22T03:03:57Z")

</div>

Guys, can you let me know when I use the following options to enable anomaly detection in Elasticsearch.  
Analytics\>Machine Learning\>Anomaly Detection\>Create job\> Categorization  
Which algorithm/model would the implementation be based on to categorize the messages (it is a log base), and which algorithm/model would the implementation be based on for anomaly detection when I select the categorization approach.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/a/f/afdc3878f009ebc386a6b9b6ac01e64c0dc25757.png)  
I needed information on that level. Ex: "Decision Tree for categorization and Feedforward Neural Networks for anomaly detection."  
I tried to look at Github but couldn't check it accurately.

Thanks in advance

---

<div class="post-metadata">

**Author:** ![sagarpatel](https://avatars.discourse-cdn.com/v4/letter/s/ebca7d/32.png) [@sagarpatel](https://discuss.elastic.co/u/sagarpatel)\
**Post date:** [March 22, 2022, 5:22am UTC](https://discuss.elastic.co/t/model-used-in-anomaly-detection/300264/2 "2022-03-22T05:22:25Z")

</div>

Please check [this](https://discuss.elastic.co/t/what-is-the-machine-learning-algorithm-in-elastic/123527) question. it will clarify most of all your question.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [March 22, 2022, 2:00pm UTC](https://discuss.elastic.co/t/model-used-in-anomaly-detection/300264/3 "2022-03-22T14:00:46Z")

</div>

The methodology for Categorization is described here: [Elastic Machine Learning Tips and Tricks - Categorization - YouTube](https://www.youtube.com/watch?v=wPd_JWWfre4&ab_channel=Elastic)

Basically, it uses an approach of:

1. removing "mutable" tokens/words from the text (IP addresses, hostnames) by focusing on dictionary words (but this [can be customized](https://www.elastic.co/blog/categorizing-non-english-log-messages-in-machine-learning-for-elasticsearch))
2. Use an algorithm similar to [Levenhstein distance](https://en.wikipedia.org/wiki/Levenshtein_distance) to determine if strings of text are similar/close to strings of text seen before.
3. If the current string is similar, put it into the same category, otherwise create a new category.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 19, 2022, 2:01pm UTC](https://discuss.elastic.co/t/model-used-in-anomaly-detection/300264/4 "2022-04-19T14:01:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
