# Create a machine learning job with aggregation

**URL:** <https://discuss.elastic.co/t/create-a-machine-learning-job-with-aggregation/253579>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [October 28, 2020, 2:40pm UTC](https://discuss.elastic.co/t/create-a-machine-learning-job-with-aggregation/253579 "2020-10-28T14:40:07Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![TheHunter1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thehunter1/32/80190_2.png) [@TheHunter1](https://discuss.elastic.co/u/TheHunter1)\
**Post date:** [October 28, 2020, 2:40pm UTC](https://discuss.elastic.co/t/create-a-machine-learning-job-with-aggregation/253579/1 "2020-10-28T14:40:07Z")

</div>

Hello everybody,

So I just began with machine learning jobs and I wanna create a job to detect port scans.  
I wanna aggregate data by `source.ip` and then by `destination.ip` and finally count the number of `destination.port`  
Could you tell me how can I make an aggregation in machine learning jobs !

Thanks.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [October 28, 2020, 3:30pm UTC](https://discuss.elastic.co/t/create-a-machine-learning-job-with-aggregation/253579/2 "2020-10-28T15:30:09Z")

</div>

Well, to answer your question, information about how to use an elasticsearch query aggregation as part of your ML job can be found here: [https://www.elastic.co/guide/en/machine-learning/7.9/ml-configuring-aggregation.html](https://www.elastic.co/guide/en/machine-learning/7.9/ml-configuring-aggregation.html)

However, you likely have a very high cardinality of IP addresses. May I suggest that you instead use [Population Analysis](https://www.elastic.co/blog/temporal-vs-population-analysis-in-elastic-machine-learning) and configure something like the following:

detector: distinct\_count(destination.port) over destination.ip  
influencers: destination.ip, source.ip

The population analysis will effectively ease the burden on the high-cardinality destination IP field and the source IP as an influencer will only get analyzed if there's an anomaly on the distinct count, as defined by the detector.

---

<div class="post-metadata">

**Author:** ![TheHunter1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thehunter1/32/80190_2.png) [@TheHunter1](https://discuss.elastic.co/u/TheHunter1)\
**Post date:** [October 28, 2020, 9:30pm UTC](https://discuss.elastic.co/t/create-a-machine-learning-job-with-aggregation/253579/3 "2020-10-28T21:30:44Z")

</div>

Thanks a lot for your reply and for your advices

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 25, 2020, 9:30pm UTC](https://discuss.elastic.co/t/create-a-machine-learning-job-with-aggregation/253579/4 "2020-11-25T21:30:46Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
