# Help: Create multi metric machine learning job

**URL:** <https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [December 4, 2020, 3:35pm UTC](https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670 "2020-12-04T15:35:41Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Abdelhalim](https://avatars.discourse-cdn.com/v4/letter/a/838e76/32.png) [@Abdelhalim](https://discuss.elastic.co/u/Abdelhalim)\
**Post date:** [December 4, 2020, 3:35pm UTC](https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670/1 "2020-12-04T15:35:41Z")

</div>

Hello everybody,

I need help to create a multi metric job.  
The idea is that I want to detect users that connect from different countries, for example if a user used to connect from Germany, if one day he connects from France, I will receive an alert.  
I tried to create the job like that:

**metric:** `Distinct count(source.geo.country_code2.keyword)` and `Distinct count(source.nat.geo.country_code2.keyword)`

**split field** : `user.name`

The result I got are like below, which is not the result that I want as it's not responding to my need

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/1/b1de67eabe690397b3bacdf28f7b785d15fe59cf.png)

Could you please tell me how can I create this job.

Best regards

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [December 4, 2020, 4:12pm UTC](https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670/2 "2020-12-04T16:12:50Z")

</div>

The likely thing you really want to do here is to leverage the `rare` detector function to find a country that is rare for a user

`rare by source.geo.country_code2.keyword partition=user.name`

See a similar example here: [Dec 4th, 2018: [EN][ML] Rarity Analysis with Machine Learning](https://discuss.elastic.co/t/dec-4th-2018-en-ml-rarity-analysis-with-machine-learning/158979)

Just want to be cognizant of the cardinality of the `user.name` field. If it is really high you'll require a lot of memory utilization for the job.

Also, from your screenshot it seems like your data has empty string for some user names. You might want to filter those out in the datafeed query (??)

---

<div class="post-metadata">

**Author:** ![Abdelhalim](https://avatars.discourse-cdn.com/v4/letter/a/838e76/32.png) [@Abdelhalim](https://discuss.elastic.co/u/Abdelhalim)\
**Post date:** [December 7, 2020, 9:19am UTC](https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670/3 "2020-12-07T09:19:39Z")

</div>

Thanks for your help @richcollier,

Could you tell me please if this query to filter empty `user.name` field is correct and if I should add an influencer to my job as I am getting a warning that my job has no influencer !

```auto
{
  "bool": {
    "must": [
      {
        "exists": {
          "field": "user.name"
        }
      }
    ]
  }
}

```

**My job Configuration:**

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/8/e/8e5554d579cb0286a55972f6c190b2c607a325b8.png)

I am getting an empty result, so don't know if there is a mistake in my configuration, or it's just no anomaly was found.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/1/3/13554fd9f1ab3a54e2376ec4d6178a94be9a0438.png)

Best regards

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [December 7, 2020, 2:13pm UTC](https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670/4 "2020-12-07T14:13:20Z")

</div>

I would make both `user.name` and `source.geo.country_code2.keyword` as influencers in your configuration.

As for your results, it's possible that you don't have an example of an anomaly in your data yet. Often, when testing, it is good to have the job learn on a good amount (weeks if possible) data, then contrive a situation (manually force the indexing of a sample document of a user connecting from a strange location).

---

<div class="post-metadata">

**Author:** ![Abdelhalim](https://avatars.discourse-cdn.com/v4/letter/a/838e76/32.png) [@Abdelhalim](https://discuss.elastic.co/u/Abdelhalim)\
**Post date:** [December 7, 2020, 2:29pm UTC](https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670/5 "2020-12-07T14:29:46Z")

</div>

Thank you very much for this valuable information  
My index has almost 2 months of data, I will try to generate later manually a connection from a rare country and see if the machine learning detects it

Thanks again 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 4, 2021, 2:29pm UTC](https://discuss.elastic.co/t/help-create-multi-metric-machine-learning-job/257670/6 "2021-01-04T14:29:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
