# Anomalie Detection : Need Help Please

**URL:** <https://discuss.elastic.co/t/anomalie-detection-need-help-please/276359>\
**Category:** Kibana\
**Tags:** elastic-stack-machine-learning\
**Created:** [June 18, 2021, 9:17am UTC](https://discuss.elastic.co/t/anomalie-detection-need-help-please/276359 "2021-06-18T09:17:44Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ydr175020221](https://avatars.discourse-cdn.com/v4/letter/y/90ced4/32.png) [@ydr175020221](https://discuss.elastic.co/u/ydr175020221)\
**Post date:** [June 18, 2021, 9:17am UTC](https://discuss.elastic.co/t/anomalie-detection-need-help-please/276359/1 "2021-06-18T09:17:44Z")

</div>

Hello,

I'm desperately trying to create a ML job to answer the following question:

How to identify a gap for a user based on elements like: Country, City, ...  
Example if an user "Max" who often connects from FRANCE, city of PARIS and uses a WINDOWS, at a given moment connects from another country, or another city, or with another OS is notified by my JOB.  
I have an index which contains all the information about user like userid, country, city, OSFamily ....  
Once the JOB started, no result displayed, I think my choices are not good.  
Here the JOB in question:

{  
"job\_id": "3",  
"job\_type": "anomaly\_detector",  
"job\_version": "7.13.1",  
"create\_time": 1624003582815,  
"custom\_settings": {  
"custom\_urls":   
},  
"description": "",  
"analysis\_config": {  
"bucket\_span": "15m",  
"detectors": [  
{  
"detector\_description": "distinct\_count(country) partitionfield=userId",  
"function": "distinct\_count",  
"field\_name": "country",  
"partition\_field\_name": "userId",  
"detector\_index": 0  
},  
{  
"detector\_description": "distinct\_count(city) partitionfield=userId",  
"function": "distinct\_count",  
"field\_name": "city",  
"partition\_field\_name": "userId",  
"detector\_index": 1  
},  
{  
"detector\_description": "distinct\_count(osFamily) partitionfield=userId",  
"function": "distinct\_count",  
"field\_name": "osFamily",  
"partition\_field\_name": "userId",  
"detector\_index": 2  
}  
],  
"influencers": [  
"country",  
"city",  
"osFamily"  
]  
},  
"analysis\_limits": {  
"model\_memory\_limit": "166mb",  
"categorization\_examples\_limit": 4  
},  
"data\_description": {  
"time\_field": "startTime",  
"time\_format": "epoch\_ms"  
},  
"model\_plot\_config": {  
"enabled": false,  
"annotations\_enabled": false  
},  
"model\_snapshot\_retention\_days": 10,  
"daily\_model\_snapshot\_retention\_after\_days": 1,  
"results\_index\_name": "shared",  
"allow\_lazy\_open": false,  
"data\_counts": {  
"job\_id": "3",  
"processed\_record\_count": 0,  
"processed\_field\_count": 0,  
"input\_bytes": 0,  
"input\_field\_count": 0,  
"invalid\_date\_count": 0,  
"missing\_field\_count": 0,  
"out\_of\_order\_timestamp\_count": 0,  
"empty\_bucket\_count": 0,  
"sparse\_bucket\_count": 0,  
"bucket\_count": 0,  
"input\_record\_count": 0  
},  
"model\_size\_stats": {  
"job\_id": "3",  
"result\_type": "model\_size\_stats",  
"model\_bytes": 0,  
"total\_by\_field\_count": 0,  
"total\_over\_field\_count": 0,  
"total\_partition\_field\_count": 0,  
"bucket\_allocation\_failures\_count": 0,  
"memory\_status": "ok",  
"categorized\_doc\_count": 0,  
"total\_category\_count": 0,  
"frequent\_category\_count": 0,  
"rare\_category\_count": 0,  
"dead\_category\_count": 0,  
"failed\_category\_count": 0,  
"categorization\_status": "ok",  
"log\_time": 1624003604016  
},  
"forecasts\_stats": {  
"total": 0,  
"forecasted\_jobs": 0  
},  
"state": "opened",  
"node": {  
"id": "ZHYsMqCORDeOcS2b3XLQog",  
"name": "node-1",  
"ephemeral\_id": "rG5ykKlSTVWlZdHbmHZHug",  
"transport\_address": "10.10.0.217:9300",  
"attributes": {  
"ml.machine\_memory": "8589328384",  
"xpack.installed": "true",  
"transform.node": "true",  
"ml.max\_open\_jobs": "512",  
"ml.max\_jvm\_size": "4294967296"  
}  
},  
"assignment\_explanation": "",  
"open\_time": "2497s",  
"timing\_stats": {  
"job\_id": "3",  
"bucket\_count": 0,  
"total\_bucket\_processing\_time\_ms": 0,  
"exponential\_average\_bucket\_processing\_time\_per\_hour\_ms": 0  
},  
"datafeed\_config": {  
"datafeed\_id": "datafeed-3",  
"job\_id": "3",  
"query\_delay": "102956ms",  
"chunking\_config": {  
"mode": "auto"  
},  
"indices\_options": {  
"expand\_wildcards": [  
"open"  
],  
"ignore\_unavailable": false,  
"allow\_no\_indices": true,  
"ignore\_throttled": true  
},  
"query": {  
"bool": {  
"must": {  
"exists": {  
"field": "userId"  
}  
}  
}  
},  
"indices": [  
"log\_users"  
],  
"scroll\_size": 1000,  
"delayed\_data\_check\_config": {  
"enabled": true  
},  
"state": "started",  
"node": {  
"id": "ZHYsMqCORDeOcS2b3XLQog",  
"name": "node-1",  
"ephemeral\_id": "rG5ykKlSTVWlZdHbmHZHug",  
"transport\_address": "10.10.0.217:9300",  
"attributes": {  
"ml.machine\_memory": "8589328384",  
"ml.max\_open\_jobs": "512",  
"ml.max\_jvm\_size": "4294967296"  
}  
},  
"assignment\_explanation": "",  
"timing\_stats": {  
"job\_id": "3",  
"search\_count": 6,  
"bucket\_count": 0,  
"total\_search\_time\_ms": 6220,  
"exponential\_average\_search\_time\_per\_hour\_ms": 6220  
}  
}  
}

Thanks. 🙂

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [June 22, 2021, 11:29am UTC](https://discuss.elastic.co/t/anomalie-detection-need-help-please/276359/2 "2021-06-22T11:29:23Z")

</div>

It seems to me you need to use [rare functions.](https://www.elastic.co/guide/en/machine-learning/current/ml-rare-functions.html)

Rare functions in the detectors allow you to bubble up rare behavior given past behavior.

Here is a [nice blog digging into it further](https://www.elastic.co/blog/detecting-rare-unusual-processes-with-elastic-machine-learning)

I think, if each user is different and should be treated separately, you will have three detectors

```json
{
"detector_description": "rare by country partitionfield=userId",
"function": "rare",
"by_field_name": "country",
"partition_field_name": "userId"
},
{
"detector_description": "rare by city partitionfield=userId",
"function": "rare",
"by_field_name": "city",
"partition_field_name": "userId"
},
{
"detector_description": "rare by osFamily partitionfield=userId",
"function": "rare",
"by_field_name": "osFamily",
"partition_field_name": "userId"
}

```

---

<div class="post-metadata">

**Author:** ![ydr175020221](https://avatars.discourse-cdn.com/v4/letter/y/90ced4/32.png) [@ydr175020221](https://discuss.elastic.co/u/ydr175020221)\
**Post date:** [June 22, 2021, 12:59pm UTC](https://discuss.elastic.co/t/anomalie-detection-need-help-please/276359/3 "2021-06-22T12:59:08Z")

</div>

Hi BenTrent,

Thank you so much for your answer, i'll try that 😃

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 20, 2021, 12:59pm UTC](https://discuss.elastic.co/t/anomalie-detection-need-help-please/276359/4 "2021-07-20T12:59:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
