# Security Analytics Recipes

**URL:** <https://discuss.elastic.co/t/security-analytics-recipes/93591>\
**Category:** Elasticsearch\
**Created:** [July 18, 2017, 1:14pm UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591 "2017-07-18T13:14:34Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![senthil\_blue](https://avatars.discourse-cdn.com/v4/letter/s/35a633/32.png) [@senthil\_blue](https://discuss.elastic.co/u/senthil_blue)\
**Post date:** [July 18, 2017, 1:14pm UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591/1 "2017-07-18T13:14:34Z")

</div>

Hi,  
I am trying to explore the example [https://github.com/elastic/examples/blob/master/Machine%20Learning/Security%20Analytics%20Recipes/dns\_data\_exfiltration/EXAMPLE.md](https://github.com/elastic/examples/blob/master/Machine%20Learning/Security%20Analytics%20Recipes/dns_data_exfiltration/EXAMPLE.md)

for Anomaly detection. However, I can't seem to get the results in the explorer. The datafeed job has been running for about a day now and it has processed about 32k records. I have followed all the instructions in that github page but I am not sure how long this job needs to run to get a result in the Anomaly explorer for exploration purpose. Can someone please help?

Thanks,  
Senthil.

---

<div class="post-metadata">

**Author:** ![sophie\_chang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sophie_chang/32/18008_2.png) [@sophie\_chang](https://discuss.elastic.co/u/sophie_chang)\
**Post date:** [July 19, 2017, 12:26pm UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591/2 "2017-07-19T12:26:31Z")

</div>

Hi

Missing anomaly results could be due to a few factors. The following checks cover the most common causes...

1. Check for sufficient data

For typical data, the job needs to run for at least 2 hours or 20 buckets (whichever is longer) for the model to initialized and results to be written - I would expect sufficient run time in this instance.

1. Check results exist

The Anomaly Explorer shows anomalies. As it is possible that the data does not contain anomalies for this time period, then check if any results are being written to elasticsearch. In Dev Tools run:

`GET _xpack/ml/anomaly_detectors/<job_id>/results/buckets`

One document should be returned per bucket time stamp, even if there are no anomalies found. Check the time stamps for results found match the Kibana time picker.

1. Check input data

The input data fields may be missing. The processing may be processing records, but there may be many missing values which mean no data is being modelled and consequently there are no results.

To check what the data feed is sending for analysis, run:

`GET _xpack/ml/datafeeds/datafeed-<job_id>/_preview`

Check that values are returned for all of the fields that are used in the job configuration.

If all of the above checks seem OK, then please send us the job info. In the Job Management page, expand the job entry and select the `JSON` tab. Please include this content for us to see.

Thanks

---

<div class="post-metadata">

**Author:** ![senthil\_blue](https://avatars.discourse-cdn.com/v4/letter/s/35a633/32.png) [@senthil\_blue](https://discuss.elastic.co/u/senthil_blue)\
**Post date:** [July 26, 2017, 4:47pm UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591/3 "2017-07-26T16:47:25Z")

</div>

Hi,  
Thanks your help. Here is the JSON.

{  
"job\_id": "suspicious\_login\_activity",  
"job\_type": "anomaly\_detector",  
"description": "suspicious login activity",  
"create\_time": 1501018984377,  
"finished\_time": 1501019008271,  
"analysis\_config": {  
"bucket\_span": "5m",  
"detectors": [  
{  
"detector\_description": "high\_count",  
"function": "high\_count",  
"partition\_field\_name": "system.auth.hostname",  
"detector\_rules": []  
}  
],  
"influencers": [  
"system.auth.hostname",  
"system.auth.user",  
"system.auth.ssh.ip"  
]  
},  
"data\_description": {  
"time\_field": "@timestamp",  
"time\_format": "epoch\_ms"  
},  
"model\_plot\_config": {  
"enabled": true  
},  
"model\_snapshot\_retention\_days": 1,  
"results\_index\_name": "shared",  
"data\_counts": {  
"job\_id": "suspicious\_login\_activity",  
"processed\_record\_count": 0,  
"processed\_field\_count": 0,  
"input\_bytes": 0,  
"input\_field\_count": 0,  
"invalid\_date\_count": 0,  
"missing\_field\_count": 0,  
"out\_of\_order\_timestamp\_count": 0,  
"empty\_bucket\_count": 0,  
"sparse\_bucket\_count": 0,  
"bucket\_count": 0,  
"input\_record\_count": 0  
},  
"model\_size\_stats": {  
"job\_id": "suspicious\_login\_activity",  
"result\_type": "model\_size\_stats",  
"model\_bytes": 0,  
"total\_by\_field\_count": 0,  
"total\_over\_field\_count": 0,  
"total\_partition\_field\_count": 0,  
"bucket\_allocation\_failures\_count": 0,  
"memory\_status": "ok",  
"log\_time": 1501019007000,  
"timestamp": -300000  
},  
"datafeed\_config": {  
"datafeed\_id": "datafeed-suspicious\_login\_activity",  
"job\_id": "suspicious\_login\_activity",  
"query\_delay": "60s",  
"frequency": "150s",  
"indexes": [  
"filebeat"  
],  
"types": [  
"doc"  
],  
"query": {  
"query\_string": {  
"query": "system.auth.ssh.event:Failed OR system.auth.ssh.event:Invalid",  
"fields": [],  
"use\_dis\_max": true,  
"tie\_breaker": 0,  
"default\_operator": "or",  
"auto\_generate\_phrase\_queries": false,  
"max\_determinized\_states": 10000,  
"enable\_position\_increments": true,  
"fuzziness": "AUTO",  
"fuzzy\_prefix\_length": 0,  
"fuzzy\_max\_expansions": 50,  
"phrase\_slop": 0,  
"analyze\_wildcard": true,  
"escape": false,  
"split\_on\_whitespace": true,  
"boost": 1  
}  
},  
"scroll\_size": 1000,  
"chunking\_config": {  
"mode": "auto"  
},  
"state": "stopped"  
},  
"state": "opened",  
"node": {  
"id": "jEnDRnUqSt22\_CsGI9BzSA",  
"name": "Oculus",  
"ephemeral\_id": "QKO13r64Sv-dN0etmOfNag",  
"transport\_address": "x.x.x.x:9300",  
"attributes": {  
"ml.enabled": "true"  
}  
},  
"open\_time": "0s"  
}  
Basically we are trying to implement this recipe here.

> <https://github.com/elastic/examples/blob/master/Machine%20Learning/Security%20Analytics%20Recipes/suspicious_login_activity/EXAMPLE.md>

The data has been ingested to ES but we are not able to get the 'datafeed' process the index.

Thanks again.

---

<div class="post-metadata">

**Author:** ![sophie\_chang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sophie_chang/32/18008_2.png) [@sophie\_chang](https://discuss.elastic.co/u/sophie_chang)\
**Post date:** [July 27, 2017, 5:45pm UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591/4 "2017-07-27T17:45:27Z")

</div>

Hi

From the details you have provided, the datafeed looks like it is currently stopped and zero records have been processed. I presume this is a config from a recently created job.

Can you please confirm what the output from the following looked like:

```auto
GET _xpack/ml/datafeeds/datafeed-suspicious_login_activity/_preview

```

I would expect this to contain data for the following fields:

```auto
...
  {
    "@timestamp": 1454803200000,
    "system.auth.hostname": "hostname1",
    "system.auth.user": "thomas",
    "system.auth.ssh.ip": "10.2.3.14"
  },
...

```

Was this job running in real-time, or did you select a start date that was historical e.g. something like 2 weeks ago.

If you are running in real-time, then the datafeed is running every 150s and selecting data from greater 60s ago. If your data takes longer to ingest than 60s, then you'll need to adjust some of the datafeed settings. You can do this in the UI. In Job Management, select the Edit icon for this job, and click on the Datafeed tab. Then edit:

```auto
"query_delay": "60s",
"frequency": "150s",

```

(You'll need to stop the datafeed first).

If you were running on historical data, (i.e. you selected a datafeed start time that was fairly far in the past), then the datafeed preview should hold the answer. Are the correct fields being returned for analysis?

Regards

---

<div class="post-metadata">

**Author:** ![senthil\_blue](https://avatars.discourse-cdn.com/v4/letter/s/35a633/32.png) [@senthil\_blue](https://discuss.elastic.co/u/senthil_blue)\
**Post date:** [July 27, 2017, 8:30pm UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591/5 "2017-07-27T20:30:34Z")

</div>

Hi,  
Thanks so much for your response.

How exactly should the indexes in ES be created . I don't see these fields in ES.

{  
"@timestamp": 1454803200000,  
"system.auth.hostname": "hostname1",  
"system.auth.user": "thomas",  
"system.auth.ssh.ip": "10.2.3.14"  
},

The provided sample dataset auth.log doesn't quite match these fields after being ingested through filebeat. I am running on historical data based on the dataset from the security recipe.

> <https://github.com/elastic/examples/blob/master/Machine%20Learning/Security%20Analytics%20Recipes/suspicious_login_activity/data/auth.log>

---

<div class="post-metadata">

**Author:** ![sophie\_chang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sophie_chang/32/18008_2.png) [@sophie\_chang](https://discuss.elastic.co/u/sophie_chang)\
**Post date:** [July 28, 2017, 9:50am UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591/6 "2017-07-28T09:50:40Z")

</div>

Could you please check on the names of the `filebeat` indices in your elasticsearch instance?

```auto
GET _cat/indices/fileb*

```

It is possible that the `filebeat` indices have the date pattern in the name e.g. `filebeat-2017.07.27`. If this is the case, then I see a problem with the datafeed configuration.

The config below:

```auto
"indexes": [
"filebeat"
],

```

... should be:

```auto
"indexes": [
"filebeat-*"
],

```

To correct this, using the UI, clone the ML job.

- In the Job Details tab  
-- Give it a new distinct name e.g. `suspicious_login_activity_2`  
-- Click on `Use dedicated index` _(this is not strictly necessary, but may avoid other issues for the purposes of troubleshooting)_
- Go to the Datafeed tab  
-- Change the Index from `filebeat` to `filebeat-*`. Make sure the `Time-field name` is selected as well as `All types`. e.g.  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/1/3/1324b81898f8d995db6204658c9b6a23ed91a766.png)  
(Note: You may have different values for time-field name and a different list of types - just pick from the list that is pre-populated).
- Go to the Datafeed Preview tab - here you should be able to see a sample of the data to be analyzed.
- Click Save
- Click Start datafeed  
-- Select to run from the beginning of the data to Now (or continue in real-time)

Hope this gets you a little closer. I'll ask the Examples team to double check the recipe is working end to end. It really should be updated to 5.5 by now.

Regards  
Sophie

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 25, 2017, 9:50am UTC](https://discuss.elastic.co/t/security-analytics-recipes/93591/7 "2017-08-25T09:50:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
