# Created ML Job with bucket span 15m and 1d

**URL:** <https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [October 17, 2018, 12:41pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839 "2018-10-17T12:41:16Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![shiv94](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@shiv94](https://discuss.elastic.co/u/shiv94)\
**Post date:** [October 17, 2018, 12:41pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/1 "2018-10-17T12:41:16Z")

</div>

Hi All,

I have created ML job of multi metric with bucket span of 15m and used the function Max(A), Max(B). Splitting it by hostname and also created other Job with bucket span of 1d. I would like to know how exactly bucket span works? does it consider maximum value of A in 15m or 1d per host. How the anomaly score is calculated based on bucket span? can anyone explain me in detail?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![walterra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/walterra/32/139867_2.png) [@walterra](https://discuss.elastic.co/u/walterra)\
**Post date:** [October 17, 2018, 12:53pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/2 "2018-10-17T12:53:13Z")

</div>

We have this blog post which covers bucket spans: [https://www.elastic.co/blog/explaining-the-bucket-span-in-machine-learning-for-elasticsearch](https://www.elastic.co/blog/explaining-the-bucket-span-in-machine-learning-for-elasticsearch)

Let us know if that answers your questions or if you need more information.

Best,  
Walter

---

<div class="post-metadata">

**Author:** ![shiv94](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@shiv94](https://discuss.elastic.co/u/shiv94)\
**Post date:** [October 18, 2018, 1:57pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/3 "2018-10-18T13:57:59Z")

</div>

Thanks for your response.  
I have one doubt, If the data is less to analyze then bucket span should be more?

---

<div class="post-metadata">

**Author:** ![walterra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/walterra/32/139867_2.png) [@walterra](https://discuss.elastic.co/u/walterra)\
**Post date:** [October 19, 2018, 8:42am UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/4 "2018-10-19T08:42:19Z")

</div>

This depends on the use case. ML can work with sparse data too. But if you expect your data to be non-sparse in a given time frame then it makes sense to tweak the bucket span in that regard.

You can use the Single Metric Wizard to experiment with different bucket spans.

For example, this dataset shows gaps when the bucket span is only 1 second:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/f/6/f69232f93955d3acac7e9682666ba1b6b3f4a69d.png)

Changing the bucket span to 1 minute for the same dataset results in a continuous line:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/0/f/0fa00a5beaa466ae72439a31d8622235d9f6f61f.png)

Note that 1 second or 1 minute are not necessarily good or bad bucket spans in general, it depends on the type of data you have and the patterns you expect to emerge. The "Estimate bucket span" button provides a helper function that will try to come up with a reasonable bucket span by analysing the source data. It is available in all job wizards.

---

<div class="post-metadata">

**Author:** ![shiv94](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@shiv94](https://discuss.elastic.co/u/shiv94)\
**Post date:** [October 19, 2018, 3:28pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/5 "2018-10-19T15:28:52Z")

</div>

I created a multi-metric job with 1 month flow of data and the processed records are 10x,xxx,xxx but the metric viewer looks like this

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/2/520f08e24c45c6093828a0ca59de9306358643cc.png)

output after viewing results

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/0/505d97e74b66303da24e56c447457f8430cf61c0.png)

I created single- metric too

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/5/0/501832289d747fb373d6fd2c5d7a9180ae360639.png)

Can you let me know why?

---

<div class="post-metadata">

**Author:** ![walterra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/walterra/32/139867_2.png) [@walterra](https://discuss.elastic.co/u/walterra)\
**Post date:** [October 23, 2018, 8:35am UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/6 "2018-10-23T08:35:16Z")

</div>

It's hard to tell this way what's wrong, can you also provide a screenshot of the results you're seeing in the previews in the job creation wizards? In addition to that it would be useful if you could post a sample document you're analyzing as well as the resulting job config JSON. Please also explain the use case you're working on, that will help me making better suggestions. Thanks, Walter

---

<div class="post-metadata">

**Author:** ![shiv94](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@shiv94](https://discuss.elastic.co/u/shiv94)\
**Post date:** [October 23, 2018, 2:44pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/7 "2018-10-23T14:44:18Z")

</div>

Hey, I am analyzing the avg of total cpu utilization on the fields system.cpu.total.pct and system.cpu.total.norm.pct by using metricbeat data. There is metric viewer for system.cpu.total.pct but not to system.cpu.total.norm.pct .

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/d/5/d58af03b9527bb5279f2bdb1de86c0ea7e0bb257.png)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/e/1/e1764343f4067212247952629fd94737976fb7aa.png)

Job Config Json

{  
"job\_id": "total",  
"job\_type": "anomaly\_detector",  
"job\_version": "6.4.1",  
"description": "",  
"create\_time": 1540299791986,  
"finished\_time": 1540303801057,  
"established\_model\_memory": 2687544,  
"analysis\_config": {  
"bucket\_span": "15m",  
"detectors": [  
{  
"detector\_description": "mean(system.cpu.total.pct)",  
"function": "mean",  
"field\_name": "system.cpu.total.pct",  
"partition\_field\_name": "beat.hostname",  
"detector\_index": 0  
},  
{  
"detector\_description": "mean(system.cpu.total.norm.pct)",  
"function": "mean",  
"field\_name": "system.cpu.total.norm.pct",  
"partition\_field\_name": "beat.hostname",  
"detector\_index": 1  
}  
],  
"influencers": [  
"beat.hostname"  
]  
},  
"analysis\_limits": {  
"model\_memory\_limit": "17mb",  
"categorization\_examples\_limit": 4  
},  
"data\_description": {  
"time\_field": "@timestamp",  
"time\_format": "epoch\_ms"  
},  
"model\_snapshot\_retention\_days": 1,  
"custom\_settings": {  
"created\_by": "multi-metric-wizard"  
},  
"model\_snapshot\_id": "1540303799",  
"results\_index\_name": "shared",  
"data\_counts": {  
"job\_id": "total",  
"processed\_record\_count": 109742474,  
"processed\_field\_count": 113691080,  
"input\_bytes": 7809229831,  
"input\_field\_count": 113691080,  
"invalid\_date\_count": 0,  
"missing\_field\_count": 215536342,  
"out\_of\_order\_timestamp\_count": 0,  
"empty\_bucket\_count": 31,  
"sparse\_bucket\_count": 3,  
"bucket\_count": 2356,  
"earliest\_record\_timestamp": 1538179200045,  
"latest\_record\_timestamp": 1540299624572,  
"last\_data\_time": 1540303799182,  
"latest\_empty\_bucket\_timestamp": 1538748900000,  
"latest\_sparse\_bucket\_timestamp": 1538720100000,  
"input\_record\_count": 109742474  
},  
"model\_size\_stats": {  
"job\_id": "total",  
"result\_type": "model\_size\_stats",  
"model\_bytes": 2687544,  
"total\_by\_field\_count": 94,  
"total\_over\_field\_count": 0,  
"total\_partition\_field\_count": 93,  
"bucket\_allocation\_failures\_count": 0,  
"memory\_status": "ok",  
"log\_time": 1540303799000,  
"timestamp": 1540298700000  
},  
"datafeed\_config": {  
"datafeed\_id": "datafeed-total",  
"job\_id": "total",  
"query\_delay": "116810ms",  
"indices": [  
"metricbeat-\*"  
],  
"types": ,  
"query": {  
"match\_all": {  
"boost": 1  
}  
},  
"scroll\_size": 1000,  
"chunking\_config": {  
"mode": "auto"  
},  
"state": "stopped"  
},  
"state": "closed"  
}  
Job results

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/1/4/14b7591ed349ba2d622fc060e0a72e867b00fce3.png)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/1/c118dc1fbc8ad7466d9effb9b0feb485a52f40ed.png)

No results found for system.cpu.total.norm.pct why?

---

<div class="post-metadata">

**Author:** ![walterra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/walterra/32/139867_2.png) [@walterra](https://discuss.elastic.co/u/walterra)\
**Post date:** [October 24, 2018, 1:38pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/8 "2018-10-24T13:38:19Z")

</div>

As far as I can tell, this is not an issue with the Machine Learning plugin by itself but rather with the data you have at hand. I just tested this with a default installation of metricbeat and there is simply no data saved to the field `system.cpu.total.norm.pct`.

metricbeat needs to be explicitly configured to write to that field, have a look at the docs here: [https://www.elastic.co/guide/en/beats/metricbeat/current/metricbeat-metricset-system-cpu.html#\_configuration\_3](https://www.elastic.co/guide/en/beats/metricbeat/current/metricbeat-metricset-system-cpu.html#_configuration_3)

once metricbeat is configured to log `normalized_percentages` that field should be populated, here's an example:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/2/7/27596fccf9f4aac53f0bd502547b6520264a912d.png)

As you can see in the screenshot above, the preview charts in the job creation wizard can help you identify if you're about to analyze the expected data, so if you don't see any data showing up there, a Machine Learning job run with that configuration will not return any results. So these preview charts, alongside with Machine Learning's Data Visualizer ([https://www.elastic.co/blog/machine-learning-data-visualizer-and-modules](https://www.elastic.co/blog/machine-learning-data-visualizer-and-modules)) can help you identify the characteristics of your source data.

---

<div class="post-metadata">

**Author:** ![shiv94](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@shiv94](https://discuss.elastic.co/u/shiv94)\
**Post date:** [October 24, 2018, 1:51pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/9 "2018-10-24T13:51:54Z")

</div>

Thank you, but I can see this field in discover. how is that possible?

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/e/7/e7386d5d931fc215f04a8673d87441a18f0bff85.png)

---

<div class="post-metadata">

**Author:** ![walterra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/walterra/32/139867_2.png) [@walterra](https://discuss.elastic.co/u/walterra)\
**Post date:** [October 24, 2018, 2:50pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/10 "2018-10-24T14:50:07Z")

</div>

As far as I can see, the field you selected in your last screenshot from Discover is a different one (`system.process.cpu.total.norm.pct`) compared to the ones you used in the Machine Learning job configs (`system.cpu.total.norm.pct`). The "process" based one is available by default.

---

<div class="post-metadata">

**Author:** ![shiv94](https://avatars.discourse-cdn.com/v4/letter/s/bc8723/32.png) [@shiv94](https://discuss.elastic.co/u/shiv94)\
**Post date:** [October 24, 2018, 3:01pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/11 "2018-10-24T15:01:20Z")

</div>

yeah sorry, got confused. I made changes in system.yml file it worked.  
Thanks your help

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 21, 2018, 3:01pm UTC](https://discuss.elastic.co/t/created-ml-job-with-bucket-span-15m-and-1d/152839/12 "2018-11-21T15:01:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
