# Best practices for preprocessing data and monitoring resource usage in predefined ML jobs (security:host)

**URL:** <https://discuss.elastic.co/t/best-practices-for-preprocessing-data-and-monitoring-resource-usage-in-predefined-ml-jobs-security-host/382320>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [September 30, 2025, 1:31pm UTC](https://discuss.elastic.co/t/best-practices-for-preprocessing-data-and-monitoring-resource-usage-in-predefined-ml-jobs-security-host/382320 "2025-09-30T13:31:32Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![ilyes](https://avatars.discourse-cdn.com/v4/letter/i/f0a364/32.png) [@ilyes](https://discuss.elastic.co/u/ilyes)\
**Post date:** [September 30, 2025, 1:31pm UTC](https://discuss.elastic.co/t/best-practices-for-preprocessing-data-and-monitoring-resource-usage-in-predefined-ml-jobs-security-host/382320/1 "2025-09-30T13:31:32Z")

</div>

**Question 1:**  
I am planning to use the predefined ML job from the _security:host_ module. To avoid overloading the ML model with too much data, what would be the best approach in terms of data preprocessing?

- Should I first store the data in a saved object before using it as the data source for the ML job?

- Or would it be better to create an ingest pipeline?

- Or should I prepare the data through a transforms job before feeding it into the ML model?

**Question 2:**  
Does Elasticsearch/Kibana provide dashboards for ML jobs? If not, how can I best monitor and evaluate the metrics (e.g., resource consumption) of an ML job so I can properly assess the performance of my tests?
