# ML jobs with missing documents

**URL:** <https://discuss.elastic.co/t/ml-jobs-with-missing-documents/239966>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [July 6, 2020, 6:56am UTC](https://discuss.elastic.co/t/ml-jobs-with-missing-documents/239966 "2020-07-06T06:56:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![saraKM](https://avatars.discourse-cdn.com/v4/letter/s/ecccb3/32.png) [@saraKM](https://discuss.elastic.co/u/saraKM)\
**Post date:** [July 6, 2020, 6:56am UTC](https://discuss.elastic.co/t/ml-jobs-with-missing-documents/239966/1 "2020-07-06T06:56:56Z")

</div>

We have some jobs running with high numbers of documents missing. As per elastic ML documents, those checks are done after buckets with the missed documents have been processed and anomaly scores are finalized and “If there is indeed missing data due to their ingest delay, the end user is notified”. The question is how we can make sure not missing any document. Increasing the query delay usually works, but we need to make sure those lost ones are processed at the end (If notified soon enough, how stopping/starting datafeed to consider those time ranges impact the ML model and results? duplicate processing?)

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 6, 2020, 3:19pm UTC](https://discuss.elastic.co/t/ml-jobs-with-missing-documents/239966/2 "2020-07-06T15:19:05Z")

</div>

In general, you want your `query_delay` to be set as high as possible to avoid missing documents due to ingest delays.

The ML job will not re-process past buckets unless you manually use the ML [Model Snapshots API](https://www.elastic.co/guide/en/elasticsearch/reference/current/ml-apis.html#ml-api-snapshot-endpoint) to revert the job to a model that was saved before your data was missed. You could pass the `delete_intervening_results` flag to delete any anomalies that surfaced since that time.

After this, you could re-start the datafeed from that moment moving forward.

---

<div class="post-metadata">

**Author:** ![saraKM](https://avatars.discourse-cdn.com/v4/letter/s/ecccb3/32.png) [@saraKM](https://discuss.elastic.co/u/saraKM)\
**Post date:** [July 6, 2020, 11:58pm UTC](https://discuss.elastic.co/t/ml-jobs-with-missing-documents/239966/3 "2020-07-06T23:58:32Z")

</div>

Many thanks for your quick response 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 3, 2020, 11:59pm UTC](https://discuss.elastic.co/t/ml-jobs-with-missing-documents/239966/4 "2020-08-03T23:59:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
