# Machine Learning datafeed skipping documents that seem to be there

**URL:** <https://discuss.elastic.co/t/machine-learning-datafeed-skipping-documents-that-seem-to-be-there/170773>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [March 4, 2019, 6:08pm UTC](https://discuss.elastic.co/t/machine-learning-datafeed-skipping-documents-that-seem-to-be-there/170773 "2019-03-04T18:08:33Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [March 5, 2019, 1:49pm UTC](https://discuss.elastic.co/t/machine-learning-datafeed-skipping-documents-that-seem-to-be-there/170773/2 "2019-03-05T13:49:18Z")

</div>

I think there might be a few things going on here.

1. There is a bug that was introduced in v6.5 (and will be fixed in v6.6.2+) that inadvertently creates an anomaly on an interim (un-finalized, or still-open bucket). See:

> [@ML alerts triggering on interim result](https://discuss.elastic.co/t/ml-alerts-triggering-on-interim-result/158408):
>
> After upgrading from 6.3 to 6.5.0 I've been getting Watcher alerts from ML jobs that are false positives. If I click on the alert link fairly quickly I see 0 hits when there should be some large number, but it also says "Interim result". If I wait a few minutes and refresh, the anomaly score goes from 99 down to \<1 as it's found more than 0 hits. Is this an indication of a performance issue on my ES cluster or some changed behavior in ES - ML - Watcher interactions with 6.5?

and the corresponding bug:

> <https://github.com/elastic/ml-cpp/issues/324>
>
> \*\*How to reproduce\*\*
> 
> 1. Create a job with simple count detector and a bucket …span of 5m
> 2. Run some data through the job up to the end of a bucket (using the \`end\` parameter of the start datafeed API)
> 3. Open the job again (it should have been auto-closed from step 2)
> 4. Call the flush API:
> 
> \`\`\`
> POST \_xpack/anomaly\_detectors/{job\_id}/flush?advance\_time={time}&calc\_interim=true
> \`\`\`
> 
> where {time} should be a timestamp into the current bucket. E.g., if \`end\` was \`2018-12-01T00:00:00Z\`, {time} should be \`2018-12-01T00:00:01Z\`
> 
> \*\*Observed Behaviour\*\*
> 
> If you get the anomaly records, you should see a record which is interim and has an actual value of \`0.0\`. This shouldn't have been created. Interestingly, calling step 4 with {time} being one millisecond forward makes that record disappear.
> 
> Also, this is broken since version \`6.4.\`

1. The auto-annotation for missing data, however, should not stumble onto this bug because it explicitly ignores interim buckets. In order to validate your datafeed timing (what bucket's it's querying and when), you could enable TRACE logging for the datafeed:

```auto
PUT _cluster/settings
{
  "transient": {
    "logger.org.elasticsearch.xpack.ml.datafeed": "TRACE"
  }
}

```

(this is a transient setting that won't survive a cluster re-start but you can always reset this back to "DEBUG" or "NORMAL" when this experiment is over)

You can also have your Watch log what it sees as well - then, in the elasticsearch.log file we should have a better understanding of when the datafeed runs and what window of time it queries - while at the same time seeing the output of your watch that is trying to also do the validation.

---

_[View the full topic](https://discuss.elastic.co/t/machine-learning-datafeed-skipping-documents-that-seem-to-be-there/170773)._
