# Anomaly Detection Kibana skipping data

**URL:** <https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200>\
**Category:** Kibana\
**Tags:** elastic-stack-machine-learning\
**Created:** [June 1, 2020, 3:39pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200 "2020-06-01T15:39:50Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![alexisdrnt](https://avatars.discourse-cdn.com/v4/letter/a/898d66/32.png) [@alexisdrnt](https://discuss.elastic.co/u/alexisdrnt)\
**Post date:** [June 1, 2020, 3:39pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/1 "2020-06-01T15:39:50Z")

</div>

I'm running a ML job with a simple `low_count` detector, with a **daily** bucket, and I have an **hourly** frequency for my datafeed.  
Everything seem to work correctly, however sometimes for a certain reason some data get skipped.  
The ML job throws an anomaly. When I look at the count everything is normal, but the anomaly still shows `actual 0`

 ![Screen Shot 2020-06-01 at 8.31.58 AM](https://us1.discourse-cdn.com/elastic/original/3X/0/5/055c6310a230b0f5b8654b084cdada7a94991478.png)  
Looking at the `view series`, I have:  
 ![Screen Shot 2020-06-01 at 8.32.53 AM](https://us1.discourse-cdn.com/elastic/original/3X/3/d/3d59c2311da3f759a6621cbaa7e40afeb2de573c.png)  
You can clearly see that the data count is not 0, but 8.

I have a **5min query\_delay** , but I checked at what time the data has been inserted, and it was in the middle of the bucket\_span.So for sure the data didn't get ingested after the ML processed it.  
Any idea?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 1, 2020, 4:34pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/2 "2020-06-01T16:34:25Z")

</div>

What version? This sounds similar to a bug that was introduced in v6.5, but fixed in v6.6. See: [Machine Learning datafeed skipping documents that seem to be there](https://discuss.elastic.co/t/machine-learning-datafeed-skipping-documents-that-seem-to-be-there/170773/2)

---

<div class="post-metadata">

**Author:** ![alexisdrnt](https://avatars.discourse-cdn.com/v4/letter/a/898d66/32.png) [@alexisdrnt](https://discuss.elastic.co/u/alexisdrnt)\
**Post date:** [June 1, 2020, 4:39pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/3 "2020-06-01T16:39:00Z")

</div>

ES and Kibana are running on `7.5.2`

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 1, 2020, 8:27pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/4 "2020-06-01T20:27:45Z")

</div>

Hmm...odd. What do you see if you run:

```auto
GET .ml-anomalies-*/_search
{
    "query": {
            "bool": {
              "filter": [
                  { "term" : { "result_type" : "bucket"}},
                  { "range" : { "timestamp" : { "gte": "now-3d" } } },
                  { "range" : { "anomaly_score" : { "gte": "90" } } }
                  ]
            }
    }
}

```

Specifically looking for this anomaly record's value of the field `event_count`

---

<div class="post-metadata">

**Author:** ![alexisdrnt](https://avatars.discourse-cdn.com/v4/letter/a/898d66/32.png) [@alexisdrnt](https://discuss.elastic.co/u/alexisdrnt)\
**Post date:** [June 1, 2020, 8:52pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/5 "2020-06-01T20:52:55Z")

</div>

When I run that I don't have any result

```auto
{
  "took" : 0,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 0,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : []
  }
}

```

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 1, 2020, 9:50pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/6 "2020-06-01T21:50:51Z")

</div>

Ok, well that doesn't make sense unless the anomaly in your screenshot is gone, or the score has changed to be below 90 now. Perhaps verify that that this is or is not the case. If the score has changed then just adjust the `"gte": "90" ` in the query

---

<div class="post-metadata">

**Author:** ![alexisdrnt](https://avatars.discourse-cdn.com/v4/letter/a/898d66/32.png) [@alexisdrnt](https://discuss.elastic.co/u/alexisdrnt)\
**Post date:** [June 1, 2020, 10:10pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/7 "2020-06-01T22:10:27Z")

</div>

The previous graph that I posted were related to `record` result\_type

---

<div class="post-metadata">

**Author:** ![alexisdrnt](https://avatars.discourse-cdn.com/v4/letter/a/898d66/32.png) [@alexisdrnt](https://discuss.elastic.co/u/alexisdrnt)\
**Post date:** [June 1, 2020, 11:04pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/8 "2020-06-01T23:04:44Z")

</div>

After looking at the `event_count`, the Anomaly Detection job missed ~2000 documents.  
I'm not sure how to determine the reason.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 2, 2020, 12:32pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/9 "2020-06-02T12:32:14Z")

</div>

Ah, yes sorry about the mistake on the `result_type`.

So, as you may well know, the `event_count` per bucket is the number of events present in the index for the bucket\_span when the datafeed's query is executed (and thus those are the documents passed along to anomaly detection). If you have occasions in which the `event_count` is less than the number of docs in that timeframe (viewed retroactively) then this really does point to an ingest delay issue - which is mitigated by increasing the `query_delay`.

---

<div class="post-metadata">

**Author:** ![alexisdrnt](https://avatars.discourse-cdn.com/v4/letter/a/898d66/32.png) [@alexisdrnt](https://discuss.elastic.co/u/alexisdrnt)\
**Post date:** [June 8, 2020, 1:44am UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/10 "2020-06-08T01:44:25Z")

</div>

Could it be another reason?  
I increased the query delay, I checked the skipped documents: The ingestion time of theses documents happens in the middle of the bucket span. Anything else I can check?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 9, 2020, 3:46pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/11 "2020-06-09T15:46:59Z")

</div>

May be a relevant detail here - but while the ingestion time does matter, however, what matters more is the timestamp field used in the index pattern. If the document has a timestamp of `2020-06-09T15:44:40.608000Z` , but gets indexed 5 minutes later, the document will still be missed by the ML datafeed if the query\_delay isn't big enough

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 17, 2020, 7:17pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/12 "2020-06-17T19:17:52Z")

</div>

By the way - this blog might be useful to those that are not sure just how much their ingest delay is:

> **[Calculating ingest lag and storing ingest time in Elasticsearch to improve...](https://www.elastic.co/blog/calculating-ingest-lag-and-storing-ingest-time-in-elasticsearch-to-improve-observability)**
>
> Relying on remote-generated timestamps for monitoring and alerting may be risky. If there is a delay between the occurrence of a remote event and the event arriving to Elasticsearch, or if the time on a remote system is set incorrectly, then...

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 15, 2020, 7:17pm UTC](https://discuss.elastic.co/t/anomaly-detection-kibana-skipping-data/235200/13 "2020-07-15T19:17:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
