# ML datafeed bucket aggregation/script challenge

**URL:** <https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [February 17, 2022, 10:45pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552 "2022-02-17T22:45:46Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [February 17, 2022, 10:45pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/1 "2022-02-17T22:45:46Z")

</div>

I would like to create an anomaly detection job for the ratio of two counts. The logic is as follows...

A = count filtered by condition X  
B = count filtered by condition Y  
C = A/B

The datafeed would need to produce C as a field to be used in the `analysis_config`.

I have been trying to figure out some combination of bucket aggregation/bucket script/scripted fields that would work as a datafeed. The first challenge is whether it is even possible to have any query that produces two bucket aggregation values each derived from a different filter condition? If I can get these two values a bucket\_script would be easy enough to do the math.

Does anyone have any tips on how something like this can be done?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [February 18, 2022, 12:12pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/2 "2022-02-18T12:12:40Z")

</div>

Hi Rob - this should help: [Analyzing a ratio of documents over time with Anomaly Detection](https://discuss.elastic.co/t/analyzing-a-ratio-of-documents-over-time-with-anomaly-detection/225255)

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [February 18, 2022, 12:54pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/3 "2022-02-18T12:54:53Z")

</div>

@richcollier that was a big help. Thanks!

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [March 6, 2022, 4:48pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/4 "2022-03-06T16:48:43Z")

</div>

Is there any way to get a field into the aggregation-based datafeed that can be used for as a `partition_field`? It is possible to add levels of aggregation that create the buckets. However I haven't been able to get the `key` for the buckets as a field that can be accessed in the `detectors` config.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [March 7, 2022, 5:50pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/5 "2022-03-07T17:50:21Z")

</div>

Robert, here's an example using the demo dataset [farequote](https://raw.githubusercontent.com/PacktPublishing/Machine-Learning-with-Elastic-Stack-Second-Edition/main/Appendix/farequote-2021.csv):

(note I put a `terms` agg of size `100` but there are only like 19 different airlines)

```auto
PUT _ml/anomaly_detectors/farequote_terms_agg
{
  "analysis_config": {
    "bucket_span": "5m",
    "detectors": [{
      "function": "mean",
      "field_name": "responsetime",  
      "partition_field_name": "airline"  
    }],
   "influencers" : ["airline"],
    "summary_count_field_name": "doc_count"
  },
  "data_description": {
    "time_field":"@timestamp"  
  },
  "datafeed_config":{
    "indices": ["farequote"],
    "aggregations": {
      "buckets": {
        "date_histogram": {
          "field": "@timestamp",
          "fixed_interval": "5m",
          "time_zone": "UTC"
        },
        "aggregations": {
          "@timestamp": {  
            "max": {"field": "@timestamp"}
          },
          "airline": {  
            "terms": {
             "field": "airline",
              "size": 100
            },
            "aggregations": {
              "responsetime": {  
                "avg": {
                  "field": "responsetime"
                }
              }
            }
          }
        }
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [March 7, 2022, 7:47pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/6 "2022-03-07T19:47:14Z")

</div>

That worked. The location in the hierarchy of aggregations is important. The key here was that it needs to be at the same level as `@timestamp`, i.e. "inside" the date histogram.

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [March 7, 2022, 9:35pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/7 "2022-03-07T21:35:41Z")

</div>

Here is the final result:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/3/c3ed11c9db8ccf037356ea80bf35da31df0db7b1.png)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 4, 2022, 9:35pm UTC](https://discuss.elastic.co/t/ml-datafeed-bucket-aggregation-script-challenge/297552/8 "2022-04-04T21:35:46Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
