# Two subaggregation in datafeed

**URL:** <https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [July 2, 2018, 1:32pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217 "2018-07-02T13:32:40Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 2, 2018, 1:32pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/1 "2018-07-02T13:32:41Z")

</div>

Hi, I am trying to do ratio in my data feed. First I created date\_histogram with max agg (couse I got an error that it's needed) and then I am doing subaggregation by attrs.src\_ca\_name and want to count ratio of success calls. Everything looks fine, no error just the data feed preview is empty. Can ML parse two subaggregation?

```
..."aggregations": {
"buckets": {
  "date_histogram": {
    "field": "@timestamp",
    "interval": "15m",
    "time_zone": "UTC"
  },
  "aggregations": {
      "@timestamp": {
        "max": {
          "field": "@timestamp"
        }
      },
    "by_src": {
      "terms": {
        "field": "attrs.src_ca_name",
        "size": 20,
        "order": {
          "_count": "desc"
        }
      },
      "aggregations": {
       "justattempts": { "filter": { "term": { "type": "call-attempt" } } },
       "ratio" : {
         "bucket_script" : {
           "buckets_path": {
              "atmptcnt": "justattempts>_count",
              "totalcnt": "_count"
           },
           "script" : "params.atmptcnt * 100 / params.totalcnt"
         }....
```

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 3, 2018, 12:37pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/2 "2018-07-03T12:37:34Z")

</div>

Yes, the datafeed can use sub-aggregations. See this blog for insight: [https://www.elastic.co/blog/custom-elasticsearch-aggregations-for-machine-learning-jobs](https://www.elastic.co/blog/custom-elasticsearch-aggregations-for-machine-learning-jobs)

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 9, 2018, 12:32pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/3 "2018-07-09T12:32:54Z")

</div>

But when I do `search` it returns correct data

```
{
  "took": 448,
   "timed_out": false,
 "_shards": {
"total": 5,
"successful": 5,
"skipped": 0,
"failed": 0
 },
"hits": {
"total": 1331042,
"max_score": 1,
"hits": [
  {
    "_index": "logstash-2018.06.19",
    "_type": "sbc_event",
    "_id": "AWQXBrJ3oA32JM6LbXea",
    "_score": 1,
    "_source": {
      "attrs": {
        "dst_ca_name": "AAAA"
      }
    }
  },
..........
 "aggregations": {
"buckets": {
  "buckets": [
    {
      "key_as_string": "2018-06-19T00:00:00.000Z",
      "key": 1529366400000,
      "doc_count": 274,
      "@timestamp": {
        "value": 1529366699000,
        "value_as_string": "2018-06-19T00:04:59.000Z"
      },
      "by_src": {
        "doc_count_error_upper_bound": 0,
        "sum_other_doc_count": 0,
        "buckets": [
          {
            "key": "BBBBB",
            "doc_count": 224,
            "justattempts": {
              "doc_count": 217
            },
            "ratio": {
              "value": 96.875
            }
          },

```

But when I run this in ML job it returns nothing.

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 9, 2018, 12:35pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/4 "2018-07-09T12:35:54Z")

</div>

Here is my whole ML datafeed ratio

```
PUT _xpack/ml/datafeeds/datafeed-ratio/
{
"job_id": "ratio_ca",
"indices": [
"logstash-2018.06.19"
],
"types": [
"doc"
],
"query": {
"bool": {
  "must": [
    {
      "terms": {"type":["call-attempt","call-end"]}
    }

  ],
  "must_not": []
}
},
"aggregations": {
"buckets": {
  "date_histogram": {
    "field": "@timestamp",
    "interval": "15m",
    "time_zone": "UTC"
  },
  "aggregations": {
      "@timestamp": {
        "max": {
          "field": "@timestamp"
        }
      },
    "by_src": {
      "terms": {
        "field": "attrs.src_ca_name",
        "size": 20,
        "order": {
          "_count": "desc"
        }
      },
      "aggregations": {
       "justattempts": { "filter": { "term": { "type": "call-attempt" } } },
       "ratio" : {
         "bucket_script" : {
           "buckets_path": {
              "atmptcnt": "justattempts>_count",
              "totalcnt": "_count"
           },
           "script" : "params.atmptcnt * 100 / params.totalcnt"
         }
       }
     }
    }
}}
   }
}
```

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 9, 2018, 12:52pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/5 "2018-07-09T12:52:02Z")

</div>

Can you paste the output from the following command in DevTools Console?

```auto
GET _xpack/ml/datafeeds/datafeed-ratio/_preview

```

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 9, 2018, 12:53pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/6 "2018-07-09T12:53:07Z")

</div>

> [@richcollier](#):
>
> GET \_xpack/ml/datafeeds/datafeed-ratio/\_preview

it's empty

```
[]

```

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 9, 2018, 12:53pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/7 "2018-07-09T12:53:11Z")

</div>

And please tell us what version of Elastic Stack you are using...

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 9, 2018, 12:54pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/8 "2018-07-09T12:54:22Z")

</div>

Version: 6.1.1

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 9, 2018, 12:56pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/9 "2018-07-09T12:56:14Z")

</div>

Ok - let me look more closely at your code and I'll see if I can replicate

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 9, 2018, 1:25pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/10 "2018-07-09T13:25:49Z")

</div>

Please show me the config of the ML job itself:

`GET _xpack/ml/anomaly_detectors/ratio_ca?pretty`

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 9, 2018, 1:27pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/11 "2018-07-09T13:27:24Z")

</div>

```
  {
 "count": 1,
"jobs": [
{
  "job_id": "ratio_ca",
  "job_type": "anomaly_detector",
  "job_version": "6.1.1",
  "description": "Ratio ca",
  "create_time": 1530537786467,
  "analysis_config": {
    "bucket_span": "15m",
    "summary_count_field_name": "doc_count",
    "detectors": [
      {
        "detector_description": "sum(ratio_ca)",
        "function": "sum",
        "field_name": "ratio_ca2",
        "detector_rules": [],
        "detector_index": 0
      }
    ],
    "influencers": []
  },
  "analysis_limits": {
    "model_memory_limit": "1024mb"
  },
  "data_description": {
    "time_field": "@timestamp",
    "time_format": "epoch_ms"
  },
  "model_plot_config": {
    "enabled": true
  },
  "model_snapshot_retention_days": 1,
  "results_index_name": "shared"
}
]
}
```

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 9, 2018, 1:31pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/12 "2018-07-09T13:31:55Z")

</div>

Hmm...your `detector` references the `field_name` of `ratio_ca2` but in your datafeed definition, the calculated field is just called `ratio`

They need to be the same

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 9, 2018, 1:55pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/13 "2018-07-09T13:55:41Z")

</div>

Thanks but it doesn't help

```
{
 "count": 1,
"jobs": [
{
  "job_id": "ratio_ca",
  "job_type": "anomaly_detector",
  "job_version": "6.1.1",
  "description": "Ratio ca",
  "create_time": 1531144344777,
  "analysis_config": {
    "bucket_span": "15m",
    "summary_count_field_name": "doc_count",
    "detectors": [
      {
        "detector_description": "sum(ratio_ca)",
        "function": "sum",
        "field_name": "ratio",
        "detector_rules": [],
        "detector_index": 0
      }
    ],
    "influencers": []
  },
  "analysis_limits": {
    "model_memory_limit": "1024mb"
  },
  "data_description": {
    "time_field": "@timestamp",
    "time_format": "epoch_ms"
  },
  "model_plot_config": {
    "enabled": true
  },
  "model_snapshot_retention_days": 1,
  "results_index_name": "shared"
}
 ]
}

```

still GET \_xpack/ml/datafeeds/datafeed-ratio/\_preview is empty

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 9, 2018, 3:02pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/14 "2018-07-09T15:02:43Z")

</div>

One other thing to notice is that you do a `terms` aggregation in the query, which implies you want separate analyses per `by_src`, but your `detector` makes no reference to this split. You might want to make your job config something like:

```auto
  "analysis_config": {
    "bucket_span": "15m",
    "summary_count_field_name": "doc_count",
    "detectors": [
      {
        "detector_description": "sum(ratio)",
        "function": "sum",
        "field_name": "ratio",
        "partition_field_name" : "by_src",
        "detector_rules": [],
        "detector_index": 0
      }
    ],
    "influencers": ["by_src"]
  },

```

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 9, 2018, 7:00pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/15 "2018-07-09T19:00:45Z")

</div>

Thank you I have set it but it's still not woking. ☹

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 10, 2018, 10:22am UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/16 "2018-07-10T10:22:57Z")

</div>

Sorry this is giving you trouble, but I cannot immediately see what your issue is and I cannot reproduce the problem given a similar situation - my setup works fine.

This hints at some subtle syntax error that's hard to spot.

May I suggest that you use the method of debugging by starting simple and progressing up to your desired end-state. So, for example, define the ML job, then define the ML datafeed without the `bucket_script` aggregation, just the `date_histogram`, the `max` on `@timestamp` and the `terms` aggregations.

Then run the datafeed `_preview` to see what you get (you should just get a bucketized count for each `by_src` similar to:

```auto
[
  {
    "@timestamp": 1486426496000,
    "by_src": "AAA",
    "doc_count": 15
  },
  {
    "@timestamp": 1486426496000,
    "by_src": "BBB",
    "doc_count": 11
  },

```

If you can get that, then move back to adding the `bucket_script` aggregation

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 11, 2018, 7:32am UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/17 "2018-07-11T07:32:10Z")

</div>

I have found an problem. It was the field

```
  "types": [
   "doc"
  ]

```

This should be the name of aggregation? If I run it without it, it works.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 11, 2018, 11:24am UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/18 "2018-07-11T11:24:43Z")

</div>

It is an index property that is a hold-over from pre-v6.x of elasticsearch and will be removed fully in v7.x:

[https://www.elastic.co/guide/en/elasticsearch/reference/6.2/ml-put-datafeed.html](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/ml-put-datafeed.html)

`types`  
(array) A list of types to search for within the specified indices. For example: []. This property is provided for backwards compatibility with releases earlier than 6.0.0. For more information, see [Removal of mapping types](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/removal-of-types.html).

---

<div class="post-metadata">

**Author:** ![Akaren](https://avatars.discourse-cdn.com/v4/letter/a/8dc957/32.png) [@Akaren](https://discuss.elastic.co/u/Akaren)\
**Post date:** [July 11, 2018, 6:22pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/19 "2018-07-11T18:22:38Z")

</div>

Thank you very much for your help!

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 30, 2018, 10:21pm UTC](https://discuss.elastic.co/t/two-subaggregation-in-datafeed/138217/20 "2018-10-30T22:21:38Z")

</div>


