# Easy way to insert top level query aggregation details back into elastic

**URL:** https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544
**Category:** Elasticsearch
**Created:** [February 11, 2016, 6:50pm UTC](https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544 "2016-02-11T18:50:08Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![Sandeep\_Takhar](https://avatars.discourse-cdn.com/v4/letter/s/a6a055/32.png) [@Sandeep\_Takhar](https://discuss.elastic.co/u/Sandeep_Takhar)
#### Post date: [February 11, 2016, 6:50pm UTC](https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544/1 "2016-02-11T18:50:08Z")

</div>

Hi.

Been following Zach's pretty cool work around getting the ebay anomaly detection algorithm to work in elastic.

> **[Implementing a Statistical Anomaly Detector in Elasticsearch - Part 1
	  	 |...](https://www.elastic.co/blog/implementing-a-statistical-anomaly-detector-part-1)**
>
> This graph shows the min/max/avg of 45 million data points (75,000 individual time series over 600 hours). There are eight large-scale, simulated disruptions in this graph...can you spot them? No? It’...

In his example he has three level grouping. I just want the top level term and ninetieth\_surprise and send that back into elastic and looking for ideas. I've searched around, but not finding much.

I have the example working with my data, but all I did was change three values to match my field names.

Here is Zach's example:

There are 5 "metrics" with 5 ninetieth\_percentiles in the json (in my json I get two ninetieth percentiles for each top level for some reason). My json return has all the bottom level buckets. I just want to insert the top level metric name and ninetieth\_percentile..which is part of the top level bucket.

{  
"query": {  
"filtered": {  
"filter": {  
"range": {  
"hour": {  
"gte": "{{start}}",  
"lte": "{{end}}"  
}  
}  
}  
}  
},  
"size": 0,  
"aggs": {  
"metrics": {  
"terms": {  
"field": "metric",  
“size”: 5  
},  
"aggs": {  
"queries": {  
"terms": {  
"field": "query",  
"size": 500  
},  
"aggs": {  
"series": {  
"date\_histogram": {  
"field": "hour",  
"interval": "hour"  
},  
"aggs": {  
"avg": {  
"avg": {  
"field": "value"  
}  
},  
"movavg": {  
"moving\_avg": {  
"buckets\_path": "avg",  
"window": 24,  
"model": "simple"  
}  
},  
"surprise": {  
"bucket\_script": {  
"buckets\_path": {  
"avg": "avg",  
"movavg": "movavg"  
},  
"script": "(avg - movavg).abs()"  
}  
}  
}  
},  
"largest\_surprise": {  
"max\_bucket": {  
"buckets\_path": "series.surprise"  
}  
}  
}  
},  
"ninetieth\_surprise": {  
"percentiles\_bucket": {  
"buckets\_path": "queries\>largest\_surprise",  
"percents": [  
90.0  
]  
}  
}  
}  
}  
}  
}

---

<div class="post-metadata">

### Author: ![Sandeep\_Takhar](https://avatars.discourse-cdn.com/v4/letter/s/a6a055/32.png) [@Sandeep\_Takhar](https://discuss.elastic.co/u/Sandeep_Takhar)
#### Post date: [February 12, 2016, 12:39am UTC](https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544/2 "2016-02-12T00:39:39Z")

</div>

One thing I'm doing is to use filter\_path. Actually Zach mentioned it in his second article and it took me a while to find it. My top level field name is different, but it looks like this:

I think I'll just use logstash to read this file that I output and dump it into elastic...I've already got a framework for doing just that and I've dealt with json objects before.

pretty=true&human=false&flat\_settings=true&filter\_path=aggregations.agent\_names.buckets.key,aggregations.agent\_names.buckets.ninetieth\_surprise.values

Don't know if it's a good way or not to do it, but I'll give it a try.

---

<div class="post-metadata">

### Author: ![Sandeep\_Takhar](https://avatars.discourse-cdn.com/v4/letter/s/a6a055/32.png) [@Sandeep\_Takhar](https://discuss.elastic.co/u/Sandeep_Takhar)
#### Post date: [February 12, 2016, 2:22am UTC](https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544/3 "2016-02-12T02:22:00Z")

</div>

Here is how I flattened out and split the resulting output as well...now I will just send to elastic using the output plugin. Again..not sure if best way, but it works:

input {  
stdin { codec =\> json }  
}

#filter{

# grok{

# match =\> ["message","%{GREEDYDATA:msg}"]

# }

# #trick to reparse the message from text file (brought in as text)

# json {

# source =\> "msg"

# }

#}

filter  
{  
mutate {  
rename =\> [  
"[aggregations][agent\_names][buckets]", "buckets"  
]  
remove\_field =\> "aggregations"  
}  
}

filter {  
split {  
field =\> "buckets"  
}  
}

filter  
{  
mutate {  
rename =\> [  
"[buckets][key]", "agent\_name",  
"[buckets][ninetieth\_surprise][values][90.0]", "ninetieth\_surprise"  
]  
remove\_field =\> "buckets"  
}  
}

output {  
stdout { codec =\> rubydebug }  
}

---

<div class="post-metadata">

### Author: ![Sandeep\_Takhar](https://avatars.discourse-cdn.com/v4/letter/s/a6a055/32.png) [@Sandeep\_Takhar](https://discuss.elastic.co/u/Sandeep_Takhar)
#### Post date: [February 12, 2016, 1:39pm UTC](https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544/4 "2016-02-12T13:39:31Z")

</div>

Here is the data I get using filter\_data, or at least part of it..than can be used with above config. In case anyone is following. I'll create a cron job now and see how the data looks like in timelion. There are plenty of null values because it is staging environment and it's not very busy at the moment.

{"aggregations":{"agent\_names":{"buckets":[{"key":"custmgtbill2srv1","ninetieth\_surprise":{"values":{"90.0":0.00917166937529728}},"ninetieth\_surprise":{"values":{"90.0":0.00917166937529728}}},{"key":"entmarketoffersvc3","ninetieth\_surprise":{"values":{"90.0":0.016666666666666666}},"ninetieth\_surprise":{"values":{"90.0":0.016666666666666666}}},{"key":"custmgtconsumersvc1","ninetieth\_surprise":{"values":{"90.0":0.017316291670002933}},"ninetieth\_surprise":{"values":{"90.0":0.017316291670002933}}},{"key":"custmgtconsumersvc4","ninetieth\_surprise":{"values":{"90.0":0.01162318909094097}},"ninetieth\_surprise":{"values":{"90.0":0.01162318909094097}}},{"key":"billpresentmentweb3","ninetieth\_surprise":{"values":{"90.0":0.0}},"ninetieth\_surprise":{"values":{"90.0":0.0}}},{"key":"custmgtconsumerweb4","ninetieth\_surprise":{"values":{"90.0":1.4005602240896359E-6}},"ninetieth\_surprise":{"values":{"90.0":1.4005602240896359E-6}}},{"key":"custmgtfulweb1","ninetieth\_surprise":{"values":{"90.0":0.0}},"ninetieth\_surprise":{"values":{"90.0":0.0}}},{"key":"custmgtfulweb3","ninetieth\_surprise":{"values":{"90.0":0.0}},"ninetieth\_surprise":{"values":{"90.0":0.0}}},{"key":"custmgtfulweb4","ninetieth\_surprise":{"values":{"90.0":0.0}},"ninetieth\_surprise":{"values":{"90.0":0.0}}}]}}}

---

<div class="post-metadata">

### Author: ![Sandeep\_Takhar](https://avatars.discourse-cdn.com/v4/letter/s/a6a055/32.png) [@Sandeep\_Takhar](https://discuss.elastic.co/u/Sandeep_Takhar)
#### Post date: [February 12, 2016, 11:22pm UTC](https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544/5 "2016-02-12T23:22:56Z")

</div>

I'll also post the curl which I use that is in a cronjob in case someone runs into this post to complete out the turn-key solution (?)

r" : {  
"and" : [ {  
"range" : {  
"@timestamp" : {  
"gte": "now-1h",  
"lte": "now"  
}  
}  
},  
{  
"terms" : { "metric\_name" : ["durationmean"]}  
} ]  
}  
}  
},  
"size": 0,  
"aggs": {  
"agent\_names": {  
"terms": {  
"field": "agent\_name",  
"size": 5000  
},  
"aggs": {  
"metric\_names": {  
"terms": {  
"field": "metric\_name",  
"size": 10000  
},  
"aggs": {  
"series": {  
"date\_histogram": {  
"field": "@timestamp",  
"interval": "minute"  
},  
"aggs": {  
"avg": {  
"avg": {  
"field": "metric\_value"  
}  
},  
"movavg": {  
"moving\_avg": {  
"buckets\_path": "avg",  
"window": 60,  
"model": "simple"  
}  
},  
"surprise": {  
"bucket\_script": {  
"buckets\_path": {  
"avg": "avg",  
"movavg": "movavg"  
},  
"script": "(avg - movavg).abs()"  
}  
}  
}  
},  
"largest\_surprise": {  
"max\_bucket": {  
"buckets\_path": "series.surprise"  
}  
}  
}  
},  
"ninetieth\_surprise": {  
"percentiles\_bucket": {  
"buckets\_path": "metric\_names\>largest\_surprise",  
"percents": [  
90.0  
]  
}  
}  
}  
}  
}  
}'

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 5, 2017, 11:16pm UTC](https://discuss.elastic.co/t/easy-way-to-insert-top-level-query-aggregation-details-back-into-elastic/41544/6 "2017-07-05T23:16:40Z")

</div>


