# Logstash - aggregate results

**URL:** <https://discuss.elastic.co/t/logstash-aggregate-results/164351>\
**Category:** Logstash\
**Created:** [January 15, 2019, 6:24pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351 "2019-01-15T18:24:45Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 15, 2019, 6:24pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/1 "2019-01-15T18:24:45Z")

</div>

Hello,

When I am using the aggregate filter in logstash (to get the total sales of each product, for example), is there a way to send the aggregated results to a different output and not line by line?

Thank you!

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 15, 2019, 6:30pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/2 "2019-01-15T18:30:04Z")

</div>

Yes, you can drop the lines that are not aggregated and just keep the aggregations. See [here](https://discuss.elastic.co/t/if-make-same-head-multiline-codec/164227) for an example. If you need more details then you need to show us your input data and what your current aggregate filter configuration looks like.

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 16, 2019, 10:28am UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/3 "2019-01-16T10:28:29Z")

</div>

Thank you!! We tried to implement the suggestion you gave. However, we have a large quantity of data and we want to read all the data and, in the end, show the aggregation results. We tried to maximize the timeout, but we never know when it will have the final results.

Here it is the code, if you can help us, we appreciate it. We just want to output the aggregation results.

> input{  
> elasticsearch{  
> "hosts" =\> "XXX"  
> "index" =\> "logs\_0"  
> schedule =\> "\* \* \* \* \*"  
> }  
> }  
> filter {  
> aggregate {  
> task\_id =\> "%{subProcessName}"  
> code =\> "map['total'] ||= 0; map['total'] += event.get('executionTime');"  
> push\_map\_as\_event\_on\_timeout =\> true  
> timeout\_task\_id\_field =\> "subProcessName"  
> timeout =\> 4  
> timeout\_tags =\> ['\_aggregatetimeout']  
> timeout\_code =\> "event.set('[@metadata][wanted]', 1)"  
> }  
> if [@metadata][wanted] != 1 { drop {} }  
> }  
> output{  
> if "\_aggregatetimeout" in [tags] {  
> elasticsearch{  
> "hosts" =\> "XXX"  
> "index" =\> "logs\_4"  
> }  
> }  
> }

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 16, 2019, 1:29pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/4 "2019-01-16T13:29:37Z")

</div>

If I am reading it correctly, that elasticsearch input re-reads the entire index once a minute. Is that right? The aggregate looks right, although a 4 second timeout is not very long. What problem are you having?

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 16, 2019, 2:33pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/5 "2019-01-16T14:33:45Z")

</div>

Yes, we notice that schedule does not make sense here, we deleted.  
The problem is that logstash takes some time to get the results from elasticsearch, and if that timeout defined is higher that the time that logstash takes to read elasticsearch, the logstash will stop before showing any aggregation results. If I considered a timeout like 10 seconds, the aggregation results will be repeated each 10 seconds. I want to have like:  
process1=\> 10,  
process2=\> 20  
and I am getting  
process1=\>5,  
process1=\>5,  
product2=\>10,  
product2=\>10  
(sometimes it loses itself in the sum and we lose data)

Any idea?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 16, 2019, 2:57pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/6 "2019-01-16T14:57:13Z")

</div>

> [@Francisca\_Lima](#):
>
> if that timeout defined is higher that the time that logstash takes to read elasticsearch, the logstash will stop before showing any aggregation results

lower, not higher. Why not use an 1800 second timeout?

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 16, 2019, 3:01pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/7 "2019-01-16T15:01:01Z")

</div>

But if logstash stops before that 1800 seconds, the results will not appear as desired because logstash execution stopped as soon as read everything from elasticsearch. Right?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 16, 2019, 3:11pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/8 "2019-01-16T15:11:03Z")

</div>

Can you use [push\_previous\_map\_as\_event](https://www.elastic.co/guide/en/logstash/current/plugins-filters-aggregate.html#plugins-filters-aggregate-push_previous_map_as_event)? That will cause logstash to flush the map when it exits.

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 16, 2019, 4:01pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/9 "2019-01-16T16:01:41Z")

</div>

We found the solution ordering the data from elasticsearch, because every time a new process came, the results were presented and the aggregations restarted. Thanks!

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 16, 2019, 5:18pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/10 "2019-01-16T17:18:03Z")

</div>

Is it possible to aggregate results (like process and total value) and then for each process do a lookup to elasticsearch and an aggregate to get as final result a list of: process, processID, total value?  
I tried to do another aggregation like this by processID, but in the final result it only appears the first aggregation and an empty process.  
I want to replicate each line of the first aggregation by the processID.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 16, 2019, 5:32pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/11 "2019-01-16T17:32:45Z")

</div>

I cannot really answer that without understanding more about the data layout.

Note that you can stash additional fields in the map, even if they are constant for a given task id, and they will get included in the aggregated event.

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 17, 2019, 1:52pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/12 "2019-01-17T13:52:53Z")

</div>

In my datasource I have like:  
process\_name ; process\_id ; value  
process1; 1; 10  
process2; 2; 20  
process1; 3; 5  
process2; 4; 10

And I need the avg by process\_name and the value by process\_id to compare, like:  
process\_name; process\_id; value; avg  
process1; 1; 10; 7.5  
process1; 3; 5; 7.5  
process2; 2; 20; 15  
process2; 4; 10; 15

So, we want to add the field avg (by process\_name), which we calculate in the aggregation, in each line. And, this avg value has to be replicated by each process\_id with the same process\_name.  
We tried to create two aggregations, with elasticsearch lookup, but no success yet.  
Can you help us?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 17, 2019, 4:11pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/13 "2019-01-17T16:11:59Z")

</div>

I see you asked a question about getting elasticsearch to do the aggregation. That is fundamentally a better approach that this. That said, if I have a csv like this

```
process_name,process_id,value
process1,1,10
process2,2,20
process1,3,5
process2,4,10

```

I can run a filter like this

```
csv { autodetect_column_names => true }
aggregate {
    task_id => "%{process_name}"
    code => "
        map['lines'] ||= 0; map['lines'] += 1;
        map['total'] ||= 0; map['total'] += event.get('value').to_i;
    "
    push_map_as_event_on_timeout => true
    timeout_task_id_field => "processName"
    timeout => 6 # seconds
    timeout_code => "event.set('average', event.get('total').to_f/event.get('lines').to_f)"
}

```

That will output the 4 lines from the csv, and then several seconds later output 2 more events having average set to 7.5 and 15.0.

The only reason I used a different value for timeout\_task\_id\_field is to make it clear what that option does.

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 17, 2019, 4:39pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/14 "2019-01-17T16:39:13Z")

</div>

Thanks! However, we want to keep process\_id in the output, and that's why it doesn't work. How we can add the process\_id and maintain the average by process\_name?  
We want to have this output:  
process\_name; process\_id; value; avg  
process1; 1; 10; 7.5  
process1; 3; 5; 7.5  
process2; 2; 20; 15  
process2; 4; 10; 15

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 17, 2019, 4:50pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/15 "2019-01-17T16:50:25Z")

</div>

Well, it could be done. You could ingest each row into an array in map, and then at the end iterate over the array and spit out each row with the average appended. But that's a terrible design (ingesting the entire data set into memory), and I'm not going to encourage you to do it by writing the code for you 😃 Doing the aggregation in elasticsearch is a far better idea.

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 17, 2019, 4:53pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/16 "2019-01-17T16:53:18Z")

</div>

Thanks! I also tried to aggregate in elasticsearch but I am facing other issue. Can you please see this topic ? [Logstash - aggregation\_fields in elasticsearch filter plugin](https://discuss.elastic.co/t/logstash-aggregation-fields-in-elasticsearch-filter-plugin/164672)  
Maybe I am missing something.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [January 17, 2019, 5:08pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/17 "2019-01-17T17:08:12Z")

</div>

Yeah, I looked at that.The configuration looks correct, given the es output you show. I don't have an elasticsearch server I can experiment against.

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 17, 2019, 5:28pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/18 "2019-01-17T17:28:39Z")

</div>

Thanks a lot!! 😃 We will keep trying!

---

<div class="post-metadata">

**Author:** ![Francisca\_Lima](https://avatars.discourse-cdn.com/v4/letter/f/edb3f5/32.png) [@Francisca\_Lima](https://discuss.elastic.co/u/Francisca_Lima)\
**Post date:** [January 17, 2019, 6:12pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/19 "2019-01-17T18:12:49Z")

</div>

We found the solution, we were using query instead of query\_template in the elasticsearch filter plugin.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 14, 2019, 6:12pm UTC](https://discuss.elastic.co/t/logstash-aggregate-results/164351/20 "2019-02-14T18:12:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
