# Logstash aggregate plugin with the same key

**URL:** <https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612>\
**Category:** Logstash\
**Created:** [July 15, 2019, 8:47pm UTC](https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612 "2019-07-15T20:47:00Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sous\_Lesquels](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sous_lesquels/32/41843_2.png) [@Sous\_Lesquels](https://discuss.elastic.co/u/Sous_Lesquels)\
**Post date:** [July 15, 2019, 8:47pm UTC](https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612/1 "2019-07-15T20:47:00Z")

</div>

Assuming I have a file like this:

```
start k1 v11
other lines
end k1 i12
other lines
other lines
start k1 v21
anything else
end k1 i22
anything else
anything else
anything else
start k1 v31
anything else
anything else
end k1 i32

```

I'd like to get events like:

```
k1 v11 i12
k1 v21 i22
k1 v31 i32

```

i.e. join `start` / `end` pairs of lines and extract `kA` and `vBC` from `start` and `iDE` from `end`.

With this config:

```
filter {
  grok {
    match => { 'message' => '^start (?<s>\w+) +(?<v>.*)' }
    add_field => { 'type' => 'v' }
  }

  if (! [type]) {
    grok {
      match => { 'message' => '^end (?<s>\w+) +(?<i>.*)' }
      add_field => { 'type' => 'i' }
    }
  }

  if (! [type]) {
    drop {}
  }

  if ([type] == 'v') {
    aggregate {
      code => "map['v'] = event.get('v')"
      map_action => 'create'
      task_id => '%{s}'
    }

    drop {}
  }

  if ([type] == 'i') {
    aggregate {
      code => "event.set('v', map['v'])"
      end_of_task => true
      map_action => 'update'
      task_id => '%{s}'
      timeout => 60
    }
  }

}

```

I'm getting this:

```
k1 v11 i12
k1 %{v} i22
k1 %{v} i32

```

I assume because keys are the same, all but the first event is getting its value dropped. Any way to fix?

I wanted to avoid use `multiline` codec here because the spacing between events can be huge (i.e. the number of `other lines` and `anything else` can be huge) and also because there can be things embedded into other lines that make the regexp to filter them out harder to write.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 15, 2019, 9:18pm UTC](https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612/2 "2019-07-15T21:18:10Z")

</div>

> [@Sous\_Lesquels](#):
>
> I assume because keys are the same, all but the first event is getting its value dropped. Any way to fix?

Add '--pipeline.batch.size 1' to the command line (or adjust it in pipelines.yml for the specific pipeline or ...)

What is happening is that the first 125 lines of the file go through the first aggregate filter before any lines hit the second aggregate. When the first type i is processed by the second aggregate it deletes the map entry for that task\_id, so the when the next two lines go through the filter has no map for the task\_id and the filter is a no-op.

Note that with the java execution engine enabled [logstash will re-order lines](https://discuss.elastic.co/t/elpased-filter-works-differently-6-8-1-vs-7-1-possible-bug/187400/4) even with a single worker thread, so a solution like this is going to be fragile.

---

<div class="post-metadata">

**Author:** ![Sous\_Lesquels](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sous_lesquels/32/41843_2.png) [@Sous\_Lesquels](https://discuss.elastic.co/u/Sous_Lesquels)\
**Post date:** [July 15, 2019, 9:23pm UTC](https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612/3 "2019-07-15T21:23:19Z")

</div>

Ah, that fixed it - thanks Badger!

Are there any performance concerns that I should be aware of with having a batch size of 1 compared to the default 125?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 15, 2019, 9:24pm UTC](https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612/4 "2019-07-15T21:24:05Z")

</div>

Batching is done for performance reasons, but I do not know how much overhead it adds to use a batch size of 1.

---

<div class="post-metadata">

**Author:** ![Sous\_Lesquels](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sous_lesquels/32/41843_2.png) [@Sous\_Lesquels](https://discuss.elastic.co/u/Sous_Lesquels)\
**Post date:** [July 15, 2019, 9:27pm UTC](https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612/5 "2019-07-15T21:27:16Z")

</div>

OK thanks Badger!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 12, 2019, 9:27pm UTC](https://discuss.elastic.co/t/logstash-aggregate-plugin-with-the-same-key/190612/6 "2019-08-12T21:27:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
