# Parse Elasticsearch json logs in filebeat

**URL:** <https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533>\
**Category:** Beats\
**Tags:** docker, journalbeat, filebeat\
**Created:** [January 5, 2023, 10:48am UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533 "2023-01-05T10:48:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![qwinkler](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/qwinkler/32/77367_2.png) [@qwinkler](https://discuss.elastic.co/u/qwinkler)\
**Post date:** [January 5, 2023, 10:48am UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/1 "2023-01-05T10:48:41Z")

</div>

Hello. I want to properly collect Elasticsearch logs. I have the following architecture.  
On the Linux node, I have Docker installed. I configured the Journald logging driver using official documentation ([Journald logging driver | Docker Documentation](https://docs.docker.com/config/containers/logging/journald/)). On this node, I'm running Elasticsearch in Docker (version 7.17.3). Since I have Journald logging, the container sends its logs to journald instead of files. Also, I have a filebeat running in a container with journald input to grab these logs:

```auto
filebeat.inputs:
- type: journald

```

Later, these logs are ingested in Kafka:

```auto
output.kafka:
  enabled: true
  hosts: "<omitted>"
  topic: "<omitted>"
  # and other Kafka parameters

```

Later these logs will be parsed by Logstash:

```auto
input {
  kafka {
      bootstrap_servers => "<omitted>"
      topics_pattern => "<omitted>"
      codec => "json"
      # and other Kafka parameters
    }
}

```

The final scheme: Docker container -\> journald -\> filebeat with journald input -\> Kafka -\> Logstash -\> Elasticsearch index.

So, I just want to properly parse logs from Elasticsearch. The problem is that Elasticsearch writes stacktrace as an array on separate lines.  
For instance, this message will be parsed and ingested in ES as a single document:

```auto
{"type":"server","timestamp":"2023-01-05T09:45:51,995Z","level":"INFO","component":"c.f.s.e.WatchRunner","cluster.name":"<omitted>","node.name":"<omitted>","message":"<omitted>","cluster.uuid":"<omitted>","node.id":"<omitted>"}

```

And this message will be parsed and ingested in ES as 4 non-json documents:

```auto
{"type": "server", "timestamp": "2023-01-05T09:45:51,996Z", "level": "WARN", "component": "c.f.s.i.InternalAuthTokenProvider", "cluster.name": "<ommited>", "node.name": "<ommited>", "message": "<ommited>", "cluster.uuid": "<ommited>", "node.id": "<ommited>" ,
"stacktrace": ["<omitted>",
"at org.elasticsearch.<omitted>",
"at org.elasticsearch.<omitted>"] }

```

So, filebeat with journald input reads it as separate messages.  
I just want filebeat to parse it and ingest it in Kafka as a **single** json message, so it will be ingested in ES later as a single document.

I tried to configure the multiline, but had no luck with it:

```auto
filebeat.inputs:
- type: journald
  id: everything
  seek: cursor
  paths: ["/var/log/journal"]
  multiline.type: pattern
  multiline.pattern: '(\"stacktrace\"|^\"(.*)\"(]| |})*)'
  multiline.negate: false
  multiline.match: after

```

The question is: how to properly parse multiline JSON logs in my case?

---

<div class="post-metadata">

**Author:** ![Ayush\_Mathur](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ayush_mathur/32/77134_2.png) [@Ayush\_Mathur](https://discuss.elastic.co/u/Ayush_Mathur)\
**Post date:** [January 6, 2023, 8:01am UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/2 "2023-01-06T08:01:36Z")

</div>

Try using

> decode\_json\_fields

processor in your journald input type. I have configured in my use case and the logs are coming in perfectly.

---

<div class="post-metadata">

**Author:** ![qwinkler](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/qwinkler/32/77367_2.png) [@qwinkler](https://discuss.elastic.co/u/qwinkler)\
**Post date:** [January 6, 2023, 1:54pm UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/3 "2023-01-06T13:54:55Z")

</div>

Thanks for the help.

I made a few more tests and it turned out that it's a bug in a filebeat with `journald` input. More details if you are interested are on GitHub: [[Filebeat] multiline doesn't work with journald input · Issue #34200 · elastic/beats · GitHub](https://github.com/elastic/beats/issues/34200)

---

<div class="post-metadata">

**Author:** ![Ayush\_Mathur](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ayush_mathur/32/77134_2.png) [@Ayush\_Mathur](https://discuss.elastic.co/u/Ayush_Mathur)\
**Post date:** [January 6, 2023, 1:58pm UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/4 "2023-01-06T13:58:55Z")

</div>

@qwinkler TBH, I didn't face any such issue in my case and the events were being processed and ingested properly. Can you please share how you configured `decode_json_fields` for journald input ?

---

<div class="post-metadata">

**Author:** ![qwinkler](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/qwinkler/32/77367_2.png) [@qwinkler](https://discuss.elastic.co/u/qwinkler)\
**Post date:** [January 6, 2023, 8:14pm UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/5 "2023-01-06T20:14:59Z")

</div>

> [@Ayush\_Mathur](#):
>
> Can you please share how you configured `decode_json_fields` for journald input ?

I don't need the `decode_json_fields` processor and there's why. As I understand, this processor ([Decode JSON fields | Filebeat Reference [8.11] | Elastic](https://www.elastic.co/guide/en/beats/filebeat/current/decode-json-fields.html)) is useful to extract json string from the **event**. For instance, you have an event:

```auto
{
  "hello": "world",
  "message": "{\"level\": \"INFO\",\"msg\": \"hello\"}"
}

```

And with the following configuration:

```auto
processors:
  - decode_json_fields:
      fields: ["message"]
      target: "log"

```

You'll get the following event in ES:

```auto
{
  "hello": "world",
  "message":"{\"level\": \"INFO\",\"msg\": \"hello\"}"},
  "log": {
    "level": "INFO",
    "msg": "hello"
  }
}

```

In my use case, the event is not "full", it's split by several messages. So, I don't need to decode JSON fields, I just need to "combine" several messages into one event.

Back to my scenario. I have 6 distinct messages in journal:

```auto
{"hello": "person"}
{"hello": "world",
"x": ["1",
"2",
"3"]}
{"hello": "2nd person"}

```

They will be sent to ES as 6 events (1 line = 1 event). So, I'll receive two JSON events (lines 1 and 6) and four string-like events (2,3,4,5), because it's an invalid JSON. I want to send only 3 events, like this:

```auto
{"hello": "person"}
{"hello": "world", "x": ["1","2","3"]}
{"hello": "2nd person"}

```

Therefore, I have to use multiline ([Manage multiline messages | Filebeat Reference [8.5] | Elastic](https://www.elastic.co/guide/en/beats/filebeat/8.5/multiline-examples.html)) to combine multiple messages into one event. Unfortunately, it doesn't work with `journald` (I'm using it), while it works with the `log` input.

---

<div class="post-metadata">

**Author:** ![TiagoQueiroz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tiagoqueiroz/32/107061_2.png) [@TiagoQueiroz](https://discuss.elastic.co/u/TiagoQueiroz)\
**Post date:** [January 13, 2023, 2:18pm UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/6 "2023-01-13T14:18:16Z")

</div>

Hi @qwinkler,

The journald input uses parsers in the same way as the filestream input, you can define them following the documentation [here](https://www.elastic.co/guide/en/beats/filebeat/current/filebeat-input-filestream.html#_parsers).

We even have a test ensuring the multiline parser works with journald input:

> <https://github.com/elastic/beats/blob/878f026103df1d4b8fa8a7152bc5bd14ff1278b8/filebeat/input/journald/input_parsers_test.go#L34-L62>

I also added this information in the GH issue.

---

<div class="post-metadata">

**Author:** ![TiagoQueiroz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tiagoqueiroz/32/107061_2.png) [@TiagoQueiroz](https://discuss.elastic.co/u/TiagoQueiroz)\
**Post date:** [January 13, 2023, 2:20pm UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/7 "2023-01-13T14:20:29Z")

</div>

There is also an example in the reference documentation: [beats/filebeat.reference.yml at main · elastic/beats · GitHub](https://github.com/elastic/beats/blob/main/filebeat/filebeat.reference.yml#L1183-L1186)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 10, 2023, 4:21pm UTC](https://discuss.elastic.co/t/parse-elasticsearch-json-logs-in-filebeat/322533/8 "2023-02-10T16:21:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
