# No idea how to setup pipeline for the xml parsing

**URL:** <https://discuss.elastic.co/t/no-idea-how-to-setup-pipeline-for-the-xml-parsing/370243>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-monitoring\
**Created:** [November 8, 2024, 6:42pm UTC](https://discuss.elastic.co/t/no-idea-how-to-setup-pipeline-for-the-xml-parsing/370243 "2024-11-08T18:42:23Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nevov21](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nevov21/32/137328_2.png) [@Nevov21](https://discuss.elastic.co/u/Nevov21)\
**Post date:** [November 8, 2024, 6:42pm UTC](https://discuss.elastic.co/t/no-idea-how-to-setup-pipeline-for-the-xml-parsing/370243/1 "2024-11-08T18:42:23Z")

</div>

Hello,

I know theres a lot of the same topics because i was trying to figure out from these how to parse logs from xml to the Elastic. What im trying to do is:

```auto
<record>
  <date>2024-11-08T18:32:56.379787054Z</date>
  <millis>1731090776379</millis>
  <nanos>787054</nanos>
  <sequence>102</sequence>
  <logger>org.forgerock.openidm.relationship.SignalPropagationCalculatorFactory</logger>
  <level>INFO</level>
  <class>org.forgerock.openidm.relationship.SignalPropagationCalculatorFactory</class>
  <method>getSignalPropagationCalculator</method>
  <thread>15</thread>
  <message>Smart-signaling disabled: false</message>
</record>

```

It should be in one log message and in separetly brackets in log. For example fieldxml.date, fieldxml.millis. I was trying to figure out with chatgpt but that even cant help me. I tried something like this:

```auto
filter {
  if "xml" in [log][file][path] {
    mutate {
      replace => { "message" => "<record>%{message}</record>" }
    }

    xml {
      source => "message"
      target => "parsed_xml"
      store_xml => true
      force_array => false
    }

    mutate {
      rename => { "[parsed_xml][date]" => "date" }
      rename => { "[parsed_xml][millis]" => "millis" }
      rename => { "[parsed_xml][nanos]" => "nanos" }
      rename => { "[parsed_xml][sequence]" => "sequence" }
      rename => { "[parsed_xml][logger]" => "logger" }
      rename => { "[parsed_xml][level]" => "level" }
      rename => { "[parsed_xml][class]" => "class" }
      rename => { "[parsed_xml][method]" => "method" }
      rename => { "[parsed_xml][thread]" => "thread" }
      rename => { "[parsed_xml][message]" => "log_message" }
    }

    mutate {
      remove_field => ["message", "parsed_xml"]
    }
  }
}

```

Thanks for helping! 🙂

---

<div class="post-metadata">

**Author:** ![rugenl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rugenl/32/12887_2.png) [@rugenl](https://discuss.elastic.co/u/rugenl)\
**Post date:** [November 8, 2024, 7:10pm UTC](https://discuss.elastic.co/t/no-idea-how-to-setup-pipeline-for-the-xml-parsing/370243/2 "2024-11-08T19:10:16Z")

</div>

Use multiline to merge the event.

---

<div class="post-metadata">

**Author:** ![Nevov21](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nevov21/32/137328_2.png) [@Nevov21](https://discuss.elastic.co/u/Nevov21)\
**Post date:** [November 9, 2024, 5:50pm UTC](https://discuss.elastic.co/t/no-idea-how-to-setup-pipeline-for-the-xml-parsing/370243/3 "2024-11-09T17:50:08Z")

</div>

Okey, but where i should put that multiline? In filebeat.yml or in mine pipeline on elastic?

---

<div class="post-metadata">

**Author:** ![gpineda\_dev](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gpineda_dev/32/137494_2.png) [@gpineda\_dev](https://discuss.elastic.co/u/gpineda_dev)\
**Post date:** [November 10, 2024, 12:19am UTC](https://discuss.elastic.co/t/no-idea-how-to-setup-pipeline-for-the-xml-parsing/370243/4 "2024-11-10T00:19:14Z")

</div>

The code you initialy posted would be the one for a logstash pipeline, not filebeat or even elasticsearch.

Usually the **data flow** is the following :

- filebeat -\> elasticsearch
- filebeat -\> logstash -\> elasticsearch
- logstash -\> elasticsearch

**Filebeat (beats) is a "standalone" binary** deployed on the host where you want to collect data while **logstash can indead collect logs but requires JVM** to run.

Once **collected, the events are sent to elasticsearch** and, if requested, **an ingest pipeline will be executed** on your event during ingestion resulting in a new document within your target index.

So in your case, since you mentioned filebeat, I assume you plan using filebeat to access the logs, then send it to logstash or directly elasticsearch.

Then to process XML log formated data with filebeat, you can indeed use [multiline](https://www.elastic.co/guide/en/beats/filebeat/current/multiline-examples.html) to **extract as message your complete xml entry** and with the [decode\_xml processor](https://www.elastic.co/guide/en/beats/filebeat/current/decode-xml.html) from filebeat or [logstash xml filter](https://www.elastic.co/guide/en/logstash/current/plugins-filters-xml.html), **parse your "message"** entry to an actual xml.

```yaml
filebeat.inputs:
- type: filestream
  id: my-filestream-id
  paths:
    - /opt/path/to/my/xml.log
  parsers:
    - multiline:
         type: pattern
         pattern: '^<record>'
         negate: true
         match: after

# xml conversion can be within the same filebeat.yaml or handed over to logstash.
# if in the same :
processors:
  - decode_xml:
      field: message
      target_field: "record"
      overwrite_keys: true

# any output (here logstash is not necessary)

```

[Here is a similar thread about it](https://discuss.elastic.co/t/filebeat-process-multilne-xml/77761)

_PS: This is my first post on the platform, so not sure if details are sufficient._

---

<div class="post-metadata">

**Author:** ![rugenl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rugenl/32/12887_2.png) [@rugenl](https://discuss.elastic.co/u/rugenl)\
**Post date:** [November 10, 2024, 2:29am UTC](https://discuss.elastic.co/t/no-idea-how-to-setup-pipeline-for-the-xml-parsing/370243/5 "2024-11-10T02:29:42Z")

</div>

I don’t have access to the config I implemented. I used a he new agent, that’s equivalent to filebeat. The config above looks like what I remember.
