# XML plugin parse

**URL:** <https://discuss.elastic.co/t/xml-plugin-parse/304124>\
**Category:** Logstash\
**Created:** [May 6, 2022, 12:28pm UTC](https://discuss.elastic.co/t/xml-plugin-parse/304124 "2022-05-06T12:28:05Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![adrianfusco](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianfusco/32/111798_2.png) [@adrianfusco](https://discuss.elastic.co/u/adrianfusco)\
**Post date:** [May 6, 2022, 12:28pm UTC](https://discuss.elastic.co/t/xml-plugin-parse/304124/1 "2022-05-06T12:28:05Z")

</div>

Hello,

I've been using the XML filtering plugin because I need to parse some XML data.

This is a simple example:

```auto
<task code="a01" status="wip"/>
<task code="a02" status="nwg"/>
<task code="a03" status="nwg">
     Description Line 1
     Description Line 2
     Description Line 3
     Description Line 4
</task>
<task code="a04" status="wip">
     <comment author="afusco">
         I've finished this part.
     </comment>
</task>

```

I'm trying to extract the code of these tasks

```auto
filter {
  xml {
    source => "message"
    store_xml => false
     xpath => [
       "/task/@code", "task_code",
       "/task/@status", "task_status",
     ]
  }
}

```

The thing is:

It's filtering correctly the lines that contains just a `<task>` tag. I can see the output is correct. But when it's processing the rest of the lines, for example, `comment` tags, it's parsing wrong.

To avoid it, I added the following simple condition just to drop the lines aren't `task` tags:

```auto
  if [message] !~ /^<task/ {
    drop { }
  }

```

But this is a workaround.

- Exists any way to just parse the desired specific tags and at this way Elasticsearch doesn't receive also the undesired data? It could be good to drop the data if it's not in the `xpath` array.

Example of the output:

When it's a `task` tag:

```auto
{
    "path" => "/var/log/xml2.log",
    "@timestamp" => 2022-05-05T16:22:24.602Z,
    "@version" => "1",
    "task_code" => [
        [0] "a01"
    ],
          "task_status" => [
        [0] "wip"
    ],

```

When it's not:

```auto
{
      "@version" => "1",
    "@timestamp" => 2022-05-06T12:27:01.559Z,
          "host" => "elastic",
       "message" => "\t<comment author="afusco">",
          "path" => "/var/log/xml2.log"
}

```

Thanks.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 6, 2022, 4:41pm UTC](https://discuss.elastic.co/t/xml-plugin-parse/304124/2 "2022-05-06T16:41:09Z")

</div>

Use a [multiline codec](https://www.elastic.co/guide/en/logstash/current/plugins-codecs-multiline.html) on the input to consume an entire XML document as a single event.

Possibly

```
  codec => multiline {
      pattern => "<task"
      negate => "true"
      what => "previous"
      auto_flush_interval => 10
  }
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 3, 2022, 4:41pm UTC](https://discuss.elastic.co/t/xml-plugin-parse/304124/3 "2022-06-03T16:41:34Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
