# Recommendations for ingesting XML stream

**URL:** <https://discuss.elastic.co/t/recommendations-for-ingesting-xml-stream/198807>\
**Category:** Logstash\
**Created:** [September 10, 2019, 1:21am UTC](https://discuss.elastic.co/t/recommendations-for-ingesting-xml-stream/198807 "2019-09-10T01:21:25Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![dpr](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@dpr](https://discuss.elastic.co/u/dpr)\
**Post date:** [September 10, 2019, 1:21am UTC](https://discuss.elastic.co/t/recommendations-for-ingesting-xml-stream/198807/1 "2019-09-10T01:21:25Z")

</div>

Hey guys,

I'd like to ingest events that are continually streaming on a TCP port.

The stream format is XML, although it's not one event per line: the opening tag and closing tag are often on different lines. Here's a snip:

```
<Current>
  <Device>0x0011ee00001ee92a</Device>
  <TimeStamp>0x2509ad93</TimeStamp>
  <E1>0x00000000050e0ebb</E1>
  <M1>0x00000001</M1>
  <S1>Y</S1>
</Current>

```

Any recommendations on convenient ways to receive this data? If it requires some transformation, what type of munging would you recommend?

Thanks!

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [September 10, 2019, 12:46pm UTC](https://discuss.elastic.co/t/recommendations-for-ingesting-xml-stream/198807/2 "2019-09-10T12:46:00Z")

</div>

Try using a multiline codec

```
 codec => multiline { pattern => '^</Current>' negate => true what => "next" auto_flush_interval => 1 }

```

That will combine lines that do not contain \</Current\> into the following line that does contain that.

---

<div class="post-metadata">

**Author:** ![dpr](https://avatars.discourse-cdn.com/v4/letter/d/8797f3/32.png) [@dpr](https://discuss.elastic.co/u/dpr)\
**Post date:** [September 10, 2019, 6:11pm UTC](https://discuss.elastic.co/t/recommendations-for-ingesting-xml-stream/198807/3 "2019-09-10T18:11:20Z")

</div>

Wow, thanks for the help! This is working very nicely!

Here's what I wound up with:

```
input {
  tcp {
    mode => "client"
    host => "my-host"
    port => "12345"
    codec => multiline {
      pattern => "<Operation1>|<Operation2>|<Operation3>|<Operation4>|<Operation5>"
      negate => "true"
      what => "previous"
    }
    type => "resource-usage"
    tags => ["xml"]
  }
}
filter {
  if "xml" in [tags] {
    xml {
      source => "message"
      target => "xml_content"
    }
  }
}

```

**Followup question:** I'd like my document to contain a field named "Operation" that indicates which XML element this stanza started with (e.g. "Operation1", "Operation2", etc as seen in the multiline pattern tag). Is there a handy way to do that in the XML filter?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [September 10, 2019, 6:51pm UTC](https://discuss.elastic.co/t/recommendations-for-ingesting-xml-stream/198807/4 "2019-09-10T18:51:12Z")

</div>

> [@dpr](#):
>
> Is there a handy way to do that in the XML filter?

Not that I know of, but you could dissect it

```
dissect { mapping => { "message" => "<%{Operation}>%{}" } }

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 8, 2019, 6:51pm UTC](https://discuss.elastic.co/t/recommendations-for-ingesting-xml-stream/198807/5 "2019-10-08T18:51:18Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
