# Parse XML log lines

**URL:** <https://discuss.elastic.co/t/parse-xml-log-lines/137276>\
**Category:** Logstash\
**Created:** [June 25, 2018, 2:28pm UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276 "2018-06-25T14:28:27Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Boukhdhira](https://avatars.discourse-cdn.com/v4/letter/b/50afbb/32.png) [@Boukhdhira](https://discuss.elastic.co/u/Boukhdhira)\
**Post date:** [June 25, 2018, 2:28pm UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/1 "2018-06-25T14:28:27Z")

</div>

I would like some recommendation on how to parse a given xml document splited into log lines with logstash.  
My document looks like this:

```
|2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<root xmlns="http://xxxxxxxxxxx/5.0"> |
|---|---|
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<file-version>2.2</file-version> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<generation-date>2018-06-19T07:27:48.900+02:00</generation-date> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<report> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<status> |
| 2018-06-19T07:07:24+02:00 127.0.0.1 ping - |<info id="PushOK" type="number">44</info> |
| 2018-06-19T07:07:24+02:00 127.0.0.1 ping - |<info id="PushFailure" type="number">0</info> |
| 2018-06-19T07:07:24+02:00 127.0.0.1 ping - |</status> |
| 2018-06-19T07:07:24+02:00 127.0.0.1 ping - |<task exec="2018-06-19T06:05:00.000+02:00" id="XXXXXXXXX_06_2018"> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<transaction id="1" type="dfzlmsi" start="2018-06-19T06:27:00.000+02:00" stop="2018-06-19T06:27:00.000+02:00" retry="0" status="failed" reason="lost device"/> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |</target> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<target id="xxxxx8514" type="X1"> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |<transaction id="1" type="dfzlmsi" start="2018-06-19T06:27:00.000+02:00" stop="2018-06-19T06:27:00.000+02:00" retry="0" status="failed" reason="lost device"/> |
| 2018-06-19T07:29:09+02:00 127.0.0.1 ping - |</target> |
| 2018-06-19T07:07:32+02:00 127.0.0.1 ping - |<taskStatus ko="290" ok="24" status="partially_failed"/> |
| 2018-06-19T07:07:32+02:00 127.0.0.1 ping - |</task> |
| 2018-06-19T07:07:32+02:00 127.0.0.1 ping - |</report> |
| 2018-06-19T07:07:32+02:00 127.0.0.1 ping - |</root>|

```

And my goal is to agregate xml report and extract data from that.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [June 25, 2018, 2:49pm UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/2 "2018-06-25T14:49:03Z")

</div>

If you format your log example as preformatted text using markdown notation or the `</>` toolbar button we'll actually be able to see what it looks like.

---

<div class="post-metadata">

**Author:** ![Boukhdhira](https://avatars.discourse-cdn.com/v4/letter/b/50afbb/32.png) [@Boukhdhira](https://discuss.elastic.co/u/Boukhdhira)\
**Post date:** [June 26, 2018, 9:07am UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/3 "2018-06-26T09:07:25Z")

</div>

Thank you for your answer. I update my post.  
My question: is it possible to apply aggregate filter to build a valid xml output and prosses it using xml filter. if it's could you please give me some exemple.

thank you in advance.  
PS: | is a tabulation \t.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [June 26, 2018, 9:20am UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/4 "2018-06-26T09:20:57Z")

</div>

I don't know if an aggregate filter would be the best option here. I've never used it. I'd probably use a multiline codec to join all log entries into a single event and then use a ruby filter to chop it up and remove the non-XML data from everything but the first line.

How do you know that a sequence of log entries like this one won't ever interlace with each other if they're logged at the same time?

---

<div class="post-metadata">

**Author:** ![Boukhdhira](https://avatars.discourse-cdn.com/v4/letter/b/50afbb/32.png) [@Boukhdhira](https://discuss.elastic.co/u/Boukhdhira)\
**Post date:** [June 26, 2018, 9:51am UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/5 "2018-06-26T09:51:03Z")

</div>

i launch logstash using one thread worker  
`

> logstash --pipeline.workers 1 -f logstash.conf

`

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [June 26, 2018, 10:58am UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/6 "2018-06-26T10:58:22Z")

</div>

I meant on the logging end. Are all lines for a given XML document guaranteed to be logged atomically with no chance of any other messages slipping inbetween? If yes, why are the timestamps different in the example above?

---

<div class="post-metadata">

**Author:** ![Boukhdhira](https://avatars.discourse-cdn.com/v4/letter/b/50afbb/32.png) [@Boukhdhira](https://discuss.elastic.co/u/Boukhdhira)\
**Post date:** [June 26, 2018, 11:23am UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/7 "2018-06-26T11:23:44Z")

</div>

there is no such risk

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 24, 2018, 11:24am UTC](https://discuss.elastic.co/t/parse-xml-log-lines/137276/8 "2018-07-24T11:24:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
