# Trying to parse xml with \\r in the xml message field

**URL:** <https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406>\
**Category:** Logstash\
**Created:** [November 13, 2018, 8:09am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406 "2018-11-13T08:09:58Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![lukas.bayard](https://avatars.discourse-cdn.com/v4/letter/l/a4c791/32.png) [@lukas.bayard](https://discuss.elastic.co/u/lukas.bayard)\
**Post date:** [November 13, 2018, 8:09am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/1 "2018-11-13T08:09:58Z")

</div>

I try to read an XML file with Logstash. But the XML is only read until the first \r. Shouldn't everything be read with Multiline? Or how can I exclude that message field?

> **XML File**
>
> ```
> <?xml version="1.0" encoding="utf-8"?><update xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xsd="http://www.w3.org/2001/XMLSchema"><protocol><year>2018</year><number>S18085936</number><assignment>D18051009</assignment><pager>106</pager><timestamp>2018-06-18T00:21:54+02:00</timestamp></protocol><keyword>1</keyword><message>Lorem ipsum dolor sit amet,  
> consetetur sadipscing elitr,
> sed diam nonumy eirmod tempor invidunt ut labore et dolore magna aliquyam erat,
> sed diam voluptua.</message><origin>HR K</origin><status><type>5</type><timestamp>2018-06-18T01:08:58+02:00</timestamp></status><object><type>Location</type><address><street>street</street><streetnumber>101</streetnumber><zip>8888</zip><city>City</city></address><name>street 10</name><coords><lat>47.00000</lat><lon>8.00000</lon></coords></object><object><type>Destination</type><name>HAUPTGEBÄUDE</name></object></update>
> 
> ```

> **My configuration looks the following:**
>
> ```
> input
> {
> file
> {
> path => "C:/Temp/SRZ/Probleme180618/test/*.xml"
> sincedb_path => "nul"
> start_position => "beginning"
> type => "xml"
> codec => multiline {
> pattern => "" 
> negate => "true"
> what => "previous"
> }
> }
> }
> filter
> {
> xml
> {
> source => "message"
> store_xml => false
> target => "protocol"
> force_array => false
> xpath => [
> "//protocol/number/text()", "protocol_number",
> "//protocol/assignment/text()", "protocol_assignment",
> "//protocol/timestamp/text()", "protocol_timestamp",
> "//protocol/calltime/text()", "protocol_calltime",
> "//status/type/text()", "status_type"
> ]
> }
> 
> }
> output
> {
> file {
> path => "C:/Temp/SRZ/output.txt"
> codec => line { format => "Nr: %{protocol_number} Assignment: %{protocol_assignment} Timestamp: %{protocol_timestamp} CallTime: %{protocol_calltime} StatusType: %{status_type}"}
> }
> stdout
> {
> codec => rubydebug
> }
> } 
> 
> ```

> **Logstash Output**
>
> ```
> {
> "@version" => "1",
> "path" => "C:/Temp/SRZ/Probleme180618/test/20180618010858_739_[SendUpdate]__T_200.xml",
> "host" => "acec-lub01",
> "type" => "xml",
> "message" => "consetetur sadipscing elitr,\r",
> "@timestamp" => 2018-11-13T08:03:37.300Z
> }
> {
> "@version" => "1",
> "protocol_timestamp" => [
> [0] "2018-06-18T00:21:54+02:00"
> ],
> "path" => "C:/Temp/SRZ/Probleme180618/test/20180618010858_739_[SendUpdate]__T_200.xml",
> "host" => "acec-lub01",
> "type" => "xml",
> "protocol_number" => [
> [0] "S18085936"
> ],
> "message" => "<?xml version=\"1.0\" encoding=\"utf-8\"?><update xmlns:xsi=\"http://www.w3.org/2001/XMLSchema-instance\" xmlns:xsd=\"http://www.w3.org/2001/XMLSchema\"><protocol><year>2018</year><number>S18085936</number><assignment>D18051009</assignment><pager>106</pager><timestamp>2018-06-18T00:21:54+02:00</timestamp></protocol><keyword>1</keyword><message>Lorem ipsum dolor sit amet, \r",
> "@timestamp" => 2018-11-13T08:03:37.253Z,
> "protocol_assignment" => [
> [0] "D18051009"
> ]
> }
> 
> ```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 13, 2018, 8:13am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/2 "2018-11-13T08:13:40Z")

</div>

That does not look like valid XML, so I am not surprised the XML filter does not work. I would recommend you either correct the input data or parse the data as text.

---

<div class="post-metadata">

**Author:** ![lukas.bayard](https://avatars.discourse-cdn.com/v4/letter/l/a4c791/32.png) [@lukas.bayard](https://discuss.elastic.co/u/lukas.bayard)\
**Post date:** [November 13, 2018, 8:15am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/3 "2018-11-13T08:15:10Z")

</div>

Wouldn't it be possible to correct the XML in the logstash before doing filtering?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 13, 2018, 10:07am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/4 "2018-11-13T10:07:23Z")

</div>

I see that you updated the data and that it now looks like valid XML. Does the file contain a single XML document spread over multiple lines or can it contain more than one?

---

<div class="post-metadata">

**Author:** ![lukas.bayard](https://avatars.discourse-cdn.com/v4/letter/l/a4c791/32.png) [@lukas.bayard](https://discuss.elastic.co/u/lukas.bayard)\
**Post date:** [November 13, 2018, 10:22am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/5 "2018-11-13T10:22:07Z")

</div>

Files are always looking the same as you see in the "XML File" example, so yes, this is a single xml document with multiple lines. As you can see at there are CR inside the XML Tag.

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [November 13, 2018, 8:40pm UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/6 "2018-11-13T20:40:28Z")

</div>

Try changing your multiline pattern to `<\?xml.*` or `<?xml` (Don't remember if it takes regular expression or not). Afterwards, before your XML filter, use the mutate filter's gsub function to remove the \r carriage return.

---

<div class="post-metadata">

**Author:** ![lukas.bayard](https://avatars.discourse-cdn.com/v4/letter/l/a4c791/32.png) [@lukas.bayard](https://discuss.elastic.co/u/lukas.bayard)\
**Post date:** [November 14, 2018, 10:25am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/7 "2018-11-14T10:25:07Z")

</div>

Thank you for your response. I have tried that with the following configuration:

> **Logstash configuration**
>
> ```
> input
> {
> file
> {
> path => "C:/Temp/SRZ/Probleme180618/test/*.xml"
> sincedb_path => "nul"
> start_position => "beginning"
> type => "xml"
> codec => multiline {
> pattern => "<\?xml.*" 
> negate => "false"
> what => "previous"
> }
> }
> }
> filter{
> mutate { 
> gsub => [ 
> "message", "[\r]", "",
> "message", "[\n]", ""
> ] 
> }
> }
> filter
> {
> xml
> {
> source => "message"
> store_xml => false
> target => "protocol"
> force_array => false
> xpath => [
> "//protocol/number/text()", "protocol_number",
> "//protocol/assignment/text()", "protocol_assignment",
> "//protocol/timestamp/text()", "protocol_timestamp",
> "//protocol/calltime/text()", "protocol_calltime",
> "//status/type/text()", "status_type"
> ]
> }
> 
> }
> output
> {
> file {
> path => "C:/Temp/SRZ/output.txt"
> codec => line { format => "custom format: %{message}" }
> }
> stdout
> {
> codec => rubydebug
> }
> } 
> 
> ```

I got the following output in the console:

> [2018-11-14T11:18:11,386][INFO][logstash.outputs.file] Opening file {:path=\>"C:/Temp/SRZ/output.txt"}  
> {  
> "type" =\> "xml",  
> "protocol\_timestamp" =\> [  
> [0] "2018-06-18T00:21:54+02:00"  
> ],  
> "@version" =\> "1",  
> "path" =\> "C:/Temp/SRZ/Probleme180618/test/20180618010858\_739\_[SendUpdate]\__T\_200.xml",  
> "@timestamp" =\> 2018-11-14T10:18:10.672Z,  
> "host" =\> "acec-lub01",  
> "protocol\_number" =\> [  
> [0] "S18085936"  
> ],  
> "protocol\_assignment" =\> [  
> [0] "D18051009"  
> ],  
> "message" =\> "\<?xml version=\"1.0\" encoding=\"utf-8\"?\>\<update xmlns:xsi="[http://www.w3.org/2001/XMLSchema-instance\](http://www.w3.org/2001/XMLSchema-instance%5C)" xmlns:xsd="[http://www.w3.org/2001/XMLSchema\](http://www.w3.org/2001/XMLSchema%5C)"\>2018S18085936D180510091062018-06-18T00:21:54+02:001Lorem ipsum dolor sit amet, "  
> }  
> {  
> "type" =\> "xml",  
> "@version" =\> "1",  
> "path" =\> "C:/Temp/SRZ/Probleme180618/test/20180618010858\_739_[SendUpdate]\_\_T\_200.xml",  
> "@timestamp" =\> 2018-11-14T10:18:10.714Z,  
> "host" =\> "acec-lub01",  
> "message" =\> "consetetur sadipscing elitr,"  
> }

and the following output in the file:

> custom format: \<?xml version="1.0" encoding="utf-8"?\>2018S18085936D180510091062018-06-18T00:21:54+02:001Lorem ipsum dolor sit amet,  
> custom format: consetetur sadipscing elitr,

So looks that the mutate is not working/configured as expected. Any hints/advice?

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [November 16, 2018, 1:28am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/8 "2018-11-16T01:28:16Z")

</div>

I'm thinking you don't want your gsub pattern in brackets.

```
filter{
	mutate { 
		gsub => [ 
			"message", "\r", "",
			"message", "\n", ""
		] 
	}
}

```

Also, once you get it working, you can combine them onto a single line as `"message", "\r|\n", ""`

---

<div class="post-metadata">

**Author:** ![lukas.bayard](https://avatars.discourse-cdn.com/v4/letter/l/a4c791/32.png) [@lukas.bayard](https://discuss.elastic.co/u/lukas.bayard)\
**Post date:** [November 16, 2018, 7:44am UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/9 "2018-11-16T07:44:04Z")

</div>

OK thx. I have also changed the "what =\> "previous" to next and now I get the 2nd line also in the output, but not the other lines:

`<?xml version="1.0" encoding="utf-8"?><update xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xsd="http://www.w3.org/2001/XMLSchema"><protocol><year>2018</year><number>S18085936</number><assignment>D18051009</assignment><pager>106</pager><timestamp>2018-06-18T00:21:54+02:00</timestamp></protocol><keyword>1</keyword><message>Lorem ipsum dolor sit amet, consetetur sadipscing elitr,`

the other 2 lines are missing:  
`sed diam nonumy eirmod tempor invidunt ut labore et dolore magna aliquyam erat,`  
`sed diam voluptua.</message>.....`

---

<div class="post-metadata">

**Author:** ![lukas.bayard](https://avatars.discourse-cdn.com/v4/letter/l/a4c791/32.png) [@lukas.bayard](https://discuss.elastic.co/u/lukas.bayard)\
**Post date:** [November 27, 2018, 4:47pm UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/10 "2018-11-27T16:47:44Z")

</div>

Does anyone have any advice?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 25, 2018, 4:47pm UTC](https://discuss.elastic.co/t/trying-to-parse-xml-with-r-in-the-xml-message-field/156406/11 "2018-12-25T16:47:46Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
