# Logstash xml parsing problems

**URL:** <https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115>\
**Category:** Logstash\
**Created:** [April 16, 2018, 7:37am UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115 "2018-04-16T07:37:49Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![povisx1](https://avatars.discourse-cdn.com/v4/letter/p/57b2e6/32.png) [@povisx1](https://discuss.elastic.co/u/povisx1)\
**Post date:** [April 16, 2018, 7:37am UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/1 "2018-04-16T07:37:50Z")

</div>

Hello,

I have the problem with xml parsing, my parsed result looks like :

`{"newMSISDN":"%{parsedMSISDN}","offset":52307874,"host":"miram.int.bite.lt","prospector":{"type":"log"},"@version":"1","@timestamp":"2018-04-16T06:03:11.780Z","beat":{"hostname":"miram.int.bite.lt","name":"miram.int.bite.lt","version":"6.2.3"},"source":"/ocs/mobicents/server/bite/log/ocs.log","message":"2018-04-16 08:01:18,089 INFO [net.bitegroup.ocs.lt.sbb.OcsSbb] Request CCR: <?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?><ccr><msisdn>37068783222</msisdn><ocsIp>10.241.53.45</ocsIp><apn>internplt</apn><sessionId>c0-10-225-64-26-epg02;1516844161;24238957</sessionId><ccrMscc><inputOctetsUsed>0</inputOctetsUsed><outputOctetsUsed>0</outputOctetsUsed><totalOctetsUsed>0</totalOctetsUsed><timeUsed>5013</timeUsed><inputOctetsUsedAfterTct>0</inputOctetsUsedAfterTct><outputOctetsUsedAfterTct>0</outputOctetsUsedAfterTct><totalOctetsUsedAfterTct>0</totalOctetsUsedAfterTct><timeUsedAfterTct>0</timeUsedAfterTct><ratingGroup>2085</ratingGroup><inputOctetsRequested>0</inputOctetsRequested><outputOctetsRequested>0</outputOctetsRequested><totalOctetsRequested>0</totalOctetsRequested><timeRequested>0</timeRequested><reportingReason>0</reportingReason><tariffChangeUsed>false</tariffChangeUsed></ccrMscc><sgsnIp>213.226.158.160</sgsnIp><ggsnIp>213.226.158.154</ggsnIp><imsi>246021005289080</imsi><imei>5359126095156720</imei><userLocationInfo>0142f62029053054</userLocationInfo><sgsnMccMnc>24602</sgsnMccMnc><chargingId>4116118880</chargingId><ip>10.20.170.183</ip><qos></qos><chargingCharacteristic>0400</chargingCharacteristic><chargingRuleName></chargingRuleName><requestType>UPDATE_REQUEST</requestType><requestNumber>453</requestNumber><creditControlFailureHandlingType>0</creditControlFailureHandlingType><ccSessionFailover>0</ccSessionFailover></ccr>","tags":["beats_input_codec_plain_applied"]}`

My code:

```
input {
  beats {
    port => 5043
  }
}

filter {
 
    xml {
      source => "message"
      remove_namespaces => "true"
      xpath => ["//ccrB/msisdn/text()", "parsedMSISDN",
                "//ccrB/ocsIp/text()", "ocsIP"
               ]
      store_xml => "false"
    }
    mutate {
      add_field =>{"newMSISDN" => "%{parsedMSISDN}"}
    }
}

output {

    file {
      path => "/home/work/logstash-out/ocs_test.out"
    }

    #stdout { codec => rubydebug }

}

```

i want to add field newMSISDN but I don't get xml value, as i understand problem is in this place

\<?xml version=\"1.0\" encoding=\"UTF-8\" standalone=\"yes\"?\>, maybe xpath don't understand these symbols \*\*\"\*\* 

Maybe someone can help me with this one , thanks 🙂

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 16, 2018, 7:46am UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/2 "2018-04-16T07:46:44Z")

</div>

Your message field starts like this:

`2018-04-16 08:01:18,089 INFO [net.bitegroup.ocs.lt.sbb.OcsSbb] Request CCR: <?xml version=\"1.0\"...`

As this is not all valid XML, you can not directly apply the XML filter to it. You will need to use a [grok](https://www.elastic.co/guide/en/logstash/current/plugins-filters-grok.html) or [dissect](https://www.elastic.co/guide/en/logstash/current/plugins-filters-dissect.html) filter to parse the components of the log so you get the XML content at the end in a separate field. Then you can use the XML filter on this field.

---

<div class="post-metadata">

**Author:** ![povisx1](https://avatars.discourse-cdn.com/v4/letter/p/57b2e6/32.png) [@povisx1](https://discuss.elastic.co/u/povisx1)\
**Post date:** [April 16, 2018, 8:22am UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/3 "2018-04-16T08:22:44Z")

</div>

Thanks for the answer 🙂

Maybe you can say more which grok filter I should use, I try to use {GREEDYDATA:msgBody}  
but this only take all message, how can I take only xml log block ? 🙂 Main problem is that logstash add these symbols, maybe I can somehow take them off ? and then as I understand also should work .

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 16, 2018, 8:37am UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/4 "2018-04-16T08:37:33Z")

</div>

Have a look at the dissect filter as it might be easier to get started with and should be sufficient here. [This blog post](https://www.elastic.co/blog/logstash-dude-wheres-my-chainsaw-i-need-to-dissect-my-logs) contains a good introduction.

---

<div class="post-metadata">

**Author:** ![povisx1](https://avatars.discourse-cdn.com/v4/letter/p/57b2e6/32.png) [@povisx1](https://discuss.elastic.co/u/povisx1)\
**Post date:** [April 18, 2018, 1:35pm UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/5 "2018-04-18T13:35:26Z")

</div>

Sorry, I don't have much time, so only now I checked your answer and I think this is not good. Is there really no other way to say for logstash to not add these symbols in this place ( \ ) :

\<?xml version=\*\*\*\*"1.0\*\*\*\*" encoding=\*\*\*\*"UTF-8\*\*\*\*" standalone=\*\*\*\*"yes\*\*\*\*"?\> 

Because in my real log there not exist these symbols , for example in my real log this place looks like this and the xpath work then:

 \<?xml version="1.0" encoding="UTF-8" standalone="yes"?\>

So is there really no way to config logstash to not add these symbols ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [April 18, 2018, 2:27pm UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/6 "2018-04-18T14:27:37Z")

</div>

> [@povisx1](#):
>
> I think this is not good

Did you try it out? If so, what did not work?

---

<div class="post-metadata">

**Author:** ![povisx1](https://avatars.discourse-cdn.com/v4/letter/p/57b2e6/32.png) [@povisx1](https://discuss.elastic.co/u/povisx1)\
**Post date:** [April 20, 2018, 12:43pm UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/7 "2018-04-20T12:43:36Z")

</div>

Hi!

I have try it and this I think for me is not correct because my log all the time is differend sometimes I don't even have xml in my log. So I now try to take value with grok regex  
So now I only have this problem, maybe you would help me, I take value like this:

![image](https://us1.discourse-cdn.com/elastic/original/3X/6/6/6654a653cba881fb053599de9480b34ca8d3905a.png)

but my result looks like this : ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/3/b36eebd9554d3479c09a584a05c196cdd40147d3.png)  
How i can get result like this : 37068783222

🙂

---

<div class="post-metadata">

**Author:** ![Jenni](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jenni/32/29684_2.png) [@Jenni](https://discuss.elastic.co/u/Jenni)\
**Post date:** [April 20, 2018, 1:04pm UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/8 "2018-04-20T13:04:09Z")

</div>

If you only want the middle part extracted, you have to put the named group only around the middle part. (Posting a screenshot instead of text doesn't make helping you easier 😉 )

`<msisdn>(?<msisdn>.*)<\/msisdn>`

(I don't know what the additional :? and ? were supposed to do.)

---

<div class="post-metadata">

**Author:** ![povisx1](https://avatars.discourse-cdn.com/v4/letter/p/57b2e6/32.png) [@povisx1](https://discuss.elastic.co/u/povisx1)\
**Post date:** [May 16, 2018, 6:13am UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/9 "2018-05-16T06:13:06Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> dissect

Hi thanks all for your help I taked all xml with grok regex like this:  
match =\> { "message" =\> "(?\<(.\*)\</cc.\>)"}

And then use xml plugin and take all values I need 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 13, 2018, 6:18am UTC](https://discuss.elastic.co/t/logstash-xml-parsing-problems/128115/10 "2018-06-13T06:18:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
