# Logstash throws runtime exception while parsing huge XML

**URL:** <https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621>\
**Category:** Logstash\
**Created:** [May 7, 2020, 8:48pm UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621 "2020-05-07T20:48:28Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Saravana\_Maadavan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saravana_maadavan/32/60232_2.png) [@Saravana\_Maadavan](https://discuss.elastic.co/u/Saravana_Maadavan)\
**Post date:** [May 7, 2020, 8:48pm UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621/1 "2020-05-07T20:48:29Z")

</div>

Hello,

In Logstash I am trying to replace a prefix in the XML which is not having any definition and then indexing it to elastic. Receiving error while parsing some huge XMLs. I would be requiring the full XML to be stored in elastic for analytics. PFB error snippet & yml file for Logstash configuration.

**Error** -

`exception=>#<RuntimeError: entity expansion has grown too large>, :backtrace=>["uri:classloader:/META-INF/jruby.home/lib/ruby/stdlib/rexml/text.rb:399:in `block in unnormalize'", "org/jruby/RubyString.java:3056:in `gsub'", "uri:classloader:/META-INF/jruby.home/lib/ruby/stdlib/rexml/text.rb:396:in `unnormalize'"`

> ```
> input {
> beats {
> port => "61000"
> client_inactivity_timeout => 3600
> }
> }
> filter{
> mutate{
> gsub => [
> "message", "<L:", "<",
> "message", "</L:", "</",
> "message", "&lt;L:", "&lt;",
> "message", "&lt;/L:", "&lt;/"
> ]
> }
> xml{
> source => "message"
> store_xml => true
> xpath => ["//RECORD/Name/text()","capability","//RECORD/DATE/text()","message_date","//RECORD/TIME/text()","message_time"]
> target => "xml_message"
> }
> 
> grok {
> match => ["message_date", "(?<month>20.{5})"]
> }
> 
> mutate {
> lowercase => ["capability"]
> add_field => {
> "message_dateTime" => "%{message_date}T%{message_time}"
> }
> remove_field => ["message_date","message_time","host"]
> }
> }
> output {
> elasticsearch {
> hosts => ["XXXXXX"]
> index => "%{capability}_%{+YYYY_ww}"
> }
> }
> 
> ```

Kindly let me know if there is a possibility to increase any parameter value to make huge XMLs parse or any workaround for this please.

Cheers,  
Maadavan

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 7, 2020, 11:53pm UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621/2 "2020-05-07T23:53:22Z")

</div>

> [@Saravana\_Maadavan](#):
>
> Kindly let me know if there is a possibility to increase any parameter value to make huge XMLs parse or any workaround for this please.

The error is occurring in the [rexml](https://github.com/ruby/rexml/blob/be62163ba12a6657679a34e472b1d29d75e0e881/lib/rexml/text.rb#L397) library, which is used by the XmlSimple library that the logstash xml filter uses. The default limit on the size of an XML entity is 10 KB. Note that, as far as I can see, the problem is not the size of the XML document, it is the size of an entity within that document.

There is no way to pass rexml configuration options to the xml filter.

The [limit](https://github.com/ruby/rexml/blob/be62163ba12a6657679a34e472b1d29d75e0e881/lib/rexml/security.rb#L19) is a class variable. So let me say that it would be a terrible, terrible idea to use a ruby filter to set it before calling the xml filter. Do not do it.

If you do not actually need to store the entire document then if you use the xpath option the XML is parsed using nokogiri instead of XmlSimple. That may not have the same limits.

---

<div class="post-metadata">

**Author:** ![Saravana\_Maadavan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saravana_maadavan/32/60232_2.png) [@Saravana\_Maadavan](https://discuss.elastic.co/u/Saravana_Maadavan)\
**Post date:** [May 8, 2020, 4:56am UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621/3 "2020-05-08T04:56:16Z")

</div>

Hello @Badger,

Thanks for your swift reply. Few queries please.

1. rexml library is used in XML filter only because ruby filter (in this case mutate gsub) is used prior to XML filter?

2. If yes for the above query , I am using it to replace the invalid namespace prefix present in "message", is there a way to achieve the same without having ruby filter before XML filter?

3. Would not be able to provide XPath for nodes to parse since many API transaction logs are being pushed to elastic using logstash, so not feasible to provide all the tag names. Any other workaround please?

4. Would there be an end difference of storing message as xml instead of text during aggregation in Kibana? Because as a text I am able to index all these XMLs but since it is huge not able to process aggregation on those messages fields.

Cheers,  
Maadavan

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 8, 2020, 12:45pm UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621/4 "2020-05-08T12:45:09Z")

</div>

> [@Saravana\_Maadavan](#):
>
> rexml library is used in XML filter only because ruby filter (in this case mutate gsub) is used prior to XML filter?

No, the rexml library is used whenever you set the store\_xml option to true.

---

<div class="post-metadata">

**Author:** ![Saravana\_Maadavan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saravana_maadavan/32/60232_2.png) [@Saravana\_Maadavan](https://discuss.elastic.co/u/Saravana_Maadavan)\
**Post date:** [May 10, 2020, 3:25pm UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621/5 "2020-05-10T15:25:53Z")

</div>

@Badger,

So no other option available to store the full XML? Only option is to strip out the XML tags is it?

Cheers,  
Maadavan

---

<div class="post-metadata">

**Author:** ![Saravana\_Maadavan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/saravana_maadavan/32/60232_2.png) [@Saravana\_Maadavan](https://discuss.elastic.co/u/Saravana_Maadavan)\
**Post date:** [May 21, 2020, 7:04am UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621/6 "2020-05-21T07:04:22Z")

</div>

Logstash Team,

How to avoid getting the below error? I cannot strip the XMLs, I would need the entire XML to be in elastic for aggregations.

> RuntimeError: entity expansion has grown too large

Cheers,  
Maadavan

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 18, 2020, 7:04am UTC](https://discuss.elastic.co/t/logstash-throws-runtime-exception-while-parsing-huge-xml/231621/7 "2020-06-18T07:04:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
