# Logstash xml input configuration: for multiple documents

**URL:** <https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119>\
**Category:** Logstash\
**Created:** [July 29, 2020, 8:07pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119 "2020-07-29T20:07:59Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 29, 2020, 8:07pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/1 "2020-07-29T20:07:59Z")

</div>

Hi Team

I'm using http\_poller to poll an end point which gives xml data as response. I'm trying to send this xml data to elasticsearch.

But when I run logstash, I see logstash is failing. Please have a look at below config.

xmldata:

few line of my xml data:

> ```
> > <?xml version="1.0" encoding="UTF-8"?>
> > <feed xmlns="http://www.w3.org/2005/Atom">
> > <generator version="1.0">Alfresco (1.0)</generator>
> > <link rel="self" href="links" />
> > <id>random_id</id>
> > <title>Activities Site</title>
> > <updated>2020-07-29T12:53:16.000-07:00</updated>
> > <entry xmlns='http://www.w3.org/2005/Atom'>
> > <title type="html"><overview></title>
> > <link rel="alternate" type="text/html" href="random link" />
> > <id>249,535,933</id>
> > <updated>2020-07-29T12:53:16.000-07:00</updated>
> > <summary type="html">
> > <![DATA[<a href="random link</a> downloaded document <a href="random link">Overview</a>]]>
> > </summary>
> > <author>
> > <name>name</name>
> > <uri>random</uri>
> > </author>
> > </entry>
> > <entry xmlns='http://www.w3.org/2005/Atom'>
> > <title type="html"><random></title>
> > <link rel="alternate" type="text/html" href="randomuri" />
> > <id>249,535,867</id>
> > <updated>2020-07-29T12:53:10.000-07:00</updated>
> > <summary type="html">
> > <![CDATA[<a href="random">Name</a> download <a href="random">intro</a>]]>
> > </summary>
> > <author>
> > <name>Name</name>
> > <uri>random</uri>
> > </author>
> > </entry>
> 
> ```

Logstash.conf:

```
input 
{
	http_poller 
	{
		urls => 
		{

		test1 =>
				{		
				url=>"randomhost"
				method => get
                user => " *********"
                password => " *******"
                headers => {
                   "Content-Type" => "text/xml; charset=UTF-8"
                   }
				}
		}
		request_timeout => 60
		schedule => { cron => "* * * * * UTC"}
        
	}
}
filter {
  xml { 
  source => "message" 
  target => "theXML" 
  
  }
}

#output { stdout { codec => rubydebug } }

output {

    elasticsearch {
      index => "logstash-xmldata"
      hosts => "http://elasticsearchhost:80"
      user => " ****"
      password => " ******"
    }
  }

```

output:

`> [2020-07-29T20:00:02,148][WARN][logstash.outputs.elasticsearch][main][41e2884444551e51d0256ad578d1476c2186e932e0995e3ce551bbd4c4286a6a] Could not index event to Elasticsearch. {:status=>400, :action=>["index", {:_id=>nil, :_index=>"logstash-xmldata", :routing=>nil, :_type=>"_doc"}, #<LogStash::Event:0x11bd3dd9>], :response=>{"index"=>{"_index"=>"logstash-xmldata", "_type"=>"_doc", "_id"=>"r1MpnHMBz8okLa_0Chk8", "status"=>400, "error"=>{"type"=>"mapper_parsing_exception", "reason"=>"object mapping for [theXML.title] tried to parse field [null] as object, but found a concrete value"}}}}`

please suggest.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 29, 2020, 8:10pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/2 "2020-07-29T20:10:28Z")

</div>

The problem appears to be with [theXML][title]

```auto
<title>Activities Site</title>

```

that is a string (a "concrete value") but the mapping in elasticsearch expects it to be an object.

Check the [mapping](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-get-mapping.html) in elasticsearch.

Read [this](https://discuss.elastic.co/t/logstash-errors-mapper-parsing-exception-vs-illegal-argument-exception/236783/3) post and then [this](https://discuss.elastic.co/t/problem-logstash-outputs-elasticsearch-could-not-index-event-to-elasticsearch-wazuh-alerts-3-x-2020-05-30/235038/6) post.

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 29, 2020, 8:38pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/3 "2020-07-29T20:38:45Z")

</div>

Hi @Badger

Got it. I'm able to ingest data to kibana by following the details in suggested posts.

But in kibana I see all the data in one xml filed ? Any suggestions on this ?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 29, 2020, 9:05pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/4 "2020-07-29T21:05:50Z")

</div>

logstash will have placed all of the parsed XML inside the top-level theXML field. If you want the object to be moved to the top level you can use a ruby filter, like [this](https://discuss.elastic.co/t/how-to-dynamically-move-nested-key-value-to-root-level/180006/2).

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 29, 2020, 9:40pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/5 "2020-07-29T21:40:58Z")

</div>

Hi @Badger

This is the kibana output without ruby filter. all the data is under **theXML.entry** field.

[

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/1/1/11f2b70b333ca0627e72245592d2be683c4912fd.png)]

This is the kibana output with the below ruby filter

> ruby {  
> code =\> '  
> event.get("theXML").each { |k, v|  
> event.set(k,v)  
> }  
> event.remove("theXML")  
> '  
> }

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/2/5/251c96f1da4ed31bdedf7132916003272c429bc3.png) . the data is under **entry field**.

Do I need make any changes to the ruby filter ? something like event.get(theXML.entry) .  
Also, how about entry.updated ? will that be parsed as well

Please suggest

Thank you

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 29, 2020, 10:20pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/6 "2020-07-29T22:20:04Z")

</div>

Well, the sample data you posted is not valid XML, and I suspect the structure is different as you get through more entries.

You might want to use the 'force\_array =\> false' option on the xml filter.

Since there are multiple \<entry\> elements that is always going to be an array. You might want to use a split filter to break those up into separate events. Maybe not, depends on your use case.

How you end up with a entry.updated array I cannot guess.

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 29, 2020, 10:32pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/7 "2020-07-29T22:32:09Z")

</div>

I posted the first few lines of the data. So, it looks like invalid. i tried converting to json (using external editors) and it worked

let me try these options and see. However, I'm still wondering about entry.updated and other similar fields.

Will let you know If i find anything interesting.

Thanks  
Rahul

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 30, 2020, 3:34pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/8 "2020-07-30T15:34:09Z")

</div>

Hi @Badger

Following filter config using split worked well.

> filter {  
> xml {  
> source =\> "message"  
> target =\> "theXML"  
> force\_array =\> false
> 
> }  
> ruby {  
> code =\> '  
> event.get("theXML").each { |k, v|  
> event.set(k,v)  
> }  
> event.remove("theXML")  
> '  
> }  
> split {  
> field =\> "entry"  
> remove\_field =\> "message"  
> }  
> }

Data in Kibana is good but I see \_jsonparsefailure tag. Is there anyway I can understand what is failing ?

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/0/30c197077c7b87653c3b39573acc3e17037c160f.png)

---

<div class="post-metadata">

**Author:** ![Jenni](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jenni/32/29684_2.png) [@Jenni](https://discuss.elastic.co/u/Jenni)\
**Post date:** [July 30, 2020, 4:01pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/9 "2020-07-30T16:01:09Z")

</div>

I think that comes from your input. You didn't set the `codec` parameter, so it tried to use its default:

- [https://www.elastic.co/guide/en/logstash/current/plugins-inputs-http\_poller.html#plugins-inputs-http\_poller-codec](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-http_poller.html#plugins-inputs-http_poller-codec)
- [https://www.elastic.co/guide/en/logstash/7.8/plugins-codecs-json.html#\_description\_173](https://www.elastic.co/guide/en/logstash/7.8/plugins-codecs-json.html#_description_173)

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 30, 2020, 4:16pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/10 "2020-07-30T16:16:41Z")

</div>

oh yea got it. Makes sense. @Jenni

Is there a way to specify xml ? I didnt see it in documentation.

---

<div class="post-metadata">

**Author:** ![Jenni](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jenni/32/29684_2.png) [@Jenni](https://discuss.elastic.co/u/Jenni)\
**Post date:** [July 30, 2020, 4:45pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/11 "2020-07-30T16:45:06Z")

</div>

I didn't see anything either. I think you can just use `plain` for this as you already have an xml filter anyway.

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 30, 2020, 4:48pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/12 "2020-07-30T16:48:56Z")

</div>

sure @Jenni

Thank you 🙂

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 31, 2020, 4:30pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/13 "2020-07-31T16:30:19Z")

</div>

> [@rahulnama](#):
>
> filter {  
> xml {  
> source =\> "message"  
> target =\> "theXML"  
> force\_array =\> false
> 
> }  
> ruby {  
> code =\> '  
> event.get("theXML").each { |k, v|  
> event.set(k,v)  
> }  
> event.remove("theXML")  
> '  
> }  
> split {  
> field =\> "entry"  
> remove\_field =\> "message"  
> }  
> }

Hi @Badger @Jenni

The above conf worked but message (field) is adding to every document (in es) with all the data. Any inputs to avoid this ?

---

<div class="post-metadata">

**Author:** ![Jenni](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jenni/32/29684_2.png) [@Jenni](https://discuss.elastic.co/u/Jenni)\
**Post date:** [July 31, 2020, 4:56pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/14 "2020-07-31T16:56:28Z")

</div>

(Edit: There was a wrong test and assumption that split doesn't call remove\_field if there was only one entry. But Badger proved me wrong below. This is long and unnessecarry, so I am getting rid of it. Have a look at the edit history of this post, if you are interested in my idiotism 🙂 )

If you move the `remove_field` option to a separate mutate filter, it should work.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 31, 2020, 5:01pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/15 "2020-07-31T17:01:46Z")

</div>

> [@rahulnama](#):
>
> xml {  
> source =\> "message"  
> target =\> "theXML"  
> force\_array =\> false
> 
> }

I would add the remove\_field =\> ["message"] to the xml filter, so that it is only removed if it is successfully parsed.

@Jenni, the split filter will not decorate the event (i.e. filter\_matched is not called) if the field [is a string that does not contain the terminator](https://github.com/logstash-plugins/logstash-filter-split/blob/925287ba5e6fb0f79f20a835a63c4e687f38e15a/lib/logstash/filters/split.rb#L80).

---

<div class="post-metadata">

**Author:** ![Jenni](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jenni/32/29684_2.png) [@Jenni](https://discuss.elastic.co/u/Jenni)\
**Post date:** [July 31, 2020, 5:05pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/16 "2020-07-31T17:05:30Z")

</div>

Ah. Sorry. Thanks. I had wrongfully assumed that add\_field would keep my array as an array.

(But a feed with only one entry could still cause problems with split because it would be a hash instead of an array, wouldn't it?)

---

<div class="post-metadata">

**Author:** ![rahulnama](https://avatars.discourse-cdn.com/v4/letter/r/8c91f0/32.png) [@rahulnama](https://discuss.elastic.co/u/rahulnama)\
**Post date:** [July 31, 2020, 6:59pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/17 "2020-07-31T18:59:44Z")

</div>

Adding the remove\_field as a separate filter worked as well. However, adding it in xml filter would make more sense.

Thank you both 🙂

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 31, 2020, 7:35pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/18 "2020-07-31T19:35:30Z")

</div>

If the field were a hash you would get

```
logger.warn("Only String and Array types are splittable. field:#{@field} is of type = #{original_value.class}")
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 28, 2020, 7:35pm UTC](https://discuss.elastic.co/t/logstash-xml-input-configuration-for-multiple-documents/243119/19 "2020-08-28T19:35:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
