# Parsing xml document using xpath

**URL:** <https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320>\
**Category:** Logstash\
**Created:** [September 15, 2015, 1:51pm UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320 "2015-09-15T13:51:00Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ushadatt](https://avatars.discourse-cdn.com/v4/letter/u/edb3f5/32.png) [@ushadatt](https://discuss.elastic.co/u/ushadatt)\
**Post date:** [September 15, 2015, 1:51pm UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320/1 "2015-09-15T13:51:00Z")

</div>

I am trying to parse the following xml data using logstash.. I am able to do it for a single document..But when I am increasing number of documents, its not working..

```
<Book:Body>
    <Book:Head>
        <bookname>Book:Name</bookname>
            <ns:Hello xmlns:ns="www.example.com">
                <ns:BookDetails>
                    <ns:ID>123456</ns:ID>
                    <ns:Name>ABC</ns:Name>
                </ns:BookDetails>
			</ns:Hello xmlns:ns="www.example.com">
    </Book:Head>
<Book:Body>

```

My config file is as given:

```
multiline {
                       pattern => "<Book:Body>"
                        what => "previous"
			negate => "true"
			}
				
                xml {
                        store_xml => "false"
                        source => "message"
			remove_namespaces => "true"
						
                        xpath =>[
                                "/Book/Book/BookDetails/ID/text()","UUID",	
				"/Book/Book/BookDetails/Name/text()","Name"
					]
			}
               
                mutate {
                        add_field => ["IDIndexed", "%{ID}"]
			add_field => ["NameIndexed", "%{Name}"]
				}
```

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 15, 2015, 2:09pm UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320/2 "2015-09-15T14:09:16Z")

</div>

Could you be a bit more specific than "it's not working"? What _do_ you get? Is there anything interesting in the logs? What do you mean by "multiple documents", multiple consecutive Book:Body elements in the same file...?

---

<div class="post-metadata">

**Author:** ![ushadatt](https://avatars.discourse-cdn.com/v4/letter/u/edb3f5/32.png) [@ushadatt](https://discuss.elastic.co/u/ushadatt)\
**Post date:** [September 16, 2015, 4:49am UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320/3 "2015-09-16T04:49:56Z")

</div>

Yes, I mean multiple consecutive Book:Body elements in the same file.. With just one entry like my example, it is parsing the two fields **ID** and **Name** and mutate filter is adding new fields..But with multiple records, it is not able to parse the message and the fields %{ID} and %{Name} appear as it is without any values..Is there something wrong with my multiline pattern or xpath?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [September 16, 2015, 6:00am UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320/4 "2015-09-16T06:00:24Z")

</div>

The multiline pattern looks okay. I suggest you simplify things by removing the xml filter and just emitting messages with the joined XML lines. What happens then if you feed Logstash a file with multiple Book:Body elements?

BTW, your example Book:Body element ends with `<Book:Body>` rather than `</Book:Body>`. I assume that was a typo?

---

<div class="post-metadata">

**Author:** ![ushadatt](https://avatars.discourse-cdn.com/v4/letter/u/edb3f5/32.png) [@ushadatt](https://discuss.elastic.co/u/ushadatt)\
**Post date:** [September 16, 2015, 6:19am UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320/5 "2015-09-16T06:19:14Z")

</div>

Actually I have tried the example again without namespace **ns** , so logstash was able to parse the document, but when I am using ns in all the tags as given in the BOOK:Body elements, it is not parsing it.. I have even used the **remove\_namespaces** tag in the xml filter.. I guess the problem is due to namespace of XML tags.. I was working with this example without namespaces:

```
    <Book>
        <bookname>Book:Name</bookname>
            <Hello>
                <BookDetails>
                    <ID>123456</ID>
                    <Name>ABC</Name>
                <BookDetails>
	</Hello>
 </Book>

```

Yeah sorry for the typo! \</Book:Body\>

---

<div class="post-metadata">

**Author:** ![Navneet\_Mathpal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/navneet_mathpal/32/3677_2.png) [@Navneet\_Mathpal](https://discuss.elastic.co/u/Navneet_Mathpal)\
**Post date:** [September 16, 2015, 11:23am UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320/6 "2015-09-16T11:23:30Z")

</div>

+1 getting the same issue ( remove\_namespaces =\> true # not removing the name spaces)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:29am UTC](https://discuss.elastic.co/t/parsing-xml-document-using-xpath/29320/7 "2017-07-06T05:29:00Z")

</div>


