# How to write a filter for a file containing a mix of xml and non xml messages

**URL:** <https://discuss.elastic.co/t/how-to-write-a-filter-for-a-file-containing-a-mix-of-xml-and-non-xml-messages/312503>\
**Category:** Logstash\
**Created:** [August 19, 2022, 8:38pm UTC](https://discuss.elastic.co/t/how-to-write-a-filter-for-a-file-containing-a-mix-of-xml-and-non-xml-messages/312503 "2022-08-19T20:38:27Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Patr123](https://avatars.discourse-cdn.com/v4/letter/p/ac91a4/32.png) [@Patr123](https://discuss.elastic.co/u/Patr123)\
**Post date:** [August 19, 2022, 8:38pm UTC](https://discuss.elastic.co/t/how-to-write-a-filter-for-a-file-containing-a-mix-of-xml-and-non-xml-messages/312503/1 "2022-08-19T20:38:27Z")

</div>

Hello,  
I have a log file which is a mix of xml and regular (non xml) lines. I need to apply grok filer + xml filter to the lines that has xml block and apply only grok filter to the regular lines. For that I need help in extracting the xml block using the logstash filter but I am not able to do that. If I use the xml filter like:  
` xml { source => "message" target => "xml" }`  
it gets applied to all the messages and the lines without xml block in it are either dropped or comes back as \_xmlparsefailure, also the lines with xml block are parsed correctly.  
I need help in figuring out how to extract the xml block and apply xml filter only if it exists in message.

Thank you.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 19, 2022, 9:21pm UTC](https://discuss.elastic.co/t/how-to-write-a-filter-for-a-file-containing-a-mix-of-xml-and-non-xml-messages/312503/2 "2022-08-19T21:21:48Z")

</div>

You need to determine whether the [message] field contains XML. The easiest way to do that is with an xml filter. If the event gets tagged with "\_xmlparsefailure" then it is not valid XML, so you can make additional processing conditional upon that.

---

<div class="post-metadata">

**Author:** ![Patr123](https://avatars.discourse-cdn.com/v4/letter/p/ac91a4/32.png) [@Patr123](https://discuss.elastic.co/u/Patr123)\
**Post date:** [August 19, 2022, 9:46pm UTC](https://discuss.elastic.co/t/how-to-write-a-filter-for-a-file-containing-a-mix-of-xml-and-non-xml-messages/312503/3 "2022-08-19T21:46:15Z")

</div>

Ahh, got it. So instead of finding a xml block and then applying the xml filter, you are suggesting finding the regular messages using the "\_xmlparsefailure" tag and then do what I need to do.

Thank you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 16, 2022, 9:46pm UTC](https://discuss.elastic.co/t/how-to-write-a-filter-for-a-file-containing-a-mix-of-xml-and-non-xml-messages/312503/4 "2022-09-16T21:46:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
