# Field Extraction/Parsing

**URL:** <https://discuss.elastic.co/t/field-extraction-parsing/171352>\
**Category:** Beats\
**Tags:** filebeat\
**Created:** [March 7, 2019, 4:27pm UTC](https://discuss.elastic.co/t/field-extraction-parsing/171352 "2019-03-07T16:27:45Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![elborni96](https://avatars.discourse-cdn.com/v4/letter/e/b19c9b/32.png) [@elborni96](https://discuss.elastic.co/u/elborni96)\
**Post date:** [March 7, 2019, 4:27pm UTC](https://discuss.elastic.co/t/field-extraction-parsing/171352/1 "2019-03-07T16:27:45Z")

</div>

Hi everyone, I'm a new user of the ELK stack.

I'm monitoring a file and I would like to extract fields (hereinafter a small extraction):

\<Item\> \<title\> Vulnerability in Cisco Products (March 6, 2019) \</ title\> \<Link\> [https://www.certnazionale.it/news/2019/03/07/vulnerabilita-in-prodotti-cisco-6-marzo-2019/](https://www.certnazionale.it/news/2019/03/07/vulnerabilita-in-prodotti-cisco-6-marzo-2019/) \</ link\> \<pubDate\> Thu, 07 Mar 2019 10:01:10 +0000 \</ pubDate\> \<dc: creator\> \<! [CDATA [National CERT]]\> \</ dc: creator\> \<Category\> \<! [CDATA [Vulnerability]]\> \</ category\> \<Category\> \<! [CDATA [cisco]]\> \</ category\> \<guid isPermaLink = "false"\> [https://www.certnazionale.it/?post\_type=news&#038;p=1946](https://www.certnazionale.it/?post_type=news&#038;p=1946) \</ guid\> \<description\> \<! [CDATA [Cisco has released several security updates on March 6, 2019 that address multiple vulnerabilities in different products.]]\> \</ description\> \</ Item\>

I would like to extract the folowing fields (with relative values):

**title=Vulnerability in Cisco Products (March 6, 2019)**

**pudDate=Thu, 07 Mar 2019 10:01:10 +0000**

**description=Cisco has released several security updates on March 6, 2019 that address multiple vulnerabilities in different products.**

Can you explain me how to extract fields at index time with Filebeat or Logstash?

Thanks

---

<div class="post-metadata">

**Author:** ![pierhugues](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pierhugues/32/48383_2.png) [@pierhugues](https://discuss.elastic.co/u/pierhugues)\
**Post date:** [March 8, 2019, 4:50pm UTC](https://discuss.elastic.co/t/field-extraction-parsing/171352/2 "2019-03-08T16:50:22Z")

</div>

Hello @elborni96, You cannot do that in Filebeat at the moment, you will need to use logstash and the [logstash-filter-xml](https://www.elastic.co/guide/en/logstash/current/plugins-filters-xml.html)

---

<div class="post-metadata">

**Author:** ![elborni96](https://avatars.discourse-cdn.com/v4/letter/e/b19c9b/32.png) [@elborni96](https://discuss.elastic.co/u/elborni96)\
**Post date:** [March 11, 2019, 8:18am UTC](https://discuss.elastic.co/t/field-extraction-parsing/171352/3 "2019-03-11T08:18:12Z")

</div>

Hi @pierhugues,

I need to understand how to parse / extract fields of events that come to me through the Filbeat, regardless of the type of file (XML, HTML, JSON, syslog ...).

Could you explain how to do it via LogStash?

Regards,  
Mirko

---

<div class="post-metadata">

**Author:** ![pierhugues](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pierhugues/32/48383_2.png) [@pierhugues](https://discuss.elastic.co/u/pierhugues)\
**Post date:** [March 11, 2019, 12:59pm UTC](https://discuss.elastic.co/t/field-extraction-parsing/171352/4 "2019-03-11T12:59:31Z")

</div>

> (XML, HTML, JSON, syslog ...).

The big picture would be something like this.

- Define different Filebeat inputs per document types and [use fields to add a field](https://www.elastic.co/guide/en/beats/filebeat/6.6/filebeat-input-log.html#filebeat-input-log-fields) to identify the type of document.
- Use the Logstash output in filebeat.
- Use the [beats input in logstash](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-beats.html)
- Define a flow to send XML events to the [logstash-xml-filter](https://www.elastic.co/guide/en/logstash/current/plugins-filters-xml.html) using [conditionals](https://www.elastic.co/guide/en/logstash/current/event-dependent-configuration.html#conditionals).

After that you can what you want with the data of the event.

---

<div class="post-metadata">

**Author:** ![elborni96](https://avatars.discourse-cdn.com/v4/letter/e/b19c9b/32.png) [@elborni96](https://discuss.elastic.co/u/elborni96)\
**Post date:** [March 11, 2019, 2:03pm UTC](https://discuss.elastic.co/t/field-extraction-parsing/171352/5 "2019-03-11T14:03:06Z")

</div>

@pierhugues

I have defined Filebeat input that monitor my file (/home/root/test.log).

After i send monitored data to Logstash using the output function.

I have configured yet on Logstash the beats input, but i need to parse the content of the received events.

These events can be xml, syslog, json or other format; so i need to understand how to extract certain fields.

On Splunk Enterprise i can parse or extract fields using the following syntax:

EXTRACT-host\_field = $regex\_to\_extract\_host

How can do the same using parsing/extraction features of Logstash?

Can i create custom searchable fields valorized from the regex results?

Thanks,  
Mirko

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 8, 2019, 2:03pm UTC](https://discuss.elastic.co/t/field-extraction-parsing/171352/6 "2019-04-08T14:03:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
