# How to extract metadata using the Webcrawler

**URL:** <https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [July 14, 2021, 2:39pm UTC](https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681 "2021-07-14T14:39:18Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Marten](https://avatars.discourse-cdn.com/v4/letter/m/4af34b/32.png) [@Marten](https://discuss.elastic.co/u/Marten)\
**Post date:** [July 14, 2021, 2:39pm UTC](https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681/1 "2021-07-14T14:39:18Z")

</div>

Hi there,  
I'm testing the App Search Webcrawler.  
Is there a way to extract more metadata than the current standard ones?  
The documentation mentions something about adding a template ([Web crawler reference | Elastic App Search Documentation [8.4] | Elastic](https://www.elastic.co/guide/en/app-search/current/web-crawler-reference.html#web-crawler-reference-meta-tags-content-extraction)) but I can't find a way to implement this.  
How can I enrich my documents with extra data without changing all my webpages?

Best Regards,

Marten

---

<div class="post-metadata">

**Author:** ![ross.bell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ross.bell/32/77559_2.png) [@ross.bell](https://discuss.elastic.co/u/ross.bell)\
**Post date:** [July 14, 2021, 3:26pm UTC](https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681/2 "2021-07-14T15:26:04Z")

</div>

Hey @Marten,

Could you link us to an example page you're crawling, and/or provide a snippet of the tags content from your crawled pages that you're using to attempt custom document attributes? The instructions you link to are indeed the way to accomplish custom document attributes.

Could you also confirm the version of Enterprise Search you're running?

Thanks  
Ross

---

<div class="post-metadata">

**Author:** ![Marten](https://avatars.discourse-cdn.com/v4/letter/m/4af34b/32.png) [@Marten](https://discuss.elastic.co/u/Marten)\
**Post date:** [July 15, 2021, 6:37am UTC](https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681/3 "2021-07-15T06:37:33Z")

</div>

Hi Ross,

Thanks for your quick reply.  
The page I want to crawl is:

> **[Particulieren (Home) | UWV | Particulieren](https://www.uwv.nl/particulieren/index.aspx)**

I'm using the hosted app search service on Elastic Cloud since yesterday, so I guess that it's the latest version.

Best Regards,

Marten

---

<div class="post-metadata">

**Author:** ![ross.bell](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ross.bell/32/77559_2.png) [@ross.bell](https://discuss.elastic.co/u/ross.bell)\
**Post date:** [July 15, 2021, 3:55pm UTC](https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681/4 "2021-07-15T15:55:16Z")

</div>

Thanks for providing the example. The documentation you originally link to is the only way currently supported. You will need to modify the crawled page(s) to include `<meta ... >` tags that the crawler will recognize and pick up as custom fields.

The good news is that we plan to introduce configurability to the crawler in the future that would not require introducing `<meta>` tags to your crawled content. However, I can't provide a date by which that would be available.

---

<div class="post-metadata">

**Author:** ![Marten](https://avatars.discourse-cdn.com/v4/letter/m/4af34b/32.png) [@Marten](https://discuss.elastic.co/u/Marten)\
**Post date:** [July 16, 2021, 7:17am UTC](https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681/6 "2021-07-16T07:17:37Z")

</div>

Great, I got it now.  
It seems that I just had to add an extra field to the schema with the name of the meta tag.  
The crawler then picked it up automatically.  
This wasn't entirely clear to me from the documentation, but it's clear to me now.

Thanks for your help,

Marten

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 13, 2021, 7:17am UTC](https://discuss.elastic.co/t/how-to-extract-metadata-using-the-webcrawler/278681/7 "2021-08-13T07:17:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
