# Extraction rules for meta tag without "elastic" using a XPath selector

**URL:** <https://discuss.elastic.co/t/extraction-rules-for-meta-tag-without-elastic-using-a-xpath-selector/386450>\
**Category:** Elastic Search\
**Tags:** open-crawler\
**Created:** [May 21, 2026, 7:05pm UTC](https://discuss.elastic.co/t/extraction-rules-for-meta-tag-without-elastic-using-a-xpath-selector/386450 "2026-05-21T19:05:16Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![creekinouryard](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/creekinouryard/32/147088_2.png) [@creekinouryard](https://discuss.elastic.co/u/creekinouryard)\
**Post date:** [May 21, 2026, 7:05pm UTC](https://discuss.elastic.co/t/extraction-rules-for-meta-tag-without-elastic-using-a-xpath-selector/386450/1 "2026-05-21T19:05:17Z")

</div>

Using the latest version of Open Crawler and having trouble extracting meta tag values  
We have a HTML page as follows:

```auto
<html lang="en-ca" xmlns=" http://www.w3.org/1999/xhtml ">
<head>
    <!--<meta http-equiv=Content-Type content="text/html; charset=windows-1252">-->
    <meta charset="utf-8" />
    <title>Example Title </title>
    <meta name="dc.date" content="2017-06-23" />
    <meta name="topic_obj_id" content="752033de-8aaf-46a4-b92e-d7a9fd0e3fad" /> 
    <meta name="location_obj_id" content="Victoria" />
</head>
<body style="font-family:Calibri; font-size:12pt">
   Content Goes Here
</body>
</html

```

We are trying to use extraction rules to get the fields called dc.date and location\_obj\_id but they are always coming up empty ""  
Extraction rules we are using are below.

Is there something we are doing wrong?

We have tried many alternative in the xpath selector but they all return with the field being empty.

This is is the current attempt.

```auto
extraction_rulesets:

    - rules:
      - action: "extract" 
        field_name: "dc_date"
        selector: /html/head/meta\[@name="dc.date"]/@content
        join_as: "string"
        source: "html"
      - action: "extract"
         field_name: "topic_obj_id"            
         selector: /html/head/meta\[@name="topic_obj_id"]/@content
         join_as: "string"             
        source: "html"

```

---

<div class="post-metadata">

**Author:** ![Asmaa\_Oufkir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/asmaa_oufkir/32/31746_2.png) [@Asmaa\_Oufkir](https://discuss.elastic.co/u/Asmaa_Oufkir)\
**Post date:** [May 29, 2026, 5:42pm UTC](https://discuss.elastic.co/t/extraction-rules-for-meta-tag-without-elastic-using-a-xpath-selector/386450/3 "2026-05-29T17:42:50Z")

</div>

Always use the form //\*[local-name()='meta' and @name='...']/@content in your Open Crawler extraction rules when dealing with XHTML or HTML5 documents containing xmlns.

This is the most portable solution, independent of specific crawler features, and it is robust regardless of the presence (or absence) of namespaces.

If you need to scale (millions of pages), it's best to pre-calculate the values ​​with a small XPath script in Python (lxml) or Java (javax.xml.xpath) to validate your rules before deploying them to Open Crawler.
