# Page not indexed if a content extraction rule with CSS selector fails if the references element is not part of the page

**URL:** <https://discuss.elastic.co/t/page-not-indexed-if-a-content-extraction-rule-with-css-selector-fails-if-the-references-element-is-not-part-of-the-page/346647>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [November 7, 2023, 6:49pm UTC](https://discuss.elastic.co/t/page-not-indexed-if-a-content-extraction-rule-with-css-selector-fails-if-the-references-element-is-not-part-of-the-page/346647 "2023-11-07T18:49:59Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![sebastianboelling](https://avatars.discourse-cdn.com/v4/letter/s/258eb7/32.png) [@sebastianboelling](https://discuss.elastic.co/u/sebastianboelling)\
**Post date:** [November 7, 2023, 6:49pm UTC](https://discuss.elastic.co/t/page-not-indexed-if-a-content-extraction-rule-with-css-selector-fails-if-the-references-element-is-not-part-of-the-page/346647/1 "2023-11-07T18:49:59Z")

</div>

Hi all,

we are using content extraction rules with CSS selectors as described here: [Web crawler content extraction rules | Enterprise Search documentation [8.11] | Elastic](https://www.elastic.co/guide/en/enterprise-search/current/crawler-extraction-rules.html#crawler-extraction-rules-css-selectors)

We've found out that a page is NOT indexed if the element referenced in the rule is NOT existing in the page. That means, the crawler is not very fault tolerant.

For example we want do extract a meta tag to the string field **displayurl** which is referenced by the following CSS selector: `html/head/link[@rel="canonical"]/@href`

How can we extract information which is not available on each page?

Segards

Sebastian

---

<div class="post-metadata">

**Author:** ![video](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/video/32/117114_2.png) [@video](https://discuss.elastic.co/u/video)\
**Post date:** [November 22, 2023, 8:44pm UTC](https://discuss.elastic.co/t/page-not-indexed-if-a-content-extraction-rule-with-css-selector-fails-if-the-references-element-is-not-part-of-the-page/346647/2 "2023-11-22T20:44:54Z")

</div>

Hi @sebastianboelling,

> We've found out that a page is NOT indexed if the element referenced in the rule is NOT existing in the page.

Am I right in assuming you don't see a document in the Elasticsearch index representing the page if the CSS selector returns an empty result?

Could you please provide a URL if it's public?

Also, could you please have a look at the [crawler event logs](https://www.elastic.co/guide/en/app-search/current/web-crawler-events-logs-reference.html) for the pages that aren't indexed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 20, 2023, 8:45pm UTC](https://discuss.elastic.co/t/page-not-indexed-if-a-content-extraction-rule-with-css-selector-fails-if-the-references-element-is-not-part-of-the-page/346647/3 "2023-12-20T20:45:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
