# Does Elastic Web Crawler supports noindex and nofollow directive

**URL:** <https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [October 11, 2023, 1:22pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821 "2023-10-11T13:22:00Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![sebastianboelling](https://avatars.discourse-cdn.com/v4/letter/s/258eb7/32.png) [@sebastianboelling](https://discuss.elastic.co/u/sebastianboelling)\
**Post date:** [October 11, 2023, 1:22pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/1 "2023-10-11T13:22:00Z")

</div>

Hi all,

does the Elastic Web Crawler supports **noindex** and **nofollow** directive? I've found this feature only on the App Search Web Crawler reference [Web crawler reference | App Search documentation [8.10] | Elastic](https://www.elastic.co/guide/en/app-search/current/web-crawler-reference.html) and not at the Elastic Web Crawler documentation.

Best regards

Sebastian

---

<div class="post-metadata">

**Author:** ![Sean\_Story](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sean_story/32/69987_2.png) [@Sean\_Story](https://discuss.elastic.co/u/Sean_Story)\
**Post date:** [October 11, 2023, 2:30pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/2 "2023-10-11T14:30:09Z")

</div>

Yes, the `noindex` and `nofollow` tags are also supported in the Elastic Web Crawler. These are documented here: [Optimizing web content for the web crawler | Enterprise Search documentation [8.10] | Elastic](https://www.elastic.co/guide/en/enterprise-search/current/crawler-content.html#crawler-content-robots-meta-tags)

---

<div class="post-metadata">

**Author:** ![sebastianboelling](https://avatars.discourse-cdn.com/v4/letter/s/258eb7/32.png) [@sebastianboelling](https://discuss.elastic.co/u/sebastianboelling)\
**Post date:** [October 13, 2023, 8:26am UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/3 "2023-10-13T08:26:52Z")

</div>

Hi @Sean_Story,

thanks. We found out that **noindex** and **nofollow** working fine. We've also found out that the Crawler does not delete a page from index if an already indexed page changes from **INDEX** to **NOINDEX**.

How can we force or achive the deletion of a page in that case?

Best regards

Sebastian

---

<div class="post-metadata">

**Author:** ![Sean\_Story](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sean_story/32/69987_2.png) [@Sean\_Story](https://discuss.elastic.co/u/Sean_Story)\
**Post date:** [October 31, 2023, 4:47pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/4 "2023-10-31T16:47:36Z")

</div>

> We've also found out that the Crawler does not delete a page from index if an already indexed page changes from **INDEX** to **NOINDEX**.

That makes sense. Crawler only deletes pages that result in 404's on re-crawl. This is to ensure that legacy documents that simply drop off a link tree don't drop out of your search capabilities.

You can manually delete these documents from the index, and they will not be picked up again, due to the NOINDEX.

---

<div class="post-metadata">

**Author:** ![sebastianboelling](https://avatars.discourse-cdn.com/v4/letter/s/258eb7/32.png) [@sebastianboelling](https://discuss.elastic.co/u/sebastianboelling)\
**Post date:** [November 7, 2023, 9:36am UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/5 "2023-11-07T09:36:06Z")

</div>

Hi @Sean_Story,

from our perspective this might be the wrong interpretation of the NOINDEX specification and meaning.

If a page says NOINDEX it should not be indexed and removed from the index. If I have a look to the Google developer docs, they do it in that way:

[Block Search Indexing with noindex | Google Search Central | Documentation | Google for Developers](https://developers.google.com/search/docs/crawling-indexing/block-indexing?hl=en)

> We have to crawl your page in order to see `<meta>` tags and HTTP headers. If a page is still appearing in results, it's probably because we haven't crawled the page since you added the `noindex` rule.

I think we should not ignore the NOINDEX if an webmaster or editor/content provider decides to set this TAG.

Could you please re check whether it is possible to implement that feature?

By the way: Does the the crawler supports **X-Robots-Tag: noindex**

Regards

Sebastian

---

<div class="post-metadata">

**Author:** ![Sean\_Story](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sean_story/32/69987_2.png) [@Sean\_Story](https://discuss.elastic.co/u/Sean_Story)\
**Post date:** [November 7, 2023, 8:56pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/6 "2023-11-07T20:56:51Z")

</div>

Hi @sebastianboelling ,

I can see how you might disagree with the interpretation, but our crawler has been intentionally built in a way that biases towards keeping data searchable as opposed to only being able to search whatever still matches the current crawl configs. For example, if your crawl depth is 2, and a page's links depth goes from 2 to 3, that page remains in the index and is not dropped unless manually deleted.

> If I have a look to the Google developer docs

It is not a goal of ours to maintain behavior parity with Google.

> Could you please re check whether it is possible to implement that feature?

If you have a support relationship with Elastic, I suggest you work with your support representative to file an Enhancement Request. This is typically how we capture an prioritize ideas from the community.

> Does the the crawler supports **X-Robots-Tag: noindex**

No. The robots meta tags that we support are documented here: [Optimizing web content for the web crawler | Enterprise Search documentation [8.11] | Elastic](https://www.elastic.co/guide/en/enterprise-search/current/crawler-content.html#crawler-content-robots-meta-tags)

---

<div class="post-metadata">

**Author:** ![sebastianboelling](https://avatars.discourse-cdn.com/v4/letter/s/258eb7/32.png) [@sebastianboelling](https://discuss.elastic.co/u/sebastianboelling)\
**Post date:** [November 8, 2023, 7:54am UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/7 "2023-11-08T07:54:33Z")

</div>

Hi @Sean_Story ,

thanks for your explantions. I raised up a support request for a feature request - as you suggested.

I maybe found out that this behavior (or a really similar one) for the App Search Web Crawler was fixed with 8.3.0. There is a knowledge base entry which describes the switch from `index` to `noindex`. [Elastic Support Hub](https://support.elastic.co/knowledge/893b3538)

Maybe you have time to look at the knowledge base article. Maybe it will help to change the behavior of the Elastic Web Crawler or to offer both variants.

Regards

Sebastian

---

<div class="post-metadata">

**Author:** ![Sean\_Story](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sean_story/32/69987_2.png) [@Sean\_Story](https://discuss.elastic.co/u/Sean_Story)\
**Post date:** [November 8, 2023, 4:36pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/8 "2023-11-08T16:36:22Z")

</div>

@sebastianboelling I owe you an apology. From that Support Hub page, I was able to track down where this behavior was changed in the App Search Crawler, and from there was able to find tests that imply that this behavior _should_ actually work as you expected in the Elastic Crawler. That is my mistake/misunderstanding, and I'm sorry for pushing against this bug report earlier.

Can you confirm that you've run a full crawl on this index, and not just a "reapply crawl rules" crawl since adding the `noindex` meta tag?

---

<div class="post-metadata">

**Author:** ![sebastianboelling](https://avatars.discourse-cdn.com/v4/letter/s/258eb7/32.png) [@sebastianboelling](https://discuss.elastic.co/u/sebastianboelling)\
**Post date:** [November 10, 2023, 1:35pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/9 "2023-11-10T13:35:41Z")

</div>

Hi @Sean_Story,

no worries. We can investigate a bit deeper. But what is the expected behavior for the Elastic Web Crawler from your point of view? In which case pages are deleted?

case 1: page with `index` meta tag is in index -\> page switches to `noindex` -\> **full** re-crawl -\> page deleted from index: _yes_ vs. _no_

case 2: page with `index` meta tag is in index -\> page switches to `noindex` -\> **partial** crawl -\> page deleted from index: _yes_ vs. _no_

case 3: page with `index` meta tag is in index -\> page is deleted -\> HTTP 404 -\> **full** re-crawl -\> page deleted from index: _yes_ vs. _no_

case 4: page with `index` meta tag is in index -\> page is deleted -\> HTTP 404 -\> **partial** crawl -\> page deleted from index: _yes_ vs. _no_

Other cases for deleting?

We only need to understand and can the investigate and decide how to delete or force deletion of pages from our indices.

Best regards

Sebastian

---

<div class="post-metadata">

**Author:** ![Sean\_Story](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sean_story/32/69987_2.png) [@Sean\_Story](https://discuss.elastic.co/u/Sean_Story)\
**Post date:** [November 10, 2023, 5:10pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/10 "2023-11-10T17:10:51Z")

</div>

> case 1: page with `index` meta tag is in index -\> page switches to `noindex` -\> **full** re-crawl -\> page deleted from index

Yes. Or at least, it should, based on what I'm seeing in the code. I believe you're reporting that this is not working?

> case 2: page with `index` meta tag is in index -\> page switches to `noindex` -\> **partial** crawl -\> page deleted from index

No. The purge phase is not run on partial crawls.

> case 3: page with `index` meta tag is in index -\> page is deleted -\> HTTP 404 -\> **full** re-crawl -\> page deleted from index

Yes.

> case 4: page with `index` meta tag is in index -\> page is deleted -\> HTTP 404 -\> **partial** crawl -\> page deleted from index

No. The purge phase is not run on partial crawls.

> Other cases for deleting?

I don't believe so.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 8, 2023, 5:10pm UTC](https://discuss.elastic.co/t/does-elastic-web-crawler-supports-noindex-and-nofollow-directive/344821/11 "2023-12-08T17:10:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
