# What happens to existing documents during a crawl?

**URL:** <https://discuss.elastic.co/t/what-happens-to-existing-documents-during-a-crawl/374366>\
**Category:** Elasticsearch\
**Tags:** crawler\
**Created:** [February 11, 2025, 1:43pm UTC](https://discuss.elastic.co/t/what-happens-to-existing-documents-during-a-crawl/374366 "2025-02-11T13:43:06Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![jkyeusun](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jkyeusun/32/141291_2.png) [@jkyeusun](https://discuss.elastic.co/u/jkyeusun)\
**Post date:** [February 11, 2025, 1:43pm UTC](https://discuss.elastic.co/t/what-happens-to-existing-documents-during-a-crawl/374366/1 "2025-02-11T13:43:06Z")

</div>

For the current version of Enterprise Search (8.17), I can't find any documentation on how already indexed documents are managed in a web crawler index, so here are my questions regarding that.

So far I've been looking at the logs to answer my own question but I'd like to know whether the behavior I observed are intended.

- If crawl rules are updated so that some documents that are already indexed are now excluded in new crawls, are those documents automatically deleted when a domain is crawled again? (in my tests, the existing document was not automatically deleted even though the document was denied from being indexed in the new crawl, according to the logs)

- If a domain is crawled again and the content of some pages haven't changed, are those pages indexed again, or those the crawler or index know not to index the page since it hasn't changed? (in my tests, all pages seemed to be indexed again even if the contents haven't changed. One way to test this is to run two crawls back-to-back to see if the logs of the second crawl indicate a document wasn't indexed again)

---

<div class="post-metadata">

**Author:** ![nfeekery](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nfeekery/32/117349_2.png) [@nfeekery](https://discuss.elastic.co/u/nfeekery)\
**Post date:** [February 11, 2025, 1:53pm UTC](https://discuss.elastic.co/t/what-happens-to-existing-documents-during-a-crawl/374366/2 "2025-02-11T13:53:56Z")

</div>

Hi @jkyeusun

> If crawl rules are updated so that some documents that are already indexed are now excluded in new crawls, are those documents automatically deleted when a domain is crawled again? (in my tests, the existing document was not automatically deleted even though the document was denied from being indexed in the new crawl, according to the logs)

If the crawl rule is properly configured, the document _should_ be deleted. Make sure you're running a _full crawl_ and not a _partial crawl_. (partial crawls don't purge documents, these are run by selecting `Crawl` -\> `Crawl with custom settings`)

> If a domain is crawled again and the content of some pages haven't changed, are those pages indexed again, or those the crawler or index know not to index the page since it hasn't changed? (in my tests, all pages seemed to be indexed again even if the contents haven't changed. One way to test this is to run two crawls back-to-back to see if the logs of the second crawl indicate a document wasn't indexed again)

The pages are indexed again. Crawler doesn't check if a page has been updated or not, it just indexes everything it finds.
