# \#webcrawler

**URL:** https://discuss.elastic.co/tag/webcrawler/153.md

[Latest](https://discuss.elastic.co/latest.md) · [Categories](https://discuss.elastic.co/categories.md) · [Tags](https://discuss.elastic.co/tags.md)

---

## [Crawl sitemap only](https://discuss.elastic.co/t/crawl-sitemap-only/375169)

<div class="topic-metadata">

**Author:** [@pngworkforce](https://discuss.elastic.co/u/pngworkforce)\
**Replies:** 7\
**Last updated:** [March 20, 2025, 7:53am UTC](https://discuss.elastic.co/t/crawl-sitemap-only/375169 "2025-03-20T07:53:55Z")

</div>

Hello! I have a question similar to the one listed here -\> Web crawler is crawling URLs that are not on the sitemap I have added a sitemap index as the sitemap to my crawler in Elastic Cloud UI. How do I instruct the …

---

## [Crawler conditional document ingest based on modified date](https://discuss.elastic.co/t/crawler-conditional-document-ingest-based-on-modified-date/373193)

<div class="topic-metadata">

**Author:** [@Dhaliwal-BCGOV](https://discuss.elastic.co/u/Dhaliwal-BCGOV)\
**Replies:** 0\
**Last updated:** [January 14, 2025, 5:30pm UTC](https://discuss.elastic.co/t/crawler-conditional-document-ingest-based-on-modified-date/373193 "2025-01-14T17:30:58Z")

</div>

Hi, Only ingest existing pages using crawler if page modified date has been updated.

---

## [Dec 12th, 2024: \[EN\] Swing through web content like a superhero](https://discuss.elastic.co/t/dec-12th-2024-en-swing-through-web-content-like-a-superhero/371474)

<div class="topic-metadata">

**Author:** [@ppf2](https://discuss.elastic.co/u/ppf2)\
**Replies:** 0\
**Last updated:** [December 12, 2024, 8:00am UTC](https://discuss.elastic.co/t/dec-12th-2024-en-swing-through-web-content-like-a-superhero/371474 "2024-12-12T08:00:00Z")

</div>

Swing through web content like a superhero Web crawlers at Elastic have undergone multiple evolutions throughout the years to adapt to the rapidly changing landscape of data ingestion (e.g., recent advancements in gen…
