# Pdf documents specified in the sitemap are not being indexed by web crawler

**URL:** <https://discuss.elastic.co/t/pdf-documents-specified-in-the-sitemap-are-not-being-indexed-by-web-crawler/361454>\
**Category:** Elasticsearch\
**Created:** [June 14, 2024, 12:23am UTC](https://discuss.elastic.co/t/pdf-documents-specified-in-the-sitemap-are-not-being-indexed-by-web-crawler/361454 "2024-06-14T00:23:19Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![vsachdeva](https://avatars.discourse-cdn.com/v4/letter/v/e56c9b/32.png) [@vsachdeva](https://discuss.elastic.co/u/vsachdeva)\
**Post date:** [June 14, 2024, 12:23am UTC](https://discuss.elastic.co/t/pdf-documents-specified-in-the-sitemap-are-not-being-indexed-by-web-crawler/361454/1 "2024-06-14T00:23:19Z")

</div>

Hello,  
I am trying to index a set of PDF documents using the web crawler. The deployment is in GCP cloud and the PDF documents are specified in the sitemap, which is documented in the robots.txt file. I am not using workspace solution. Do I need to define an attachment processor in the ingestion pipeline? Thanks

---

<div class="post-metadata">

**Author:** ![vsachdeva](https://avatars.discourse-cdn.com/v4/letter/v/e56c9b/32.png) [@vsachdeva](https://discuss.elastic.co/u/vsachdeva)\
**Post date:** [June 14, 2024, 3:58am UTC](https://discuss.elastic.co/t/pdf-documents-specified-in-the-sitemap-are-not-being-indexed-by-web-crawler/361454/2 "2024-06-14T03:58:00Z")

</div>

The log explorer is showing the message: Unexpected content type application/pdf for a crawl task with type=content  
for each pdf document in the sitemap.

---

<div class="post-metadata">

**Author:** ![vsachdeva](https://avatars.discourse-cdn.com/v4/letter/v/e56c9b/32.png) [@vsachdeva](https://discuss.elastic.co/u/vsachdeva)\
**Post date:** [June 14, 2024, 4:01am UTC](https://discuss.elastic.co/t/pdf-documents-specified-in-the-sitemap-are-not-being-indexed-by-web-crawler/361454/3 "2024-06-14T04:01:32Z")

</div>

I believe the issue should be resolved as per the documentation specified here:

> **[Web crawler reference | App Search documentation \[8.14\] | Elastic](https://www.elastic.co/guide/en/app-search/current/web-crawler-reference.html#web-crawler-reference-binary-content-extraction)**

I will update the web crawler configuration and provide an update.
