# Cannot crawl a website

**URL:** <https://discuss.elastic.co/t/cannot-crawl-a-website/279825>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [July 28, 2021, 8:29am UTC](https://discuss.elastic.co/t/cannot-crawl-a-website/279825 "2021-07-28T08:29:50Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Marten](https://avatars.discourse-cdn.com/v4/letter/m/4af34b/32.png) [@Marten](https://discuss.elastic.co/u/Marten)\
**Post date:** [July 28, 2021, 8:29am UTC](https://discuss.elastic.co/t/cannot-crawl-a-website/279825/1 "2021-07-28T08:29:50Z")

</div>

Hello,

I'm trying to crawl a website (www.werk.nl) but the crawler gets a 599 error on robots.txt.  
If I use my browser, I can open robots.txt and the website without problems.  
What can be the issue?

Marten

---

<div class="post-metadata">

**Author:** ![Marten](https://avatars.discourse-cdn.com/v4/letter/m/4af34b/32.png) [@Marten](https://discuss.elastic.co/u/Marten)\
**Post date:** [July 28, 2021, 11:52am UTC](https://discuss.elastic.co/t/cannot-crawl-a-website/279825/2 "2021-07-28T11:52:33Z")

</div>

Btw, I'm using AppSearch 7.13.4

---

<div class="post-metadata">

**Author:** ![Byron\_H](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/byron_h/32/82245_2.png) [@Byron\_H](https://discuss.elastic.co/u/Byron_H)\
**Post date:** [July 29, 2021, 12:08pm UTC](https://discuss.elastic.co/t/cannot-crawl-a-website/279825/3 "2021-07-29T12:08:06Z")

</div>

Hi Marten, a 599 status code can indicate a network connection timeout error. Can you validate that your server is not taking an undue amount of time when responding to App Search's request to robots.txt?

---

<div class="post-metadata">

**Author:** ![Marten](https://avatars.discourse-cdn.com/v4/letter/m/4af34b/32.png) [@Marten](https://discuss.elastic.co/u/Marten)\
**Post date:** [July 30, 2021, 1:02pm UTC](https://discuss.elastic.co/t/cannot-crawl-a-website/279825/4 "2021-07-30T13:02:52Z")

</div>

Hi Byron,

Please define undue amount.  
If I access robots.txt from my browser, I get a response immediately.  
In the logging of the crawler I see that the log entry with the status code 599 has exactly the same @timestamp as the log entry with the crawler start, meaning that the the crawler doesn't seem to wait very long for the robots.txt.

Marten

---

<div class="post-metadata">

**Author:** ![Marten](https://avatars.discourse-cdn.com/v4/letter/m/4af34b/32.png) [@Marten](https://discuss.elastic.co/u/Marten)\
**Post date:** [August 2, 2021, 7:41am UTC](https://discuss.elastic.co/t/cannot-crawl-a-website/279825/5 "2021-08-02T07:41:53Z")

</div>

I will transfer this issue to the support portal, this one can be closed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 30, 2021, 7:42am UTC](https://discuss.elastic.co/t/cannot-crawl-a-website/279825/6 "2021-08-30T07:42:50Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
