# Web Crawler Failed HTTP request: Unable to request "\< domain \>" because it resolved to only private/invalid addresses

**URL:** <https://discuss.elastic.co/t/web-crawler-failed-http-request-unable-to-request-domain-because-it-resolved-to-only-private-invalid-addresses/270631>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [April 19, 2021, 10:00pm UTC](https://discuss.elastic.co/t/web-crawler-failed-http-request-unable-to-request-domain-because-it-resolved-to-only-private-invalid-addresses/270631 "2021-04-19T22:00:12Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jerrac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerrac/32/52980_2.png) [@jerrac](https://discuss.elastic.co/u/jerrac)\
**Post date:** [April 19, 2021, 10:00pm UTC](https://discuss.elastic.co/t/web-crawler-failed-http-request-unable-to-request-domain-because-it-resolved-to-only-private-invalid-addresses/270631/1 "2021-04-19T22:00:13Z")

</div>

When I try to run the web crawler against a site we host, it fails with this error:

> Failed HTTP request: Unable to request "\< domain \>" because it resolved to only private/invalid addresses

The site in question would resolve to a 10.n.n.n ip address. Is the crawler configured to reject that? Is there a way to override that behavior?

---

<div class="post-metadata">

**Author:** ![jerrac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerrac/32/52980_2.png) [@jerrac](https://discuss.elastic.co/u/jerrac)\
**Post date:** [April 19, 2021, 10:12pm UTC](https://discuss.elastic.co/t/web-crawler-failed-http-request-unable-to-request-domain-because-it-resolved-to-only-private-invalid-addresses/270631/2 "2021-04-19T22:12:35Z")

</div>

Not sure it's related, but if I target my personal site, not hosted internally, it fails as well.

In the logs I see:

> Allow none because robots.txt responded with status 599

and

> Failed HTTP request: Remote host terminated the handshake

That also happens if I target [Elastic Blog: Stories, Tutorials, Releases | Elastic Blog](http://elastic.co/blog).

I double checked my personal site's robot.txt file. It's the default Drupal 8 robots.txt file. So there shouldn't be anything in it that would completely block the crawler.

Anyway, I'm glad this still beta. 🙂

---

<div class="post-metadata">

**Author:** ![orhantoy](https://avatars.discourse-cdn.com/v4/letter/o/c5a1d2/32.png) [@orhantoy](https://discuss.elastic.co/u/orhantoy)\
**Post date:** [April 20, 2021, 9:02am UTC](https://discuss.elastic.co/t/web-crawler-failed-http-request-unable-to-request-domain-because-it-resolved-to-only-private-invalid-addresses/270631/3 "2021-04-20T09:02:08Z")

</div>

> [@jerrac](#):
>
> Is the crawler configured to reject that? Is there a way to override that behavior?

Yes, that's the current, default behavior and it will become configurable in the next minor release.

As for the other issue you're experiencing, it sounds like you can't crawl any site at all, is that correct?

---

<div class="post-metadata">

**Author:** ![jerrac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jerrac/32/52980_2.png) [@jerrac](https://discuss.elastic.co/u/jerrac)\
**Post date:** [April 20, 2021, 3:29pm UTC](https://discuss.elastic.co/t/web-crawler-failed-http-request-unable-to-request-domain-because-it-resolved-to-only-private-invalid-addresses/270631/4 "2021-04-20T15:29:56Z")

</div>

> [@orhantoy](#):
>
> Yes, that's the current, default behavior and it will become configurable in the next minor release.

Nice.

> [@orhantoy](#):
>
> As for the other issue you're experiencing, it sounds like you can't crawl any site at all, is that correct?

Yep, can't crawl my personal site, or [Elastic Blog: Stories, Tutorials, Releases | Elastic Blog](http://Elastic.co/blog). Haven't tried any other sites yet.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 18, 2021, 3:30pm UTC](https://discuss.elastic.co/t/web-crawler-failed-http-request-unable-to-request-domain-because-it-resolved-to-only-private-invalid-addresses/270631/5 "2021-05-18T15:30:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
