# Web Crawler Setup: "Content Verification" for Domain fails

**URL:** <https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [December 28, 2021, 4:00pm UTC](https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055 "2021-12-28T16:00:44Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![b0r1sp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/b0r1sp/32/98555_2.png) [@b0r1sp](https://discuss.elastic.co/u/b0r1sp)\
**Post date:** [December 28, 2021, 4:00pm UTC](https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055/1 "2021-12-28T16:00:44Z")

</div>

Hi,

I'd like to setup a domain for a web crawler - it fails in the last step "Content Verification" with:

"The web server at [http://www.test.de](http://www.test.de) redirected us to a different domain URL ([http://www.test.com/](http://www.test.com/)). If you want to crawl this site, please use [http://www.test.com](http://www.test.com) as the domain name."

Problem is, I need to crawl pages which are located at a subdirectory of [test.de](http://test.de), say [test.de/ressources](http://test.de/ressources) and I can't touch any DNS or server-settings for this domain to change the redirect behaviour of root or "/".

What I need is to setup a domain in App Search independently of it's redirect behaviour, say via a flag or something - is there a setting I overlook in the docs?

Is there another possibility beside using another tool for crawling and uploading json-files?

Thanks for helping.

---

<div class="post-metadata">

**Author:** ![b0r1sp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/b0r1sp/32/98555_2.png) [@b0r1sp](https://discuss.elastic.co/u/b0r1sp)\
**Post date:** [December 28, 2021, 4:14pm UTC](https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055/2 "2021-12-28T16:14:59Z")

</div>

Some more details:

- curl of [http://www.test.de/ressources](http://www.test.de/ressources) succeeds with 200.
- curl of [http://www.test.de](http://www.test.de) redirects with 301 to [http://www.test.com](http://www.test.com).

(Note: I use "[test.de](http://test.de)" for demonstration purposes of the problem, unfortunately I cannot disclose the real domain)

- I use a self hosted stack, version 7.16.2

---

<div class="post-metadata">

**Author:** ![Irina\_Truong](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/irina_truong/32/112040_2.png) [@Irina\_Truong](https://discuss.elastic.co/u/Irina_Truong)\
**Post date:** [January 5, 2022, 11:01pm UTC](https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055/3 "2022-01-05T23:01:58Z")

</div>

This is a little clunky, but you can add your domain via API, to bypass verification. Here is the API documentation:

> **[Web crawler API (beta) reference | Elastic App Search Documentation \[7.16\] |...](https://www.elastic.co/guide/en/app-search/current/web-crawler-api-reference.html#web-crawler-apis-post-crawler-domains)**

Example request to add a domain:

```auto
curl -X POST http://[ENTERPRISE-SEARCH-URL]/api/as/v0/engines/[ENGINE-NAME]/crawler/domains -H "Content-Type: application/json" -d '
{
  "name": "http://www.test.de"
}'

```

Once this is done, you can add the entrypoint in the App Search UI. Alternatively, the entrypoint can be added using API:

```auto
curl -X POST http://[ENTERPRISE-SEARCH-URL]/api/as/v0/engines/[ENGINE-NAME]/crawler/domains/[DOMAIN-ID]/entry_points -H "Content-Type: application/json" -d '
{
  "value": "/ressources/"
}'

```

The value of `DOMAIN-ID` is the `id` returned in the first response, when adding the domain.

---

<div class="post-metadata">

**Author:** ![b0r1sp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/b0r1sp/32/98555_2.png) [@b0r1sp](https://discuss.elastic.co/u/b0r1sp)\
**Post date:** [January 13, 2022, 7:51pm UTC](https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055/4 "2022-01-13T19:51:11Z")

</div>

Thank you 🙂 I'd propose a UI-change to achieve this in a more comfortable way.

---

<div class="post-metadata">

**Author:** ![b0r1sp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/b0r1sp/32/98555_2.png) [@b0r1sp](https://discuss.elastic.co/u/b0r1sp)\
**Post date:** [January 13, 2022, 7:54pm UTC](https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055/5 "2022-01-13T19:54:13Z")

</div>

Furthermore, adding a domain via curl and ui should have the same result - I guess the ui is using a different api?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 10, 2022, 7:55pm UTC](https://discuss.elastic.co/t/web-crawler-setup-content-verification-for-domain-fails/293055/6 "2022-02-10T19:55:15Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
