# Use web crawler beta app search behind corporate proxy

**URL:** <https://discuss.elastic.co/t/use-web-crawler-beta-app-search-behind-corporate-proxy/279599>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [July 26, 2021, 7:39am UTC](https://discuss.elastic.co/t/use-web-crawler-beta-app-search-behind-corporate-proxy/279599 "2021-07-26T07:39:15Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![markn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/markn/32/74417_2.png) [@markn](https://discuss.elastic.co/u/markn)\
**Post date:** [July 26, 2021, 7:39am UTC](https://discuss.elastic.co/t/use-web-crawler-beta-app-search-behind-corporate-proxy/279599/1 "2021-07-26T07:39:15Z")

</div>

Hi,

we run ECE on premise in our data center. We have deployed an Enterprise Search App Search engine. We want to use the web crawler functionality. But the connection towards public internet pages runs via a corporate proxy.

I could not find this in the documentation nor on this forum.

Error message in logs: "Allow none because robots.txt responded with status 599".

How can we configure the web crawler to use a proxy for internet connectivity?  
Thanks.

Kind regards,  
Mark

---

<div class="post-metadata">

**Author:** ![markn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/markn/32/74417_2.png) [@markn](https://discuss.elastic.co/u/markn)\
**Post date:** [July 28, 2021, 3:46pm UTC](https://discuss.elastic.co/t/use-web-crawler-beta-app-search-behind-corporate-proxy/279599/2 "2021-07-28T15:46:14Z")

</div>

Hi, maybe I did not explain it well enough.  
We want to crawl websites like [https://www.unive.nl](https://www.unive.nl)  
But to reach that website we have to go via the corporate proxy in our datacenter.  
I think the web crawler is not aware of that proxy and tries to resolve www.unive.nl directly and that will fail in our datacenter.

My question is: is it possible to configure the web crawler so it uses the proxy to go outside the datacenter?

Thanks.

Cheers,  
Mark

---

<div class="post-metadata">

**Author:** ![oleksiy-elastic](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleksiy-elastic/32/49074_2.png) [@oleksiy-elastic](https://discuss.elastic.co/u/oleksiy-elastic)\
**Post date:** [July 29, 2021, 1:44pm UTC](https://discuss.elastic.co/t/use-web-crawler-beta-app-search-behind-corporate-proxy/279599/3 "2021-07-29T13:44:16Z")

</div>

Hello Mark,

Thank you for trying out the Crawler for your project! Unfortunately, there is no support for running behind a proxy yet. I'll add it to our roadmap since I suspect there will be more potential customers who may need this kind of mode of operation for the product.

In the meantime, there are some options, but they are completely unsupported by Elastic: If you run Enterprise Search outside of ECE (you need more control over the environment around the product to apply those kinds of solutions), you may be able to coerce it to using the proxy by applying socksify or a similar OS-level TCP connection routing mechanism: [socksify(1) - Linux man page](https://linux.die.net/man/1/socksify). Alternatively, there is transparent proxying support in many proxy servers (see [https://wiki.squid-cache.org/Features/Tproxy4](https://wiki.squid-cache.org/Features/Tproxy4) for example), but it requires some really deep understanding of linux firewalls, etc to implement.

I hope this helps.

---

<div class="post-metadata">

**Author:** ![markn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/markn/32/74417_2.png) [@markn](https://discuss.elastic.co/u/markn)\
**Post date:** [August 2, 2021, 6:07am UTC](https://discuss.elastic.co/t/use-web-crawler-beta-app-search-behind-corporate-proxy/279599/4 "2021-08-02T06:07:57Z")

</div>

Thanks for putting it on the roadmap!

In general I think when Elastic should support proxy by default in all their products. Lot of enterprise companies will have to run behind a proxy. But that's just my humble opinion.

Thanks for the alternatives, unfortunately those are a no go for us.

Kind regards,  
Mark

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 30, 2021, 6:08am UTC](https://discuss.elastic.co/t/use-web-crawler-beta-app-search-behind-corporate-proxy/279599/5 "2021-08-30T06:08:36Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
