# Web crawler (github.com/elastic/crawler) search behind corporate proxy

**URL:** <https://discuss.elastic.co/t/web-crawler-github-com-elastic-crawler-search-behind-corporate-proxy/375668>\
**Category:** Elastic Search\
**Created:** [March 10, 2025, 7:32pm UTC](https://discuss.elastic.co/t/web-crawler-github-com-elastic-crawler-search-behind-corporate-proxy/375668 "2025-03-10T19:32:42Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vikram\_Tiwari](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikram_tiwari/32/141931_2.png) [@Vikram\_Tiwari](https://discuss.elastic.co/u/Vikram_Tiwari)\
**Post date:** [March 10, 2025, 7:32pm UTC](https://discuss.elastic.co/t/web-crawler-github-com-elastic-crawler-search-behind-corporate-proxy/375668/1 "2025-03-10T19:32:42Z")

</div>

I am trying to run the [GitHub - elastic/crawler](http://github.com/elastic/crawler). It works great for public websites.

To make it work for a customer, we had them remove limits from a proxy server so that we could scrape content from their website. However, I am not sure how to make sure that crawler uses that proxy server URL as the gateway to get to the customer's website.

---

<div class="post-metadata">

**Author:** ![nfeekery](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nfeekery/32/117349_2.png) [@nfeekery](https://discuss.elastic.co/u/nfeekery)\
**Post date:** [March 11, 2025, 10:57am UTC](https://discuss.elastic.co/t/web-crawler-github-com-elastic-crawler-search-behind-corporate-proxy/375668/2 "2025-03-11T10:57:03Z")

</div>

Hi @Vikram_Tiwari

The Crawler has proxy configurations that you can configure. See these example configs: [crawler/config/crawler.yml.example at d3f1bd30eb791a218c62a0c32f06a3c6bbf880e9 · elastic/crawler · GitHub](https://github.com/elastic/crawler/blob/d3f1bd30eb791a218c62a0c32f06a3c6bbf880e9/config/crawler.yml.example#L78-L83)

Can you check if configuring these allows crawling through the proxy server?

---

<div class="post-metadata">

**Author:** ![Vikram\_Tiwari](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vikram_tiwari/32/141931_2.png) [@Vikram\_Tiwari](https://discuss.elastic.co/u/Vikram_Tiwari)\
**Post date:** [March 11, 2025, 5:49pm UTC](https://discuss.elastic.co/t/web-crawler-github-com-elastic-crawler-search-behind-corporate-proxy/375668/3 "2025-03-11T17:49:44Z")

</div>

Awesome! This fixed it. I was expecting it to be at crawler docker level but this is much better.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 8, 2025, 5:49pm UTC](https://discuss.elastic.co/t/web-crawler-github-com-elastic-crawler-search-behind-corporate-proxy/375668/4 "2025-04-08T17:49:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
