# Suggestions for places to start for a crawler?

**URL:** <https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023>\
**Category:** Elasticsearch\
**Created:** [June 16, 2010, 4:23pm UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023 "2010-06-16T16:23:26Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Joe\_Bowman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joe_bowman/32/3330_2.png) [@Joe\_Bowman](https://discuss.elastic.co/u/Joe_Bowman)\
**Post date:** [June 16, 2010, 4:23pm UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/1 "2010-06-16T16:23:26Z")

</div>

Hello,

I'm working on a hosted search solution, currently using Bing's API as  
a source for providing domain targeted web search results. I'll  
continue to offer this as the free part of a freemium model, but I'm  
now researching how to set up my own crawler and search backend, which  
will eventually give my customers more control over crawling and how  
search results are displayed.

Everything I've read has me liking ElasticSearch for the search  
engine. It fits the infrastructure design I have, which is completely  
horizontally scalable. I'm also using MongoDB for data storage for  
other components right now, so the fact ElasticSearch is JSON oriented  
makes it a good fit as well.

I'm really new to this area of development, just dripping my toes so  
to speak, so I thought I'd ask the list.

My requirements at this point are fairly simple. I need a crawler that  
can honor robots.txt files, and be restricted by domain. I'm not  
interested in creating another internet wide crawler, I want to crawl  
the sites my customers configure me to. I'd like to find a product/  
approach that's proven and stable. My current platform is built  
primarily using python, but while I'm extremely rusty I can probably  
brush up on C if necessary. I'd like to avoid Java unless something  
that can run within the same JVM as ElasticSearch possibly. Really not  
looking for the overhead of additional JVMs, and I'm not much of a  
java developer, but that's a preference not a requirement. I also have  
a strong belief in right tool for the job.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 16, 2010, 7:37pm UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/2 "2010-06-16T19:37:40Z")

</div>

Hi,

So, elasticsearch does not do crawling itself, the crawling job you will  
need to do and index data into elasticsearch. Nutch, as a side note, is one  
such crawler, and I had a chat with some of its developers and they were  
keen on getting Nutch to index data into elasticsearch. I am not familiar  
with other crawlers, but there are probable several others, the integration  
point would be to get the crawled data indexed into elasticsearch.

-shay.banon

On Wed, Jun 16, 2010 at 7:23 PM, Joe Bowman [bowman.joseph@gmail.com](mailto:bowman.joseph@gmail.com) wrote:

> Hello,
> 
> I'm working on a hosted search solution, currently using Bing's API as  
> a source for providing domain targeted web search results. I'll  
> continue to offer this as the free part of a freemium model, but I'm  
> now researching how to set up my own crawler and search backend, which  
> will eventually give my customers more control over crawling and how  
> search results are displayed.
> 
> Everything I've read has me liking Elasticsearch for the search  
> engine. It fits the infrastructure design I have, which is completely  
> horizontally scalable. I'm also using MongoDB for data storage for  
> other components right now, so the fact Elasticsearch is JSON oriented  
> makes it a good fit as well.
> 
> I'm really new to this area of development, just dripping my toes so  
> to speak, so I thought I'd ask the list.
> 
> My requirements at this point are fairly simple. I need a crawler that  
> can honor robots.txt files, and be restricted by domain. I'm not  
> interested in creating another internet wide crawler, I want to crawl  
> the sites my customers configure me to. I'd like to find a product/  
> approach that's proven and stable. My current platform is built  
> primarily using python, but while I'm extremely rusty I can probably  
> brush up on C if necessary. I'd like to avoid Java unless something  
> that can run within the same JVM as Elasticsearch possibly. Really not  
> looking for the overhead of additional JVMs, and I'm not much of a  
> java developer, but that's a preference not a requirement. I also have  
> a strong belief in right tool for the job.

---

<div class="post-metadata">

**Author:** ![Joe\_Bowman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joe_bowman/32/3330_2.png) [@Joe\_Bowman](https://discuss.elastic.co/u/Joe_Bowman)\
**Post date:** [June 17, 2010, 5:03pm UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/3 "2010-06-17T17:03:52Z")

</div>

Sorry if I was unclear, I do understand the integration point, I  
thought though that Elasticsearch does the indexing as the data is fed  
into it? I think that's what you're saying though. I really am looking  
for suggestions about the specific crawler piece. My own research  
everyone points to Nutch, and based off of your post, unless someone  
adds something to change my mind, I believe I will go with that when  
I'm ready to start that process of my product.

Thanks.

On Jun 16, 3:37 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Hi,
> 
> So, elasticsearch does not do crawling itself, the crawling job you will  
> need to do and index data into elasticsearch. Nutch, as a side note, is one  
> such crawler, and I had a chat with some of its developers and they were  
> keen on getting Nutch to index data into elasticsearch. I am not familiar  
> with other crawlers, but there are probable several others, the integration  
> point would be to get the crawled data indexed into elasticsearch.
> 
> -shay.banon
> 
> On Wed, Jun 16, 2010 at 7:23 PM, Joe Bowman [bowman.jos...@gmail.com](mailto:bowman.jos...@gmail.com) wrote:
> 
> > Hello,
> 
> > I'm working on a hosted search solution, currently using Bing's API as  
> > a source for providing domain targeted web search results. I'll  
> > continue to offer this as the free part of a freemium model, but I'm  
> > now researching how to set up my own crawler and search backend, which  
> > will eventually give my customers more control over crawling and how  
> > search results are displayed.
> 
> > Everything I've read has me liking Elasticsearch for the search  
> > engine. It fits the infrastructure design I have, which is completely  
> > horizontally scalable. I'm also using MongoDB for data storage for  
> > other components right now, so the fact Elasticsearch is JSON oriented  
> > makes it a good fit as well.
> 
> > I'm really new to this area of development, just dripping my toes so  
> > to speak, so I thought I'd ask the list.
> 
> > My requirements at this point are fairly simple. I need a crawler that  
> > can honor robots.txt files, and be restricted by domain. I'm not  
> > interested in creating another internet wide crawler, I want to crawl  
> > the sites my customers configure me to. I'd like to find a product/  
> > approach that's proven and stable. My current platform is built  
> > primarily using python, but while I'm extremely rusty I can probably  
> > brush up on C if necessary. I'd like to avoid Java unless something  
> > that can run within the same JVM as Elasticsearch possibly. Really not  
> > looking for the overhead of additional JVMs, and I'm not much of a  
> > java developer, but that's a preference not a requirement. I also have  
> > a strong belief in right tool for the job.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [June 17, 2010, 5:09pm UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/4 "2010-06-17T17:09:25Z")

</div>

On Thu, 2010-06-17 at 10:03 -0700, Joe Bowman wrote:

> Sorry if I was unclear, I do understand the integration point, I  
> thought though that Elasticsearch does the indexing as the data is fed  
> into it? I think that's what you're saying though. I really am looking  
> for suggestions about the specific crawler piece. My own research  
> everyone points to Nutch, and based off of your post, unless someone  
> adds something to change my mind, I believe I will go with that when  
> I'm ready to start that process of my product.

Although you mentioned that you weren't keen on Java, you didn't specify  
which languages you would like to use.

Perl has some very good crawling libraries, which you could throw  
together to create your own crawler. I don't see one that supports  
robots.txt out of the box, but with three lines of code, you could  
change Web::Scraper to do so:

> **[Web::Scraper - Web Scraping Toolkit using HTML and CSS Selectors or XPath...](https://metacpan.org/release/MIYAGAWA/Web-Scraper-0.32/view/lib/Web/Scraper.pm)**
>
> Web Scraping Toolkit using HTML and CSS Selectors or XPath expressions

Perl also has an interface to Elastic Search:

> **[ElasticSearch - An API for communicating with ElasticSearch - metacpan.org](https://metacpan.org/release/DRTECH/ElasticSearch-0.16/view/lib/ElasticSearch.pm)**
>
> An API for communicating with ElasticSearch

So if you're familiar with Perl, then combining the two parts above  
would be quite simple.

Clint

---

<div class="post-metadata">

**Author:** ![venkatnehatha](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/venkatnehatha/32/3205_2.png) [@venkatnehatha](https://discuss.elastic.co/u/venkatnehatha)\
**Post date:** [June 7, 2011, 4:31pm UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/5 "2011-06-07T16:31:37Z")

</div>

Hi Clinton,

I have also same question. Thank you for your replies. Can you suggest PHP, Java based crawlers those suits for Elastic Search integration.

Thanks in advance.

-Nehatha

---

<div class="post-metadata">

**Author:** ![fashionalwallet](https://avatars.discourse-cdn.com/v4/letter/f/839c29/32.png) [@fashionalwallet](https://discuss.elastic.co/u/fashionalwallet)\
**Post date:** [June 10, 2011, 12:30am UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/6 "2011-06-10T00:30:42Z")

</div>

- deleted -

---

<div class="post-metadata">

**Author:** ![ctjmorgan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ctjmorgan/32/3073_2.png) [@ctjmorgan](https://discuss.elastic.co/u/ctjmorgan)\
**Post date:** [November 16, 2011, 9:04pm UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/7 "2011-11-16T21:04:38Z")

</div>

We are looking at using Nutch (Java Crawler) for our efforts. The below source I put together to write data directly from Nutch to Mongodb in the same fashion as the way SolrIndexer works in Nutch.

> **[ctjmorgan/nutch-elasticsearch-indexer](https://github.com/ctjmorgan/nutch-elasticsearch-indexer)**
>
> nutch-elasticsearch-indexer - Allow the indexing of Nutch crawl data directly into elasticsearch. This is similar in nature to that of the SolrIndexer that comes with Nutch which let you index dir...

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:48am UTC](https://discuss.elastic.co/t/suggestions-for-places-to-start-for-a-crawler/3023/8 "2017-07-06T03:48:22Z")

</div>


