# Can Apache Nutch be used with Elasticsearch to index web crawl content?

**URL:** <https://discuss.elastic.co/t/can-apache-nutch-be-used-with-elasticsearch-to-index-web-crawl-content/93>\
**Category:** Elastic Community and Ecosystem\
**Created:** [May 1, 2015, 10:20am UTC](https://discuss.elastic.co/t/can-apache-nutch-be-used-with-elasticsearch-to-index-web-crawl-content/93 "2015-05-01T10:20:56Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![shaunak](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shaunak/32/6643_2.png) [@shaunak](https://discuss.elastic.co/u/shaunak)\
**Post date:** [May 1, 2015, 10:20am UTC](https://discuss.elastic.co/t/can-apache-nutch-be-used-with-elasticsearch-to-index-web-crawl-content/93/1 "2015-05-01T10:20:56Z")

</div>

In general, you are free to use any web crawl product to fetch URL content. Apache Nutch is certainly one of the more popular open source web crawl products in the market. In the case of Apache Nutch (starting in Nutch 1.7+), there is an `ElasticSearchWriter` class written by the Apache team for integration with ElasticSearch.

However, **please keep in mind that the Nutch plugin class (`ElasticSearchWriter`) mentioned above is not written, owned, tested or certified by Elasticsearch so it is not an integration module we support**.

If you run into issues with the Nutch Elasticsearch plugin, please [file a ticket with the Apache Nutch team](https://issues.apache.org/jira/browse/Nutch). There are also additional resources such as [Nutch mailing lists](https://nutch.apache.org/mailing_lists.html) available.

Instead of relying on the Nutch plugin, you can optionally write custom code to pull data out of the default Apache Nutch storage and invoke the Elasticsearch API to create the index. This gives you full control of the ingest/import pipeline given that 3rd party plugins may break and may not be updated to work with the latest Elasticsearch versions.

---

<div class="post-metadata">

**Author:** ![rodrigo](https://avatars.discourse-cdn.com/v4/letter/r/4af34b/32.png) [@rodrigo](https://discuss.elastic.co/u/rodrigo)\
**Post date:** [July 15, 2015, 9:42am UTC](https://discuss.elastic.co/t/can-apache-nutch-be-used-with-elasticsearch-to-index-web-crawl-content/93/2 "2015-07-15T09:42:12Z")

</div>

Yes, and it works pretty well once you get all the kinks ironed out in the configuration files. I got a news crawler indexing to ES 1.4 a while ago, had to fight with it for a few hours but I can dig that code if it helps. Did not get it to work with Nutch 2.x though, but there are a few tutorials out there.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:45pm UTC](https://discuss.elastic.co/t/can-apache-nutch-be-used-with-elasticsearch-to-index-web-crawl-content/93/3 "2017-07-06T13:45:11Z")

</div>


