# Indexing Nutch Results

**URL:** <https://discuss.elastic.co/t/indexing-nutch-results/5241>\
**Category:** Elasticsearch\
**Created:** [August 24, 2011, 5:00pm UTC](https://discuss.elastic.co/t/indexing-nutch-results/5241 "2011-08-24T17:00:29Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Adam\_Estrada](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adam_estrada/32/2704_2.png) [@Adam\_Estrada](https://discuss.elastic.co/u/Adam_Estrada)\
**Post date:** [August 24, 2011, 5:00pm UTC](https://discuss.elastic.co/t/indexing-nutch-results/5241/1 "2011-08-24T17:00:29Z")

</div>

Has anyone been able to do this? I am using Nutch to crawl the web and  
now I would like to store my results in ElasticSearch rather than in  
Solr.

Thoughts?

Adam

---

<div class="post-metadata">

**Author:** ![Tomislav\_Poljak](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomislav_poljak/32/1177_2.png) [@Tomislav\_Poljak](https://discuss.elastic.co/u/Tomislav_Poljak)\
**Post date:** [August 29, 2011, 3:55pm UTC](https://discuss.elastic.co/t/indexing-nutch-results/5241/2 "2011-08-29T15:55:10Z")

</div>

Hi,

2011/8/24 Adam Estrada [estrada.adam@gmail.com](mailto:estrada.adam@gmail.com):

> Has anyone been able to do this? I am using Nutch to crawl the web and  
> now I would like to store my results in Elasticsearch rather than in  
> Solr.
> 
> Thoughts?

I don't think such an integration exists at the moment, but if you  
check Nutch-Solr integration code you can see Nutch-ES integration  
would be very similar. Nutch integrates with Solr through 2  
points/commands (in bin/nutch script): Solr indexing and Solr  
de-duplicatoin.  
...  
elif ["$COMMAND" = "solrindex"] ; then  
CLASS=org.apache.nutch.indexer.solr.SolrIndexerJob  
elif ["$COMMAND" = "solrdedup"] ; then  
CLASS=org.apache.nutch.indexer.solr.SolrDeleteDuplicates  
...

Solr indexing of Nutch crawled content is implemented through  
SolrIndexerJob and if you check the code  
(org.apache.nutch.indexer.solr.SolrIndexerJob) you will see it uses  
SolrJ (actually CommonsHttpSolrServer) to post the data to Solr for  
indexing. So, ElasticSearchIndexerJob needs to be implemented (similar  
to SolrIndexerJob; also extends IndexerJob) where SolrJ code would be  
replaced with ES indexing client code (for example Java API indexing  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/java-api/index_.html))

Other part of Nutch-Solr integration is deduplication  
([http://wiki.apache.org/nutch/bin/nutch\_dedup](http://wiki.apache.org/nutch/bin/nutch_dedup)) where duplicate  
documents are removed from the index based on either the same contents  
(via MD5 hash) or the same URL. Here iteration through documents and  
duplicate deletes (job queries the solr server and removes duplicates)  
implemented with SolrJ needs to replaced with ES Java API's  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/java-api/search.html) and  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/java-api/delete.html)

Hope this helps.

Tomislav

> Adam

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:55am UTC](https://discuss.elastic.co/t/indexing-nutch-results/5241/3 "2017-07-06T03:55:43Z")

</div>


