# Index content to elasticsearch cluster from sitemap

**URL:** https://discuss.elastic.co/t/index-content-to-elasticsearch-cluster-from-sitemap/65717
**Category:** Elasticsearch
**Created:** [November 10, 2016, 7:43pm UTC](https://discuss.elastic.co/t/index-content-to-elasticsearch-cluster-from-sitemap/65717 "2016-11-10T19:43:39Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![Srinivasan\_Ramaswamy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/srinivasan_ramaswamy/32/74106_2.png) [@Srinivasan\_Ramaswamy](https://discuss.elastic.co/u/Srinivasan_Ramaswamy)
#### Post date: [November 10, 2016, 7:43pm UTC](https://discuss.elastic.co/t/index-content-to-elasticsearch-cluster-from-sitemap/65717/1 "2016-11-10T19:43:39Z")

</div>

I have bunch of sitemaps with a list of urls and last modified time, that i want to fetch (get the html) and parse (extract title, links, text, etc) the content and finally index it to elasticsearch. In future, I might have to deal with PDF, Doc and other kinds of content present in some of the urls, as well.

So far I looked at Nutch, Scrapy and Storm crawler. I am trying to keep it simple with room for further improvement in future. I would like to go with a solution thats widely adopted and has a good support. Does any one have any recommendation ?

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [December 8, 2016, 7:43pm UTC](https://discuss.elastic.co/t/index-content-to-elasticsearch-cluster-from-sitemap/65717/2 "2016-12-08T19:43:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
