# Swifttype Crawl Rate

**URL:** <https://discuss.elastic.co/t/swifttype-crawl-rate/169709>\
**Category:** Elastic Search\
**Tags:** elastic-site-search\
**Created:** [February 24, 2019, 9:04am UTC](https://discuss.elastic.co/t/swifttype-crawl-rate/169709 "2019-02-24T09:04:08Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![legislat.io](https://avatars.discourse-cdn.com/v4/letter/l/f9ae1b/32.png) [@legislat.io](https://discuss.elastic.co/u/legislat.io)\
**Post date:** [February 24, 2019, 9:04am UTC](https://discuss.elastic.co/t/swifttype-crawl-rate/169709/1 "2019-02-24T09:04:08Z")

</div>

I'm trialing the service to see if it will help a project we're working on. but I don't seem to get the crawler working across the site.  
I fed it the seed and a sitemap, but it only crawls 20 pages in 12 hours, vs the many thousand pages that exist.  
Is there a better way to get the crawler working? or crawl with another tool then upload a url list?

---

<div class="post-metadata">

**Author:** ![goodroot](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/32/38389_2.png) [@goodroot](https://discuss.elastic.co/u/goodroot)\
**Post date:** [February 24, 2019, 5:56pm UTC](https://discuss.elastic.co/t/swifttype-crawl-rate/169709/2 "2019-02-24T17:56:08Z")

</div>

Hello, Chris!

I took a look at the various sitemaps for [https://legislat.io](https://legislat.io). It looks as though there are ~10 pages listed within the sitemaps.

We have a [documentation page](https://swiftype.com/documentation/site-search/crawler-troubleshooting) that may help.

The key item within that page is the concept of discovery. The crawler follows links within pages, unless directed otherwise, and in doing so "crawls/discovers" your pages.

The sitemaps that your robots.txt file references look like so:

1. [https://legislat.io/sitemap-1.xml](https://legislat.io/sitemap-1.xml), 3 URLs
2. [https://legislat.io/image-sitemap-1.xml](https://legislat.io/image-sitemap-1.xml), 4 images.
3. [https://legislat.io/news-sitemap.xml](https://legislat.io/news-sitemap.xml), 3 URLs.

If, like the crawler, we follow the links within the items in your sitemap, we aren't left with many pages.

I would look at two potential action items:

1. Update the sitemaps so that they are comprehensive.
2. Verify that the website's structure is "hierarchical" in nature, and that links are available as part of a "discovery tree", akin to a "path through your content".

Hopefully this is helpful.

Enjoy the rest of the weekend!

Kellen

---

<div class="post-metadata">

**Author:** ![legislat.io](https://avatars.discourse-cdn.com/v4/letter/l/f9ae1b/32.png) [@legislat.io](https://discuss.elastic.co/u/legislat.io)\
**Post date:** [February 24, 2019, 6:45pm UTC](https://discuss.elastic.co/t/swifttype-crawl-rate/169709/3 "2019-02-24T18:45:13Z")

</div>

Thanks Kellen,  
The site we're indexing is [http://www.legislation.gov.uk/ukpga](http://www.legislation.gov.uk/ukpga)  
The .Io is what we're building out, very much work in progress.

For some reason it just gets the links on that page and stops. Is there a way to feed in a list of links?  
Chris

---

<div class="post-metadata">

**Author:** ![goodroot](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/32/38389_2.png) [@goodroot](https://discuss.elastic.co/u/goodroot)\
**Post date:** [February 24, 2019, 6:59pm UTC](https://discuss.elastic.co/t/swifttype-crawl-rate/169709/4 "2019-02-24T18:59:01Z")

</div>

Hey Chris --

Ah, I see - lots more links on that page. 😁

There are a couple ways you could proceed...

1. Add URLs [via the dashboard](https://swiftype.com/documentation/site-search/crawler-troubleshooting#add-url-button).

2. Add URLs [via the API](https://swiftype.com/documentation/site-search/api-crawler-operations#create_domain).

3. Reformat the sitemap. The URL you linked has a page called "sitemap", but it isn't technically a sitemap that adheres to the [sitemap XML format](https://www.sitemaps.org/protocol.html). Discovery, as linked in the above reply, is still relevant here. The crawler is likely having troubles discovering the other pages.

Keep me posted!

Kellen

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 24, 2019, 6:59pm UTC](https://discuss.elastic.co/t/swifttype-crawl-rate/169709/5 "2019-03-24T18:59:01Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
