Map a site
Use map to list a site’s URLs before you read any of them: the links on the start page and the entries of the sitemaps the site declares, each with where it was found. A map reads one page body at most, so it answers in seconds. You then choose which URLs to extract or batch.
Try it without an account
Hosted Octocrawl maps a site with no key. A map counts as one page of the daily allowance, 20 a day per address.
curl -s https://api.octocrawl.dev/v1/map \
-H 'content-type: application/json' \
-d '{"url": "https://modelcontextprotocol.io", "search": "transports"}'From an agent connected to https://mcp.octocrawl.dev/mcp (see Connect MCP), ask:
Use Octocrawl map on https://modelcontextprotocol.io with search "streamable http" and list the URLs it found.On your own computer, with no daily limit and nothing to sign up for:
npx octocrawl map https://modelcontextprotocol.io --search transports --limit 3The SDKs call the same route: new W2L({ baseUrl }).map(url, options) in TypeScript (@octocrawl/sdk) and W2L(...).map(url, **options) in Python (octocrawl-client, options in snake_case such as include_subdomains).
Expected output
On 2026-10-09 at 12:06 UTC the curl command above returned 11 URLs in 1.2 seconds (hosted revision h2card, source ff6a5a5). Here it is cut to its first two links, with id, url, sources.startPage, the sitemap’s files, identity, warnings, the refusal counts that were 0 and the refusal samples left out:
{
"status": "completed",
"stoppedBy": null,
"links": [
{
"url": "https://modelcontextprotocol.io/community/working-groups/transports",
"via": ["sitemap"],
"sitemapFile": "https://modelcontextprotocol.io/sitemap.xml",
"lastmod": "2026-08-26T03:23:04.413Z",
"robots": "allowed"
},
{
"url": "https://modelcontextprotocol.io/specification/2024-11-05/basic/transports",
"via": ["sitemap"],
"sitemapFile": "https://modelcontextprotocol.io/sitemap.xml",
"lastmod": "2026-10-08T12:58:40.318Z",
"robots": "allowed"
}
],
"sources": {
"sitemap": {
"mode": "include",
"sources": ["robots"],
"listed": 356,
"accepted": 11,
"truncated": null,
"error": null
}
},
"refused": { "duplicate": 18, "hostDenied": 8, "searchFiltered": 347 },
"elapsedMs": 1200
}The site’s own pages may change, so a later run can return other URLs.
linkscome in a fixed order: the start URL, then the start page’s links as the page lists them, then sitemap entries not already found.viasays where a URL was found:start,link(an<a href>on the start page) orsitemap. A URL found both ways lists both.sitemapFileandlastmodname the sitemap that listed the URL and the date it gave, as written there.title, when present, is the start page’s own<title>, a link’s anchor text, or a news sitemap’s title. A map never opens a page to fetch its title.sourcesshows what was read: the start page (status, links found) and each sitemap file (theSitemap:lines of robots.txt, or/sitemap.xmlwhen there are none).refusedcounts what was left out and why, every reason each time, 0 included. Up to 20 examples are kept for URLs folded into a similar one, on other hosts, and disallowed by robots.txt. Here 347 URLs did not matchsearch, 18 repeated a URL already seen, and 8 were on other hosts (among themblog.modelcontextprotocol.io,github.comanddiscord.gg).
Over MCP the answer is compact by default: each URL with its title, and counts of what was returned and refused. Set debug: true for the full map above.
Narrow the list
| Option | What it does |
|---|---|
search |
Keeps URLs whose address or title contains every word, ignoring case. It filters and does not rank. |
limit |
The most URLs to return: 5000 by default and at most on hosted Octocrawl, up to 100000 on your computer. A map that reaches it is completed with stoppedBy: "limit". |
sitemap |
include by default. skip reads only the start page; only reads only the sitemaps. |
includeSubdomains |
Also keeps docs., blog. and other hosts under the start domain. Off by default. |
crawlEntireDomain |
Keeps URLs outside the start URL’s path. Off by default. |
includePaths, excludePaths |
Regular expressions on the path; an exclude wins. |
ignoreQueryParameters |
Treats URLs that differ only in their query as one. Off by default. |
timeout |
Milliseconds for the whole map: 60000 by default and at most on hosted Octocrawl. At the deadline you get what was found. |
A map stays inside the start URL’s path, or the path it redirects to. On the same day, starting at https://modelcontextprotocol.io/docs returned 117 URLs, all under /docs, and counted 248 links and sitemap entries elsewhere on the site as subtreeDenied; starting at https://modelcontextprotocol.io returned 358. Start at the site’s root, or set crawlEntireDomain, to list the whole site.
If the list is short or empty
- The site has no sitemap: a map then lists only the start page’s links. To go further, crawl the site on your computer (
npx octocrawl crawl), which reads each page it finds. start_page_client_rendered: the start page builds its links with JavaScript, which a map does not run. Scrape the page withformats: ["links"]on your computer, or with a key that has the browser lane, where a browser can render it.partialwithmap_timeout: the deadline ended the map. The links found so far are in the answer; raisetimeoutor narrow the map.refused.robots: robots.txt disallows those URLs, and a map does not list them. Hosted Octocrawl always obeys robots.txt. On your computer,ignoreRobotsTxtreturns them with their verdict, except where a rule names Octocrawl itself.failed: nothing was found, and something did not finish: the start page could not be read or robots.txt disallows it (sources.startPage.failureReasonsays which), a sitemap could not be read (a file insources.sitemap.filesmarkedunreadableorrefused, orsources.sitemap.error), too many sitemap files, or the deadline came first.warningsnames each. When the start page fails but a sitemap lists URLs, the map ispartialand keeps them.
See the reference for every route, and limits and result states for the hosted allowances.