Changelog
What changed in Octocrawl, newest first. Unreleased is on main and on this site, and not yet in a published package; each numbered version is on npm and PyPI.
Unreleased
- octocrawl.dev has one top navigation on every page (
apps/public-web/scripts/siteNav.mjs): Product (what Octocrawl does, the ways to use it, and connecting an agent over MCP in one line), Docs (the guide and reference pages), Free tiers and Changelog, with GitHub and Try it beside them. The menus are<details>, so they open without a script;docs-assets/nav.jsadds hover, one menu at a time, Escape and, on narrow screens, a Menu button. A menu prints open in the site’s glyph style: the card appears a text line at a time under a dotted print head, its entries in turn, and its monospaced words decode from noise; with reduced motion it opens at once. The new Changelog page (/changelog/) is built fromCHANGELOG.md, its links to repository files pointing at GitHub, and is in the sitemap and llms.txt;.gcloudignoreletsCHANGELOG.mdinto the Cloud Build upload. - A page with little text (at most 4,000 characters) that still shows its data is on the way (a visible element marked
aria-busy, a visible short “Loading…” or “Fetching results…” text, not on a button or a link) is waited for up to 8 s in all on the browser rung and on a provider’s, instead of being read after the usual settle of at most 1.5 s; a request withwaitFororactionskeeps its own wait instead. The trace records the wait (loading_wait), and a page read as content while it still showed the sign carries apage_still_loadingwarning. A provider page read that way in the PA 4 Steel run had answeredsuccesswith “Inventory Search Results Fetching…” for its content. - The published packages point at the site: npm
octocrawl,@octocrawl/cli,@octocrawl/sdkand@octocrawl/mcpgethttps://octocrawl.devas their homepage, the GitHub issues asbugsand search keywords; PyPIoctocrawl-clientgets the site, the docs and the issues as project URLs and keywords; each package README names the site, and@octocrawl/mcp’s README offers the hosted server (https://mcp.octocrawl.dev/mcp) for trying it without installing. They take effect with the next release. - octocrawl.dev, from the 2026-10-09 site check: the home page carries its FAQ as FAQPage data, built from the same list the FAQ shows; every sitemap address has a
lastmod(the day its content last changed, kept inapps/public-web/scripts/docsPages.mjsand checked against git by a test); the hashed scripts and stylesheets, and files asked for with their?v=content version, are cached for a year (immutable) instead of an hour; the Codex and Cursor MCP links on Connect MCP point at those docs’ current addresses; the docs introduction has its own heading instead of the home page’s; the header’s Docs link no longer carries the outside-link arrow; and the home page description fits in 155 characters. - The ladder goes on to its next rung for more failures a stronger rung may answer (ROADMAP PA item 4): a 403 or 405 answered without a gate it recognises, a page a browser rendered with no main content it could verify, and the HTTP rung’s refused connection (its next rung is the browser, whose own network failure ends the run);
ladder_stepwithescalatesays which. When the stronger rungs fail without a page, the page stepped past is the answer (ladder_evidence_kept). On the PA 9 gap tasks such failures had ended the run before the browser or the provider was asked. A timeout, a rate limit (429), an error status that is the page’s own answer (404, 410, 5xx), a failure of the saved login’s rung and a failure after content an earlier rung found still end it. - The home page shows what Octocrawl does beyond the one-page preview, how to start and what is free, in three sections between How it works and the FAQ. What it does: six cells say what runs on hosted Octocrawl (scrape, map, the Evidence Record) and what runs only on your computer (crawl and batch, Monitors, your own Chrome), each in two or three short points, with a filter; a cell pointed at fills with orange from the left and its words turn white. Get started: a Use it from window with CLI, MCP, TypeScript, Python and REST tabs (the CLI, TypeScript and Python lines were run as written on 2026-10-09; the MCP command is the one Connect MCP marks verified on 2026-10-06), under a heading where a glyph octopus drawn after the brand mark lives in a small sea: it swims, rests, sleeps, changes colour, waves, peeks at the heading, blows bubbles, reads a page that drifts away as Markdown, chases or greets a passing fish, squirts ink and flees from a pointer that comes close, comes to a click, and keeps clear of the words. Free tiers: the four allowances the services enforce (five previews a visitor a day, 20 pages a day per address without a key, 1,000 to start with one, no daily limit on your own computer) pass one at a time over the glyph Earth artwork (
scene-earth.webp), the scroll settling on a tier when it comes to rest; each tier lights that many of the planet’s own painted marks (5, 20, 1,000, or every mark on the planet), found by a worker the way the hero finds its glyphs. The glyph band under the hero rolls with two slow waves. Motion runs only while on screen; with reduced motion everything shows still, and without a script every code line and every tier still shows. - Every paid provider call is on the page’s Evidence Record (ROADMAP PA item 4):
access.paidCallslists, in order, the provider, the rung, the ADR 0005 capabilities its session was created with, the price ceiling reserved, what the spend ledger charged, the price the provider stated (null when none), and what Octocrawl made of the page the call returned by its own checks (outcome,reason; null when it returned none), never the provider’s word, withanswermarking the call whose page is the record’s.access.grantnames the access grant they were made under by the SHA-256 of its text, its tier and its attestation time. A provider called with no budget is settled in an uncapped ledger of the run’s own, so it is recorded too, and a page keeps the calls of a read that did not become its answer: one given up for another egress, the run whose stopped page the person read in their Chrome, and a batch or crawl run that threw after a paid call. Both fields are optional in the v1 schema, so earlier records stay valid. - Provider spend under a ledger (ROADMAP PA item 4): an access grant’s
tariffsgive each provider’s prices per call and per hour and what bounds a call’s time (maxSessionMs,minBilledMs,billingIncrementMs), from which its price ceiling is computed; a provider without one is not called, and a per-GB price is refused (bandwidth cannot be bounded from here). A task’s pages, retries and providers reserve each call’s ceiling from one ledger before the call and settle it after (at the reported price, or at the ceiling when none was reported, a call that threw or was cut included), so concurrent workers cannot passperRunUsd, and a resumed or appended run opens the ledger with what the task was already charged (stored on each attempt aschargedUsd);perRequestUsdnow caps one page’s calls and a single scrape’s. Under a tariff a provider call opens a connection and a session of its own, created with the provider’s own timeout and its proxies off, and released when the call ends. A run stops at its cap (cost) instead of at the first unpriced provider page (cost_unknown). Answers addusage.externalCostChargedUsd(the run’s total) and aspend_settledtrace event per call;externalCostUsdstays the exact cost or null. The ladder CLI takes the grant’s tariffs too. - The my-browser lane says more exactly why a page was not read: a page still on a site’s check when the wait ends is
blockedwith that check (a Cloudflare block page answered 403 had beenfailed/timeout), and a Chrome that refuses a command (a tab it will not open) isfailed/connection_error, nottimeout. A page the site leads elsewhere on it each time Octocrawl takes the tab back (www.linkedin.com/mynetwork/ in the 2026-10-09 acceptance run) isfailed/redirect_limit, nottimeout. Its warning says the person got through only when the page showed a check and they acted in its tab, not when the check cleared by itself. A page read in the person’s Chrome counts ashanded_to_persononly when it showed a check and they acted in its tab; a check that cleared by itself isuser_browser. - The main content of a Wikipedia article leaves out its hidden categories (
.mw-hidden-catlinks, the maintenance categories the site’s stylesheet hides from every reader), which ended the Markdown after the visible categories; those stay, and so does the whole page’s copy withonlyMainContent: false(#289, the other half of the Wikipedia main-content gap). - The main content of a page leaves out the navigation a site lays out inside its content column: an element with
role="navigation", and MediaWiki’s portlets and menus (.mw-portlet,.vector-menu, the language menu#p-lang-btn), which Wikipedia’s Vector 2022 skin puts in<main>beside the heading. A scrape of a Wikipedia article withonlyMainContent(the default) now opens with the article, where it opened with the interlanguage list (“22 languages” and a link per language) and the page tabs; the “See also” portal box goes too.onlyMainContent: falseandincludeTagskeep them, as they keep<nav>(the Wikipedia main-content gap seen on parity case A16). - A list continued in the person’s Chrome counts its items again:
actions.lists[].itemsis the last page’s count anditemsReadthe kept pages’ count plus each page the person showed (the reader’s tab counts the step’sitemSelectoron every page), where before both werenullafter a continuation.itemsReadis a sum only when the addresses show no page can be counted twice (the kept pages each at its own, the check’s page and the pages shown at none of them, the check not at the list’s own address) and no page shows again what another shows (its items’ whole text, as the list merge tells it, since a result set tied to the session that made it comes back at new addresses), and staysnullotherwise (also when the step’sitemSelectoruses a form the extractor does not take, such as:has(), so pages cannot be compared), as for a pager that reloads its items in place at a redirected address. The counts are not added toactions.scrapes. Found by the real-site run on indeed.com (2026-10-09): pages 2 to 5 went through, and the record could not count them. - A page handed to the person whose own script rewrites its address once it has come is still the page asked for: the handoff reader hears the tab’s navigations (
Page.frameNavigated,Page.navigatedWithinDocument) and takes an address the page set in place by itself, on the document that came at the page asked for, before the person clicked or typed on that document (its user activation, asked of Chrome in Octocrawl’s own world when the address changes), when it keeps the path and the value of every parameter both name, and drops no parameter whose value is a number (a page or an offset). Found on indeed.com (ROADMAP PA item 3’s real-site run, 2026-10-09): its list page drops its paging tokenppand adds the job shown asvjk, so the reader took the real page for one elsewhere, took the tab back twice and gave up. A page parameter that changed (another page of a list), or an address the person’s own click moved in place (Next, a sort), is still not the page: the tab is taken back. - Where a pool egress leaves from (ROADMAP PA item 3, “every result records … the location”): with
W2L_EGRESS_ECHO_URLbesideW2L_EGRESS_PROXIES, each proxy is asked the echo URL through itself (once, again after it failed or after 10 minutes), and the Evidence Record’saccess.egressgainsexit: { ip, country, observedAt }on every page read through it (anegress_exittrace event; the country only when the service gives a two-letter one). Without the echo URL, when it did not answer, or for the environment proxy,exitisnull. Optional in the v1 schema file. The server refuses the echo URL without egress proxies or when it is not http(s). - A list that stopped at a check goes on in the person’s own Chrome (ROADMAP PA item 3): a batch whose only step is
paginatewith anitemSelectoris handed over for the items whose list stopped at a check; the page the check was on opens in their Chrome, they get through it and page on by clicking Next themselves, Octocrawl only reads that tab (a read-only script; it clicks nothing), and the pages its own browser read before the check are merged with those into one list, each page once (actions.scrapes[].bysays who read a page; thelist_continuedtrace counts the kept pages and names the lane that read them). The item becomes the whole list withactions.lists[].continued = { from, pages, by: "user_browser" }and thelist_continuedtrace event;stoppedBysays how the reading ended (endonce Next has been unusable on the last page read for 5 s,max, ordeadlineafter 60 s without a new page or at the handoff’swaitMs, withlist_not_exhausted); a person who got through but showed no page after the check’s within that time leaves the item stopped, for a later handoff to go on from. A page off the site, on a login path, with a password field or not answered 2xx is not read, as for any page read in their Chrome. The CLI prompt and the MCPhand_off_batchtext say so. Any other batch with steps is still not handed over. - A paginate step stops at a check the site puts up where the next page should be (ROADMAP PA item 3): each page is put to the gate classifier’s decisive marks as it is read, and an interstitial (Cloudflare, a PerimeterX press-and-hold, a verification form) ends the step as
challengewithout being read, told or counted; a page of nothing but a CAPTCHA widget, which only the extractor tells from a page with little on it, is marked the same way by the page’s own verdict after the steps (a page with records on it, or with other content, is a page whatever widget it carries). The pages before it keep their records inlist, the result isblockedwith the check’s reason,actions.lists[].challengenames the page and the URL, thelist_challengetrace event says what the gate saw, andlist_not_exhaustedsays the list stopped at a check. In a batch the pages before it stay in the task’s checkpoint (going on from them once the check is handled is still to come); a page with no record on it is not kept there. Before, the check’s page was read as an empty page of the list and the checkpoint was cleared. - With egress proxies, a page whose request never went out (robots.txt or a policy refused it, a lockdown found no cached copy) no longer has its egress probed: only a page that failed on the network with no HTTP answer from any rung does. The probe reaches no site, but it was a CONNECT to the proxy for nothing (PR #247’s follow-up).
- A task’s cookie session file and egress binding are removed before its terminal event goes out, so a webhook receiver or an events stream told of the end finds them gone; before, they went after the event (PR #242’s follow-up).
- A list task cut at page N resumes there (ROADMAP PA item 3): a batch with a
paginatestep keeps each page the step reads in the task’s checkpoint the moment it is read, and when the batch resumes on the next start it passes over those pages along the site’s own Next links (pages have no address of their own in a click-driven list), reads the pages after them and merges every page once; a kept page is known again by its address, its items, or their links alone when a price or date changed meanwhile (rows with no links whose text changed are read again, repeat, and count twice towardmaxPages);actions.lists[].resumedand thelist_resumedtrace event say how many pages came from the checkpoint. The lane tells each page throughExecutionContext.onListPageand takes them back throughlistResume. A paginate step that fails at page N (a click that never lands) now keeps the records of the pages it read inlist, with the failure inactions.failed. - The Evidence Record’s
accessblock names the egress a page left through and the task session it was read with (ROADMAP PA item 3):egressis{ proxy, source, switchedFrom }(the proxy’shost:port, never its credentials;pool,environmentordirecton a request that got a page response; the pool egress the task last moved off before the page was read, elsenull),nullwhen nothing says: a vendor’s service, the person’s own browser, or a lane that stopped before any page request;sessionis{ id }, never the cookies,nullwhen none. Both are optional in the v1 schema file, so earlier records stay valid. - The browser-compatible transport has a default host list (ROADMAP PA item 2): a local server whose access grant names
compatible_transportand that sets noW2L_COMPAT_HOSTSuses it for the five hosts G1’s two-window acceptance showed it helps (research/access/benefit-hosts.v1.json: fred.stlouisfed.org, www.idealo.de, www.investing.com, www.ironmountain.com, www.wayfair.com).W2L_COMPAT_HOSTS=noneturns it off; naming hosts replaces the list. The list is a constant in the code, kept equal to the JSON record by a test, so every build has it. - The README, the Chrome connection hint, the CLI prompt and the MCP
scrapetool say that while Chrome’s remote debugging is on (which the handoff,octocrawl login importand the my-browser lane need), every page seesnavigator.webdriverastrue, Octocrawl connected or not (seen on Chrome 153 with the chrome://inspect switch and on 154 with --remote-debugging-port), so a bot check may refuse the person’s Chrome; turn it off when done. - The my-browser lane waits 10 minutes, not 120 s, for the person to allow the sites in Chrome; a refusal says what Octocrawl’s page last answered; a batch waiting for the approval reports
waitingForApproval: true; and the server logs each step of the approval (my_browser_approvalon stderr). Three real-page runs had timed out at this step with nothing to tell a slow click from one not recognised. - One plain choice of how pages are reached (ROADMAP PA item 7):
"access": "standard"(no rung that costs a third party),"enhanced"(what the server’s access grant of tier enhanced approves, its providers in mode standard too; refused by name without one) or"my-browser"(the person’s own Chrome, aslane: "my-browser") on scrape and batch,standardorenhancedon crawl; in the MCPscrape,batch_scrapeandcrawltools and--accesson the CLI. Batches and crawls keep the choice. The/fcshim maps Firecrawl’sproxyonto it (basicis standard;stealthandautoare enhanced) instead of refusing it. - The
my-browserlane for batches (ROADMAP PA item 8):"lane": "my-browser"onPOST /v1/batches, the MCPbatch_scrapetool andoctocrawl batch --lane my-browserread every page in the person’s own Chrome, one at a time. The person allows every site of the batch (host and port) once for the run, in the page Octocrawl opens there; a page on another site, or after Revoke, is not read. Never cached. Refused with a webhook and withmaxConcurrencyabove 1, besides what a scrape on the lane refuses. - The
my-browserlane (ROADMAP PA item 8), for one page:"lane": "my-browser"onPOST /v1/scrape, the SDK, the MCPscrapetool andoctocrawl scrape --lane my-browserreads the page in the person’s own Chrome on a server on their machine. After Chrome’s Allow, Octocrawl opens a page of its own there listing the site and the task; only the person’s click on Allow reading these sites lets it read that site without a further click, and closing that page or clicking Revoke stops it. Recorded as lanemy_browser(a new value oflane), never cached. Refused by name on other servers and with actions, a screenshot, lockdown or a mode other than standard. Every Evidence Record’saccessgainscompletion(unattended,authorized_session,user_browser,handed_to_person, or null when no page was read). Batches and MCPbatch_scrapefollow. - Managed sessions (
/v1/sessions/*) no longer answer with the profile’s path on the server (profileDir) or a CDP endpoint (cdpEndpoint); every route returns the public fields only. A hosted engine refuses them all (409), in the engine itself as well as at the hosted gate, and makes no browser profile (ROADMAP PA, G4).
0.3.1 — 2026-10-06
The published packages (octocrawl, @octocrawl/cli, @octocrawl/sdk, @octocrawl/mcp, octocrawl-client) at 0.3.1: everything below since 0.3.0 on 2026-10-05, and @octocrawl/mcp now carries mcpName for the official MCP Registry.
-
The published
@octocrawl/mcpcarriesmcpName: io.github.77777R7/octocrawland the repository hasserver.jsonfor the official MCP Registry (registry.modelcontextprotocol.io): the npm package over stdio, withW2L_API_URLandW2L_API_TOKEN, and the hosted remotehttps://mcp.octocrawl.dev/mcp.scripts/release-version.mjskeepsserver.json’s versions equal to the packages’. Publishing to the registry takes a release that carriesmcpName(0.3.1 or later), thenmcp-publisher login githubandmcp-publisher publish. -
docs/hosted-api.md: the key and waitlist scripts need
NODE_USE_ENV_PROXY=1behind a proxy, since Node’sfetchignoresHTTPS_PROXY; and where to keep a freshly issued key. -
Egress proxies (ADR 0005
egress_sessions, ROADMAP PA item 3):W2L_EGRESS_PROXIESnames the operator’s own http(s) proxies. A batch or crawl keeps one for its run (kept inegress.jsonin its directory, so a resumed task goes on through it) and moves to the next healthy one only when the proxy itself fails (after a page got no HTTP answer, a probe finds the proxy does not answer or refuses its credentials), at most twice a run, with a new cookie session; the page is read again there and its trace says so (egress_switched). A block, a challenge, a 429 or a connection the site reset never moves it. Cookie session files are kept per route, so a task resumed on another route starts a new session; modeauthednever uses the pool. A failed proxy cools down for 10 minutes; scrapes and maps take the next healthy one. Each proxy reads robots.txt and sitemaps for its own pages; per-host pacing stays shared. Needs the grant; refused on a hosted server; credentials are never recorded. -
The public site and the README say hosted Octocrawl is live (PH phase 1, step 5). Connect MCP leads with
https://mcp.octocrawl.dev/mcp: one picker for the hosted URL (Claude Code, Cursor and OpenCode verified on 2026-10-06; Codex from its docs), then how to add a key (--headerin Claude Code,headersin Cursor and OpenCode,--bearer-token-env-varin Codex), the first task, what the hosted service does not do, then “Run it on your computer” with a second picker for the stdio server, and self-hosting. The home page’s Get code panel is “Use it in your code or agent”: the cURL listing callsapi.octocrawl.dev, the MCP and prompt steps name the hosted URL first; the FAQ, the quota advice and the waitlist band (now “Ask for a hosted Octocrawl key”) follow. Limits gains a hosted allowance table, Privacy what the hosted service records, Terms and the acceptable-use policy cover it. The README’s MCP section is hosted, on your computer, self-hosted. llms.txt and every docs footer say the same. -
Hosted Octocrawl: the health route is
GET /health./healthzon a Cloud Run URL is answered by Google’s front end with its own 404 page and never reaches the container (seen on the first deploy, 2026-10-06); the runbook’s check uses the new path. -
Hosted Octocrawl, phase 1, the deployment (ROADMAP PH):
Dockerfile.hosted-apiandcloudbuild.hosted-api.yamlbuild theoctocrawl-apiimage (Chromium included, task root in memory and swept);cloudflare/hosted-api-proxy/answers onapi.octocrawl.devandmcp.octocrawl.devand forwards to Cloud Run with the shared secret, redirecting only plain http;scripts/hosted/issue-key.mjsissues, lists and revokes keys (a key is printed once, its HMAC is the Firestore document id);docs/hosted-api.mdis the runbook: deploy without traffic, move traffic under a tag, the Firestore TTL policy, the Worker, the outside checks, the rollback, and the two-week cost record. -
Hosted Octocrawl, phase 1 (ROADMAP PH), the service:
npm run hosted:api(packages/mcp/src/hostedApiCli.ts) runs the API’s hosted mode and a remote MCP endpoint at/mcpin one process. It servesPOST /v1/scrape,POST /v1/mapand the MCP toolsscrape,mapandscrape_product; every other route and tool is refused by name with a hint to run Octocrawl locally. A caller withAuthorization: Bearer <key>is looked up by the key’s HMAC in Firestore (hostedApiKeys/{digest}: enabled, plan, dailyLimit, browser); a caller without one is keyless, counted by the address the Cloudflare Worker reports (proven by the shared secret), withinW2L_KEYLESS_DAILYpages a day (20) on the HTTP lane alone (fastModeis pinned; a screenshot is refused by name). Every start consumes the caller’s and the service’s (W2L_SITE_DAILY, 1500) daily counters in one Firestore commit, the public preview’s pattern; over the allowance the answer is 429quota_exhaustedwithRetry-Afterto 00:00 UTC; per minute, 10 starts keyless and 60 with a key. Every request is also limited per address (120 a minute) before any key lookup, the limiter and key cache remember at most 20,000 callers, a file a call reads is capped at 5 MiB, nothing is stored for reuse (storeInCacheis pinned off) and scrape and map records and saved files are swept from the task root after 10 minutes, since Cloud Run keeps it in memory.W2L_HOSTED_STORE=memoryruns it without Firestore for a local check. The deployment (Dockerfile, Cloud Build, the Worker routes, the key script) is the next change. -
A batch’s or crawl’s cookie session (
egress_sessions) now survives a restart: it is written tocookie-session.jsonin the task’s directory after every change (created 0600, replaced whole) and read back when the task resumes, with the same session id; the file is deleted when the task completes, fails or is cancelled. A run paused by shutdown keeps it. -
The public site no longer contradicts the published packages (PH phase 1, step 1): the Get code panel, the Extract page guide and the reference say the API and MCP server come from
npx octocrawl serveandnpx -y @octocrawl/mcp, not from a repository checkout; the reference names@octocrawl/sdkandoctocrawl-clientinstead of calling the SDK a private workspace package; the introduction drops its source commits and “local machine” narration and names the four MCP clients; the footer of every docs page, the client picker, llms.txt and the Connect MCP page say a hosted URL is coming (with the early-access link) instead of “hosted MCP is paused”; the Monitor guide says it runs from a checkout and that a hosted Monitor is a later phase. The early-access form takesconnect-mcpas a trigger (/?from=connect-mcp#waitlist). -
ROADMAP: hosted Octocrawl restarted as phase PH (decided 2026-10-06). The hosted API and remote MCP leave the Paused table: phase 1 serves scrape and map at
api.octocrawl.devandmcp.octocrawl.devin hosted mode, keyless within a per-IP allowance and metered by credits with a key, in Firecrawl’s shape; batch, crawl, Monitor and PA’s access routes are phase 2. The local path stays free and complete. P4 prices hosted credit packs first; “Free core and Pro” says so. The Paused row becomes “A hosted browser cluster beyond Cloud Run’s instance cap” (hosted_browser_clusterin@w2l/http-coreand ADR 0005 follow it). The research behind it isdocs/launch/2026-10-06-hosted-feasibility.md. -
Cookie sessions for batches and crawls (ADR 0005
egress_sessions, ROADMAP PA item 3): under a grant that names it, the cookies a task’s pages set are sent again to their site on its later pages, by the HTTP rung (each redirect hop included), the compatible transport and the browser, which starts from the session’s cookies and leaves its own there. Matching follows RFC 6265 (tough-cookie 6.0.2, a new dependency). The session lives in memory for one run of the task; values are never recorded (session_cookiesnames a random session id and counts), and no page read with a session is cached. Without the grant nothing changes. -
The README links to octocrawl.dev at the top (the preview, the documentation and Connect MCP, each with
utm_source=githubso the site’s page events can tell these visits apart), its Quick Start no longer says the preview has no permanent URL, and “For local MCP use” starts from the published packages (npx octocrawl serve, thennpx -y @octocrawl/mcpin the client); the checkout’s managed local service stays as the Monitor → HTTPS delivery path. -
The public site’s Connect MCP page now starts from the published packages: step 1
npx octocrawl serve, step 2 addnpx -y @octocrawl/mcpto the client, with the Claude Code command, the Cursor and OpenCode configs (each run here and connected) and the Codex command (its documented syntax; not run here). The first task is ascrapeof the first-use page. The repository checkout’s managed local service (127.0.0.1:8791/mcp,w2l-local) moves to a short section at the end. The home page’s Get code panel says the same: the cURL and MCP listings start withnpx octocrawl serve, and the prompt asks for “Octocrawl’s scrape tool” instead ofw2l-local. -
The Evidence Record states how each page was reached: a new
accessfield gives the route (http,http_compat,browser,enhanced_browser,authed_browser,user_browser,vendor), the client that sent the requests and its version when known (undici,impit0.14.5,playwright,patchright, the person’s browser, the vendor’s id), the HTTP transport’s browser profile, and the run’s third-party spend (0when none was paid,nullwhen a vendor stated no price). It is read from the page’s own trace, so batch items and crawl pages carry it too; a page served from the cache states the route and cost of the fetch it reuses. A result no lane produced (a rung the deadline cut, that failed, or whose identity was refused; a cache-only miss) has anullroute and client rather than a guess. It is optional in the v1 schema file, so records written before it stay valid. Each ladder attempt insummary.attempts(full responses) also carries its place in the run (ordinal), when its rung was asked and answered (startedAt,endedAt), and a provider rung’svendorId. -
The public site’s
automatedflag on page events and preview outcomes now also covers monitors that do not call themselves bots (Dataprovider.com, DomainMonitor) and browser strings no one runs any more (iOS before 15, Chrome before 110). In the first week of logging, most page views came from such clients: they opened the page and never touched it, and the funnel counted them as visitors. Current browsers, Chrome on iOS and Edge included, are unchanged. -
robots.txt by who chose the URL (decided 2026-10-05). On a local server a URL the request names (a scrape, a batch entry, the CLI’s URL list, MCP
scrapeandbatch_scrape,/fc/v1/scrape) is fetched when robots.txt disallows it or cannot be read; robots.txt is still read and recorded, with arobots_overriddenwarning and the new Evidence Record fieldrobotsDecision.overrideBasis(user_named_url,robots_override,ignore_robots_txt; optional in the v1 schema file). The links a crawl or map discovers and a Monitor’s re-reads still obey it, and a hosted server obeys it for every URL. A crawl or map takesignoreRobotsTxton a local server (CLI--ignore-robots-txt, MCP, and Firecrawl v2’s name on/fc/v1/crawl): a crawl fetches the pages and sitemap files robots.txt disallows, a map returns those URLs withrobots: "disallowed"or"unreachable". A recordedrobotsOverridenow sets an unreachable robots.txt aside as well. robots.txt is matched with the product tokenOctocrawladded to every User-Agent: a group forOctocrawl(orw2l-research) is the site owner’s targeted opt-out, which a named URL andignoreRobotsTxtdo not set aside; only a recordedrobotsOverridedoes. Crawl-delay pacing and the 429 cooldown are unchanged. A hosted engine refusesrobotsOverride,robotsOverridesandignoreRobotsTxtby name, the hosted MCP engine included. -
The HTTP lane no longer reports a SHA-256 of an empty body when no response arrived (a refused connection):
rawSha256is null. -
Repository cleanup (ROADMAP P0): the earlier product documents (
PHASE1_ENGINEERING_NOTES.md,PRODUCT_PLAN_V2.md,PRODUCT_STRATEGY.md) and the Render Blueprint move from the root todocs/archive/; the hosted MCP pilot’s Render and WorkOS setup moves fromdocs/mcp-first-use.mdtodocs/archive/hosted-mcp-pilot.md, and its code (packages/mcp/src/host.ts,npm run hosted:mcp) is marked experimental, saying so when it starts. The README gains a table of the local ports (8787 API, 8791 local MCP, 8788 first-use webhook receiver, 8798 site preview). -
Releases are published from a tag
vX.Y.Zby theReleaseworkflow, through npm and PyPI trusted publishing (no stored token, npm provenance), after the type check, the tests and the package install check.scripts/release-version.mjssets and checks the one version the published packages share;scripts/check-packages.mjsreads the expected CLI version from its manifest instead of a fixed 0.3.0. -
The README starts from the published packages (
npx octocrawl,pip install octocrawl-client,@octocrawl/sdk,@octocrawl/mcp). A newInstall checkworkflow runsnpx -y octocrawl@latest scrape https://example.comfrom an empty npm cache on macOS, Windows and Linux (scripts/install-check.mjs, limit 5 minutes) andpip install octocrawl-client, each week and on demand. -
Security, before the first publish:
- A local server on an address other than 127.0.0.1, localhost or ::1 (
--host 0.0.0.0, a LAN address,W2L_API_HOST) refuses to start without a token. Before, it answered anyone, other machines and rebinding web pages included, and fetched the person’s localhost and network for them. - Mode
authedrefusesexecuteJavascript(a script could read the session’s cookies and storage) and, on a batch, awebhook(pages read with the session are not sent to another address). A batch with a webhook is no longer handed to the person (its stopped items carry no handoff, andhandOffBatchrefuses it): a page read in their own Chrome is read signed in as them, and was sent to the webhook as a handoffpageevent. Click, write, press, scroll and the list steps stay available. The MCP tools say so, and tell the model to ask the person before using a saved login. - A job id that is not one the server issues (
..%2F…) names nothing: crawl and batch routes answer 404 instead of opening or creating a task store outside the task root. - The loopback server refuses a request body that is not JSON (
content-type: application/json), so a page on another local port cannot send one as a form or text; a POST without a body (curl -X POST …/cancel) still needs no type. - The one-off commands (
octocrawl scrape,batch,crawl,map) ignoreW2L_API_HOST: they listen nowhere. - Paid browser services are used only when named in
W2L_VENDORS(browserbase,steel) with their key; aBROWSERBASE_API_KEYorSTEEL_API_KEYin the shell alone no longer sends pages to them. - A task root the engine creates is readable by the person alone (0700) and holds a
.gitignore, so a repository it sits in does not take in pages and job databases; the folder of the saved logins is 0700 too. - The CLI’s README no longer says every fetch declares its identity: the default mode sends the Chrome user agent and client hints, and
--mode researchnames itself. The research user agent links to the Octocrawl repository.
- A local server on an address other than 127.0.0.1, localhost or ::1 (
-
The packages to be published carry the Octocrawl name:
@octocrawl/cli(also as the unscopedoctocrawl, sonpx octocrawl scrape <url>runs it),@octocrawl/sdk,@octocrawl/mcpand the Python clientoctocrawl-client(import octocrawl_client; the nameoctocrawlon PyPI belongs to another project). The command isoctocrawland the MCP serveroctocrawl-mcp, and the CLI’s messages, the API’s hints and the MCP tool descriptions nameoctocrawlcommands; the MCP server introduces itself asoctocrawl. The workspace keeps its@w2l/*package names, theW2L_*environment variables and the.w2l/directories. Nothing is published yet. -
A page that is a list of items is now read as its content with the default
onlyMainContent: true, not failed asempty_unverified(EXTRACTOR_VERSIONis nowextract-tf/14). Before, two kinds of list page failed. A list whose items carry no link, such as quotes with their authors, was not a listing of cards to the last-resort fallback, which also wanted a page with exactly one h1. A grid of cards that each declare a schema.org Product in microdata was routed as one product page, whose recommendation pruning then cut the cards. The extractor now takes the list thelistformat’s detection finds when no other region was found: at least three of its items must hold 40 characters of text each, all different, and the region widens to the last h1 before them. Such a page is still checked for a wall as one with nothing found (lastResorton the extractor’s output), so a login form with a list beside it staysblocked, and a handoff in the person’s Chrome keeps waiting at a captcha with a list beside it. Three or more outermost Product scopes of one tag and the same classes now route as a collection, unless the page declares a Product in JSON-LD, has a buy box or a lone h1 with a price shown outside the cards, or shows the cards as recommendations (an h1 inside one, a recommendation heading before them, or an element around them named for recommendations). Found by the real-site cases AC01, AC05, AC07, AC08 and AC09, which ran their steps and then failed on the read. -
A rendered answer (the browser or a provider lane’s) that its own extraction found thin and unsure (300 main-content tokens or fewer, confidence 0.3 or less) now carries the
low_content_yieldwarning, with an agent hint on what to pass (waitFor,actions,onlyMainContent: false); its status stands. The warning was the http lane’s alone, so IMF’s datamapper, whose figures are drawn by script, came back assuccesswith 226 tokens of social and navigation links and no caveat. The 300 is set from the browser lane’s yield on the real-site set (record): of its 73 rendered pages, only a tag page whose extraction took its sidebar falls under it. -
A step cursor this API did not issue (made up or cut short) on
GET /v1/batches/:id/errors,/v1/batches/:id/items,/v1/crawl/:id/pagesor/v1/crawl/:id/errorsis now HTTP 400invalid_request“cursor is not one this API issued”, as on the crawl status routes; it was a 500internal_error. -
A crawl’s or a map’s sitemap reader now keeps the Crawl-delay a host’s robots.txt declares for the request’s identity between its own requests to that host: a sitemap index’s files came the policy’s 250 ms apart. The delay counts from the reader’s last request to the host and is waited before it takes its turn on the origin scheduler, so it neither slows another job’s requests to the host nor waits for a gap in them; the pages after the files are paced by the crawl’s frontier as before. On www.cbs.nl (
Crawl-delay: 1) its index and first file are now 1.0 s apart, 0.25 to 0.29 s before (record). -
Lists and quotes are indented 32 levels deep at most, so the Markdown of a list nested thousands deep grows with its text, not with the square of its depth: each level indented every line inside it, and 2,000 nested
<ol><li>agave 6 million characters. The blocks of a list item or quote deeper than 32 levels are written as blocks of the 32nd level’s item or quote, with no marker or>of their own and a blank line between each two, so none runs into another (a line into a table below it, a line before---into a heading); their text is kept, in order. The same 2,000 levels give 194,000 characters, and twice the depth gives twice the Markdown. Rendered with markdown-it, random pages nested 33 to 70 levels deep show every word of the page in order, with no code block, heading, table or cell the page does not have (2,122 of 2,122). Up to 32 levels of the Markdown’s own nesting nothing changes: the same Markdown and tables for 30,000 random inputs, 3,860 random pages nested up to 32 levels and 120 locally captured pages (a list written directly in a list is one level deeper in the Markdown, inside the item before it).EXTRACTOR_VERSIONis nowextract-tf/13. -
Deep lists and quotes are written with less memory: instead of a prefix kept for each line at each level (the entry below), a block is kept as the tree of its items and quotes and written out once, each line getting the prefixes of the items and quotes it is in as it is written. Peak memory for 2,000 nested levels, less the 110 MB of an idle process:
<ol><li>a103 MB (276 MB as a string rewritten at each level, 203 MB with a prefix per line), a quote of a paragraph 170 MB (359 and 385 MB), a quote of a list item 315 MB (446 and 697 MB), list items of three paragraphs 241 MB (362 and 520 MB); the time stays as it was with the prefixes. The Markdown and tables do not change: the same for 120,000 random inputs and 120 locally captured pages, soEXTRACTOR_VERSIONstays. -
Lists and quotes nested deep are written in time that grows with their Markdown, not faster: each list item and quote wrote the text of everything inside it again to indent it, so the time grew with the cube of the depth while the Markdown (indented at each level) grows with its square. A block now keeps its lines, and a list item or quote adds its indent, marker or
>to each line, written out once at the end. 2,000 nested<ol><li>atook 1.8 s and take 0.1 s; 2,000 quotes of a list item, 6.5 s and 0.4 s. The Markdown and tables do not change: the same for 120,000 random inputs of nested lists, quotes, code, tables and line-start characters, and for 120 locally captured pages, soEXTRACTOR_VERSIONstays. Converting the 120 pages takes as long as before (4,310 and 4,302 ms, median of five). -
Blocks nested thousands deep (
<div>, lists, quotes, layout tables, an emphasis or<span>around blocks,<pre>content) no longer run the Markdown converter out of stack: 2,000 nested lists or layout tables, 4,000 quotes or<b><div>, or 8,000<div>s threwRangeError: Maximum call stack size exceeded, which failed the page. The block walk keeps its place on a stack of its own, each level with what it writes once its children are (a list item’s marker, a quote’s>, a closing paragraph break), and the text of a<pre>is read the same way. A table’s own rows and cells are now found by looking through it once, not by searching all of its descendants and dropping a nested table’s: a table nested 5,000 deep took 3.1 s and takes 33 ms. The Markdown and tables do not change: the same for 60,000 random inputs of nested blocks, tables, templates, svg and math, and for 120 locally captured pages, soEXTRACTOR_VERSIONstays. Converting the 120 pages takes about as long (3,678 and 3,734 ms, median of five). -
Inline elements nested thousands deep (
<sup>,<sub>,<span>,<i><b>,<code>and the like) no longer run the Markdown converter out of stack: 8,000 nested<sup>or<span>, or 4,000<i><b>, threwRangeError: Maximum call stack size exceeded, which failed the page. The inline walk, the search for a block inside an element and the text of an element with hidden parts now keep their place on a stack of their own; 20,000 levels convert as 3 do. The Markdown does not change: the same for 20,000 random pages, their fragments and 120 locally captured pages, soEXTRACTOR_VERSIONstays. Converting the 120 pages takes about 1% longer. Blocks nested that deep (lists, quotes, tables) still run out of stack. -
Nested emphasis that ends in punctuation right before a letter (
<i><b>"y"</b></i>z) is now written so CommonMark reads it as emphasis: the punctuation goes after the closing markers of both runs (***"y***"z); it was written***"y"***z, which shows its markers as text. An emphasis run’s content that ends with another run passes on how that run is written before a letter, so the outer run moves the same punctuation after its own marker. Checked by rendering the Markdown with markdown-it and comparing each page’s visible text with Chromium’s: on random pages of two nested runs with punctuation at their edges and a letter after them, 0 of four runs of 1,500 differ (319, 117, 309 and 318 before); on random pages of emphasis, punctuation, code and links, 85, 72, 76 and 76 differ (86, 73, 76 and 76 before); no page differs that matched before. The Markdown of 120 locally captured pages does not change.EXTRACTOR_VERSIONis nowextract-tf/12. -
Nested emphasis that starts with punctuation right after a letter (
x<i><b>"y"</b></i>) is now written so CommonMark reads it as emphasis: the markers of both runs make one delimiter run, read by the letter before it, so the punctuation goes before them all (x"***y"***); it was writtenx***"y"***, which shows its markers as text. After a space, or at the start of a link’s text, nothing moves. Checked by rendering the Markdown with markdown-it and comparing each page’s visible text with Chromium’s on random pages of nested emphasis, punctuation, code and links: 103, 96, 95 and 106 of four runs of 1,500 differ (116, 109, 108 and 119 before); no page differs that matched before. The Markdown of 120 locally captured pages does not change.EXTRACTOR_VERSIONis nowextract-tf/11. -
The Markdown converter writes emphasis, escapes and adjacent runs so CommonMark reads them as the page shows them (
EXTRACTOR_VERSIONis nowextract-tf/10):- A
<b>,<strong>,<em>or<i>that holds blocks keeps its emphasis on each paragraph in it, as a browser shows it; it was dropped. - White space at the edges of emphasis goes outside the markers, a full-width space included, and a backslash at the end of emphasised text is escaped, so the markers are read as emphasis (
**indent**, not** indent**). - Emphasis next to punctuation keeps the punctuation outside the markers where CommonMark would not otherwise read them as emphasis (
a"**x**"b); a<,&or&#so moved next to what follows is escaped, as it would make a tag or an entity with it. - A backslash that would escape the
]or)closing a link or image is escaped ([C:\\](…)). - Text is escaped only where CommonMark would read it as Markdown:
*,_and`that could open emphasis or code,[and]that could make a link,<that could start a tag, a backslash before punctuation, and what would start a heading, list, quote, rule or table at a line’s start (1.alone on a line gives1\., which CommonMark would read as an empty list item). - Adjacent runs of one emphasis that are plain text are joined (
<b>a</b><b>b</b>gives**ab**), in linear time, as is code; runs holding their own Markdown or a character that pairs across the join (<b><i>x</i></b><b><i>y</i></b>,<b><</b><b>span></b>) are written side by side, the second with underscores.
These six fixes were written for the htmlparser2 parser and are carried over onto parse5 unchanged. Checked by rendering the Markdown with markdown-it and comparing, for every character that is not white space, its bold, italic, code and link with Chromium’s on random pages of emphasis, code, links, blocks and characters Markdown reads (with raw HTML read as HTML): 257, 279, 262 and 262 of four runs of 1,500 differ (779, 780, 782 and 784 before); no page differs that matched before. Of 120 locally captured pages, the Markdown of 55 changes (49 for their main content); their tables do not.
- A
-
EXTRACTOR_VERSIONisextract-tf/9: the entries down to the table span parsing below landed onmaintogether, afterextract-tf/8, so each version they name is that one step; a Monitor whose Markdown changes only through them recordsextraction_reprocessedonce. -
In a template’s content, a table’s text still pending when the input ends or a
</template>closes the template is written after that token is done, as in Chromium, which queues the write: an option the token closes is copied into its select’s<selectedcontent>without the text, which then goes where the table’s rules put it (fostered out of the table, or kept in it when all whitespace).<template><select><button><selectedcontent></selectedcontent></button><option selected>w8<table>w10</template>gives a<selectedcontent>ofw8and the table, where it also heldw10. At the end of the input the text is written once no template is left open, so an option around the templates, at a page’s or a fragment’s own level, still has it when it closes; with any other token the copy has the text, as before. Template content is not read by the Markdown, links, images, attributes or main-content selection, so only thehtmlformat shows this andEXTRACTOR_VERSIONstaysextract-tf/9. Against Chromium, with each input given its own time limit, random inputs of selects, tables, templates and formatting elements now match wherever Chromium gives a result (the one input per 5,000 that differed now matches), with no regressions; the 120 captured pages build the same trees. -
Every option a copy into a
<selectedcontent>holds is handled in turn as inserted, as Chromium handles each node of an insertion: one the copy of an earlier one already took out of the tree is still selected and copied, once, if it has aselectedattribute, but is then not the select’s (never selected by default, and not copied into a<selectedcontent>inserted later); while it is copied the select has no selection of its own, so an option its copy holds is selected by default if it is the first enabled one left. Before, such an option was skipped, so a list box whose selected option holds selected options and<selectedcontent>elements kept the copy of the first of them where Chromium goes on to the last:<select size=3><button><selectedcontent></selectedcontent></button><option selected>A<span><option selected>B<button><selectedcontent></selectedcontent></button></option><option selected>Q</option></span></option></select>now showsQ. Against Chromium, with each input given its own time limit (some hang Chromium), random select inputs now differ only where they did before (a<selectedcontent>holding a table’s fostered text, 1 in 5,000 in one run), and hand-written cases of taken-out copies with and withoutselected, disabled first options and later<selectedcontent>elements match; the 120 captured pages build the same trees.EXTRACTOR_VERSIONis nowextract-tf/9. -
An option’s or
<selectedcontent>'s select is found up its ancestors, not the stack of open elements, so one fostered out of a table, or moved by the adoption agency, is read where it is, as in Chromium. An option fostered out of a table in a<selectedcontent>is the select’s even once a copy took the table out of the tree:<select><selectedcontent>w<table><option><option><option>leaves the<selectedcontent>empty, each option copied in turn taking out what was there, where the later options were left in it. Against Chromium, random inputs of the selects sets (options, optgroups, datalists, tables, templates and<selectedcontent>written freely, and realistic customizable selects) now differ in at most 1 of 5,000 per run, where 0 or 1 did before, with no input that matched before differing now; what still differs differed before too (a<selectedcontent>holding a table’s fostered text, a list box with several<selectedcontent>, where Chromium’s own results disagree). The 120 captured pages build the same trees.EXTRACTOR_VERSIONis nowextract-tf/9. -
An option that holds a select, or a
<selectedcontent>with options, is copied into its select’s<selectedcontent>elements as any other, where it was left alone; the options the copy holds are then the select’s, by the rules a parsed option follows. In the document, a copied option with aselectedattribute is selected and copied in turn, which takes it out again, as in Chromium (<select size=3><button><selectedcontent></selectedcontent></button><option selected>A<selectedcontent><option selected></option></selectedcontent></option></select>leaves the first<selectedcontent>empty); in a template’s content and in a fragment the option is only copied. Each such copy is of an option nested deeper, so this always ends, at the deepest selected copy; a chain deeper than 100 is parsed by linkedom instead, and each step up the tree it takes counts against the parser’s budget. Chromium itself gives no result for some of these: its DOMParser hangs on such an option in a select that is not a list box, and its renderer crashes on some list boxes with several<selectedcontent>. Against Chromium, random inputs that write options and<selectedcontent>freely into selects now differ in 0 or 1 of 5,000 per run (0 to 6 before), with no input that matched before differing now; realistic selects still match in every case, and the 120 captured pages build the same trees.EXTRACTOR_VERSIONis nowextract-tf/9. -
An option written in a select’s
<selectedcontent>is read as Chromium reads it, where it was left alone. It is one of the select’s options, and a copy into the<selectedcontent>that holds it takes it out of the select. In a page, Chromium copies an option as soon as it is selected, while still empty, so<select><selectedcontent><option>A</option>z</selectedcontent></select>gives<selectedcontent>z</selectedcontent>; when the selected option is taken out, the first enabled option left is selected without a copy (a<selectedcontent>written later takes it), and parts taken out of the tree are no longer the select’s. In a template’s content and in a fragment the copy is made only when the option closes, so the same input givesAz, and the next option is then selected. An option in an<optgroup>in another<optgroup>is not one of the select’s, a select in another select (reached through a table) copies nothing, and a<selectedcontent>in a<datalist>is still the select’s, as in Chromium. Against Chromium, random inputs that write options, optgroups and<selectedcontent>freely into selects now differ in 0 to 7 of 5,000 per run (40 to 45 before), with no input that matched before differing now; the realistic set matches in every case (one<selectedcontent>in a<datalist>per 5,000 differed before in one run), and the 120 captured pages build the same trees. Each visit to a<selectedcontent>and each element walked in an option counts against the parser’s budget, so a page that would make a select re-copy many times falls back to linkedom.EXTRACTOR_VERSIONis nowextract-tf/9, asincludeTagsselections and microdata see the changed copies. -
A
<select>'s<selectedcontent>elements hold a copy of its selected option, as in Chromium (the standard’s customizable select): the option with theselectedattribute (the last one), or else the first that is not disabled (none in a list box,sizeabove 1; no copies in amultipleselect). The copy replaces what the<selectedcontent>held and is made when the option closes (also at the end of the page); in a page, a<selectedcontent>written after the selected option takes it when it is inserted, in a template’s content and in a fragment it does not. Every copied node counts against the parser’s element budget (comments now count too), so a page that would copy a large option into many<selectedcontent>elements is parsed by linkedom instead. An option written in a<selectedcontent>or in another option, and an option that holds a select or a<selectedcontent>holding an option, are left as they were: Chromium reads them through a chain of removals and resets this does not follow. The main content’s Markdown is unchanged (a select is left out of it), and links and images drop the repeats, butincludeTagsmatches the copies as a browser’s DOM would (includeTags: ['.price']over a select’s options now also gets the selected option’s price from its<selectedcontent>), and so do product facts read from microdata, theattributesand thehtmlformats. Against Chromium, random pages and fragments of selects with<button><selectedcontent></selectedcontent></button>, selected and disabled options, optgroups, datalists, templates, links, images,multipleandsizenow match in every case (28 to 38 of 5,000 per run differed before), and inputs that write options into<selectedcontent>differ no more than before; the 120 captured pages give the same output.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
<select>is read by the HTML Standard’s rules since July 2025 (whatwg/html#10548), as Chromium has read it since version 134. The “in select” modes are gone: a select and its options hold what is written in them (<div>,<b>,<p>, tables, svg, links, images), where they kept only the text and<option>/<optgroup>/<hr>. A<select>bounds every scope, so a<p>or<li>written in it does not close one outside;</select>closes what is open in it; an<input>or a second<select>still ends it, a<textarea>or<keygen>no longer does. Markdown skips the select itself, but what follows it can change (<p>a <select><b>x</select>b</p>isa **b**, nota b), and links and images in a select’s options are now collected. Against Chromium, random pages and fragments of select, option, form-control, table, template and formatting tags now match in every case (1,480 to 1,530 of 5,000 per run differed before); the 120 captured pages give the same output.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
An end tag in svg or math is matched to an open element by its exact name, as in Chromium. In svg it first takes svg’s spelling (
</foreignObject>,</clipPath>,</linearGradient>,</textPath>…); in math it keeps its own. An end tag that meets an HTML element first is read by the HTML rules, where a name svg respelled matches nothing, so it closes nothing. The standard (and parse5) compares names lowercased, so</foreignObject>in an svg closed an HTML<foreignobject>(or<clippath>…) or math’s around it, and what followed left the svg or math, where Chromium keeps it: in<p>a<lineargradient><svg></lineargradient>b</p>thebis in the svg, which shows no text, so the Markdown isa, notab. Two more places where parse5 also took an svg or math element for an HTML one now read HTML elements only, as the standard and Chromium do: an end tag written in HTML inside svg or math (a</desc>or</mi>closed the svg<desc>or math<mi>around it, and what followed left it), and the insertion mode chosen after a table, cell or select closes (an svg<tfoot>made the rest be read as a row group’s, so a second<table>was dropped). Against Chromium, random pages of table, template, svg and math tags (withforeignObject,clipPathandlinearGradientin either case) now match in every case (36 to 50 of 5,001 per run differed before), and random fragments in every case but one written as<body>…</body>, which is read as a selection; the 120 captured pages give the same output.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
A
<title>,<base>,<basefont>,<bgsound>or<noframes>at a template’s top level switches it to the body’s rules, as in Chromium, which reads every start tag there by the head’s rules only for<link>,<meta>,<script>,<style>and<template>. The standard reads these five by the head’s rules too, which left the template’s mode as it was, so rows, cells and columns after them were kept where Chromium drops them and keeps their text. A fragment (main content, a selection) is read as a template’s content, so its Markdown changes when its top level has such a tag before rows:<title>Rows</title><tr><td>a</td><td>b</td></tr>isab, notaandbas paragraphs. Against Chromium, random fragments and pages of table, template and head tags now match in every case (788 to 850 of 5,000 fragments and 236 to 303 of 5,001 pages per run differed before); the 120 captured pages give the same output.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
A fragment (main content, a selection) that starts with a
<col>ends as in Chromium’stemplate.innerHTML: a table’s text still pending at its end, in a<template>in it, is dropped, where the standard writes it.<col><template><table>abcis<col><template><table></table></template>; the formatting elements the text re-opens before the table are still re-opened. Such text is always in a nested template, which no output reads, so no output changes andEXTRACTOR_VERSIONstaysextract-tf/9. Against Chromium’stemplate.innerHTML, random fragments of table and template tags now match in all 50,000 of ten runs (1 to 6 of 5,000 per run differed before). -
A
<form>in a<template>is read as Chromium reads it, where its parser differs from the standard (and parse5). A<form>written in a table inside a template is kept where it is written and closed at once: Chromium drops such a form only when a form is open outside any template, the standard whenever a template is open. A</form>in a template closes its form as any other end tag closes its element, so not past a<p>,<div>,<li>or other special element still open in it, where the standard closes those first:<template><form><p></form>xkeepsxin the paragraph. Neither becomes the page’s form. Against Chromium, random pages of template, table and form tags now match in every case (11 to 18 of 1,500 per run differed before); with svg and math as well, what still differs is Chromium’s</foreignObject>past an HTML<foreignObject>. The Markdown, links, images and attributes leave template content out, so no output changes andEXTRACTOR_VERSIONstaysextract-tf/9. -
A table’s end tags in a
<template>that is in a table stay in the template, as in a browser. parse5’s table scope stopped only at<table>and<html>, where the standard’s stops at<template>too, so a</table>,</tr>or row group end tag in such a template closed the cells, rows, groups and table outside it, and the template’s text became page text:<table><tr><td>a</td><template><td>hidden</td></table>x</template><td>b</td></tr><tr><td>c</td><td>d</td></tr></table>gave a one-cell table followed byxbcdinstead of the tablea | b,c | d. Against Chromium, random pages of table and template tags now match in every case but two kinds parse5 and the standard read alike: a<form>in a table in a template, which Chromium keeps, and Chromium’s newer<select>. The 120 captured pages have no<template>and give the same output.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
The trees of the last two pages parse5 built are kept until the next task, not one. The browser lane reads two versions of each page by turns, the page as rendered (main-content selection, the Markdown,
tables) and its body as received (links, images, attributes), and each read pushed the other’s tree out: withonlyMainContent: falseit built the rendered page three times and the body twice. Run as the browser lane runs them, the 120 locally captured pages, withonlyMainContent: false,imagesandtables, now build 234 trees instead of 576 (about two per page, one for each version) and take 15.5 to 16.6 s (23.6 to 25.8 s before), with the same output; with the main content only, each version was already built once (234 trees both times). A loop over pages keeps two trees at most. -
The tree parse5 builds for a page is kept until the next task, so a page read several times in one request (main-content selection, the Markdown,
tables, links, images) is parsed once and copied into a new linkedom document each time. Main-content selection, both Markdowns,tables, links and images of the 120 locally captured pages together take 17.1 to 17.5 s again (25.6 to 30.3 s with a parse each, 17.2 s before parse5), with the same output. Only one tree is held, and only until the next task. One page’s own parse still takes about 1.7 times what linkedom’s took: that is parse5’s tokenizer and the garbage it makes; building the linkedom nodes straight from parse5 (a tree adapter) instead of copying was not faster (3.7 s against 3.6 s for the 120 pages), as creating linkedom nodes costs the same either way. -
Pages are parsed by the HTML standard’s tree construction: parse5 builds the tree, which is copied node by node into linkedom (
dom.tsparse, used by main-content selection, links, images, attributes, adapters, the Markdown andtables). htmlparser2, linkedom’s own parser, built misnested markup otherwise than a browser, which the table tag pass and rebuild of the entries above patched for tables only; both are removed. Now formatting elements are reopened as a browser reopens them (<b>1<p>2</b>3</p>gives**1**and**2**3), a page written without<html>or<body>, or with content after</body>, keeps all of it, and a<tbody>opens where a browser opens one. A fragment, such as the main content, is read as a<template>'s content, so one row or cell of a layout table stays one; a selection given as<body>…</body>is read without its<body>tag, whose attributes are kept, soincludeTags: ["tr.athing"]keeps its rows. A<noscript>'s content, text to a browser running scripts, is kept as its elements, so the images and links in it are still the page’s. A page whose reopened formatting elements would outgrow its tags (four elements, or twenty element-name reads, per<: the standard reopens every<b>a block closed in each block after it, so 3,000 differently attributed<b>before 3,000 paragraphs, 56 KB, would make 9 million elements and run out of memory) is parsed by linkedom instead; the 120 locally captured pages use at most 12% of that budget. Three changes to parse5 8.0.1 (pinned): in a row it closed the row at a</tbody>,</tfoot>or</thead>whose group is not open, where the standard and Chromium ignore the tag; it moved a node’s children one by one with a linear search for each (a page of 200,000 lines without a doctype, or 80,000 under a misnested<b>, took seconds), and now moves them together; and a tag of more than 256 attributes (pages have at most 47; parse5 and linkedom check each against those before it, so 100,000 took half a minute) sends the page to linkedom. Checked against Chromium’s tree on random whole pages of misnested tags: every word has the ancestors Chromium gives it in all of 10,500 pages of 7 runs (formatting elements, tables, svg and math, templates, raw-text elements, pages without a doctype, a bare doctype, or no<html>), where 365 to 1,130 per run of 1,500 differed before; the table oracles of the entries above match as before. Still different: Chromium’s newer<select>parsing (content in a select, which the Markdown skips),</foreignObject>in svg past an HTML<foreignObject>(Chromium ignores it, the standard closes it), and misnested markup in a<template>in a table (parse5 differs from the standard there; 12 of 3,000 random sequences). On the 114 HTML pages of 120 locally captured ones the Markdown, tables, main content’s Markdown, links and images are unchanged (the other 6 are gzip data saved as.html, whose NUL bytes are now dropped); thehtmlformat gains the<tbody>a browser opens (42 pages) and keeps what follows</body>in the body (109 pages). Parsing takes about 1.7 times as long (3.1 s against 1.8 s for the 120 pages, 102 MB), selection plus both Markdowns about 1.35 to 1.5 times, and extraction (whole and with a selection), the three Markdowns, tables, links, images and thehtmlformats together 1.7 times (46.1 s against 27.1 s).parse58.0.1 is a dependency of@w2l/extract-tf, andhtmlparser2no longer is.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
svg and math in tables are read as a browser reads them where the tag pass still differed. Integration points are now per namespace: svg’s
<foreignObject>,<desc>and<title>and math’s<mi>,<mo>,<mn>,<ms>and<mtext>, so a<td>in math’s<foreignObject>or svg’s<mi>is theirs, not a cell. Closing an svg or math writes out the end tag of every svg and math element open in it, innermost first, so htmlparser2 does not close an outer<math>for an inner one, nor a cell for an svg’s own<td>; such end tags never carry the pass’s internal names (it wrote</^foreignobject>, which htmlparser2 kept as a comment). htmlparser2’s implied closes inside svg and math (a<tr>there closes a<tr>before it) and its view of a self-closing slash (it ignores one under an element namedmi,title, … in any namespace) are followed, writing out the end tag where it would keep the element open. An end tag in svg that closes no svg element, of a name svg writes in camel case (</foreignObject>,</clipPath>, …), is ignored, as Chromium ignores it. And an end tag of an element other than a block, list item, heading,<p>,<form>or formatting element no longer passes a special element (<div>,<p>,<li>, …) in a table:<td><span><div>x</span>y</div>keepsxytogether, as a browser does. Checked against Chromium on random tag sequences in a table, written as whole pages: every table matches in 2 runs of 2,000 with svg, math,<foreignObject>and<mi>(13 and 12 differed before), 4 runs of 2,000 adding<g>,<mo>,<mtext>and<desc>(7, 6, 6 and 4 before), 3 runs of 1,000 of loose svg and math tags (26, 28 and 25 before) and 3 runs of 2,000 with self-closing svg and math elements (0, 1 and 1 still differ, as they reopen<b>; 51, 52 and 48 before). Still different: an svg<title>,<style>or<script>holding tags (htmlparser2’s tokenizer reads its content as text, as for HTML’s; 1 of 2,000 such sequences), formatting elements a browser reopens after a misnested end tag (its adoption agency), and an HTML<title/>in a body (a browser reads the rest of the page as its text). 120 locally captured pages need no edit and give the same output; 20,000 random sloppy tables give the same output as before.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
A whole page (one with an
<html>tag or a doctype) is parsed with its table tags outside any table ignored, as a browser ignores them: a stray<tr>,<td>,<th>, row group, caption or column after a table ended, or on a page without tables, is dropped and its content stays where it is written. Before, htmlparser2 kept them, so<table>…</table><tr><td>x</td><td>y</td></tr>gave the paragraphsxandywhere a browser showsxy. Inside a<template>, svg or math they stay, and a fragment keeps them (given to the converter, or the main content thehtmlformat returns), as the main content of a layout table can be one of its rows or cells. A</template>for a template opened before a table it left open now closes the template and the table; before, it was dropped and the template held the rest of the page (<template><table><tr><td>x</template>followed by a table gave an empty Markdown). Checked against Chromium’s text of whole pages with a table followed by random stray table tags, text,<span>,<b>,<p>and tables: the text and its breaks match in all of 8,000 pages (183, 191, 195 and 324 of four runs of 2,000 differed before). 120 locally captured pages need no edit and give the same output; 20,000 random sloppy tables (fragments) give the same output as before.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
Pages are parsed with their table tags read as a browser reads them (
dom.tsparse, so main-content selection, links and images as well as the Markdown andtables). htmlparser2 applies an end tag to the nearest open element of its name wherever it is, so a cell’s own</td>after a<td>written in it, or a</div>or</body>written in a table, closed cells, rows and whole tables further out:<table><tr><td>a</td><td>b</td></tr></body><tr><td>c</td><td>d</td></tr></table><p>after</p>lost the second row and the paragraph after the table. Before linkedom parses, htmlparser2’s own tokenizer now reads the tags and follows the stack of elements a browser has open from the outermost table in (with htmlparser2’s implied closes such as<p>closing a<p>, and svg and math content read as a browser reads it: a self-closing<svg/>closes, an HTML element such as<div>or an end tag that closes the cell ends the svg), and the source is edited only where the two differ: an end tag a browser ignores there is dropped, and so is an end tag with space after</, a comment to a browser; the end tags of the cells, rows, row groups and captions a browser closes at a table tag are written out, with the<tr>it opens for a cell written directly in a row group (htmlparser2 closed the<thead>there) and a<tbody>where rows directly in the table follow a closed implied one, so they stay two row groups; a<table>where rows belong ends the table there (inside a<template>a browser ignores it), and the table tags left of it outside any table are dropped; a column group ends at anything but a column. Outside tables nothing changes, and a<td>or<tr>outside any table stays, as main content can be one cell of a layout table. The converter no longer takes an svg’s<tr>or<td>for a row or cell, and moves an svg written at a table’s own level before the table. Checked against Chromium on random tag sequences inside a table, every table matches: 2 runs of 3,000 with cells, rows, row groups, captions,<div>,<span>,<p>and stray</body>(1,267 and 1,257 differed before), 3 runs of 3,000 adding column groups,<b>,<li>,<form>and</html>(643, 654 and 462 before), 2 runs of 3,000 adding<template>(509 and 511 before), and runs of<p>/<li>/<dd>implied closes (2,000; 308 before), end tags with space after</(2,000; 758 before), self-closing<svg/>in cells (400; 71 before), svg icons with<title>and MathML<mi>in cells (400; 79 before) and icons with<title>,<desc>,<use/>,<foreignObject>and MathML (1,000; 275 before) and an svg in an svg’s<foreignObject>(400; 143 before). Still different: 4 and 26 of 3,000 sequences that also close tables (409 and 754 before), as their pages start<!doctype html><div>with no<html>or<body>, so the converter reads only that first element once the table and the<div>have ended (an older converter issue, open separately; written with<html><body>, none differ), and svg and math holding table tags and stray end tags (13 of 2,000, 394 before; 26 of 1,000 sequences of loose svg,<foreignObject>, math and<mi>tags, 244 before; see the svg and math entry above). Random tables with structure written in cells match in all of 1,939 and 1,765 (545 and 824 before). 120 locally captured pages need no edit and give the same output; 20,000 random sloppy tables give the same output as before. The tokenizer pass skips a page without<table>, keeps indexes into its stack so no step walks it, and takes 93 ms on 1.4 MB of 100,000 unclosed cells.htmlparser2(already installed with linkedom) is now a declared dependency of@w2l/extract-tf.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
Tables (Markdown and the
tablesformat) are rebuilt as a browser’s parser builds them where linkedom kept tags where they are written. A<thead>,<tbody>,<tfoot>,<tr>,<td>,<th>,<caption>or<col>inside a cell closes that cell there, also inside a<span>or<div>(htmlparser2 closed a cell only at a<tr>or<td>that is its direct child, so<td>a<th>bnested the<th>in the<td>). An element where a row group, row or cell belongs (a<div>around rows, a caption in a<tbody>, a cell directly in a<tbody>) and text there move as a browser moves them, the text before the table. A<table>there ends the table, and it and what follows come after the table. A<template>stays whole and its rows are not the table’s. Before, such a cell’s text held the rows written in it, and those rows were also rows of the table:<td>a<thead><tr><td>x</td><td>y</td></tr></thead></td>gave the cella x y. Only tables built otherwise are rebuilt: 120 locally captured pages with 298 tables give the same Markdown and tables as before. Checked against Chromium on random tables: with structure written in cells (inside a<span>or<div>where it is a<td>or<tr>), the cells, caption and text before the table match in all of 1,905, 1,939 and 1,765 tables of three runs (1,782, 1,818 and 1,765 differed before); of 20,000 tables written with unclosed cells, comments and<form>wrappers, all 4,718 whose output changed now match. A<td>or<tr>written directly in a cell whose own end tag follows still differed (545 of 1,939 random tables): see the next entry.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
Table rows (Markdown and the
tablesformat) are in the order browsers lay them out: the first<thead>first and the first<tfoot>last (an empty one included), wherever they are written, and a later<thead>or<tfoot>where it is written, as CSS lays out only the first as the header or footer. Before, rows kept the HTML order, so a<tfoot>written before the<tbody>became the GFM header row and a<thead>written after it was a body row (headerRows0). A row inside a<div>in a<thead>belongs to that<thead>, and each run of rows directly in the table is its own row group, as the browser’s parser wraps each in a<tbody>. On 8,306 random tables with<thead>,<tbody>,<tfoot>and bare rows, thetablesgrid now matches the cell positions Chromium lays out in every one (2,169 differed before).EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
Table rowspans (Markdown and the
tablesformat) cover the rows browsers give them. A rowspan now counts every row it spans, also one whose cells end before its column, where it used to wait for the next row that reached the column:<tr><td>a</td><td>a2</td><td rowspan="3">b</td></tr><tr><td>c</td></tr>followed by two full rows putbin the fourth row and shifted its last cell one column right. A rowspan no longer runs past the end of its row group into the next<tbody>, and two spans over one slot both count every row. On 4,047 random tables of<tbody>groups and bare rows, thetablesgrid now matches the cell positions Chromium lays out in every one (1,897 differed before).EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
Table spans (Markdown and the
tablesformat) are read as browsers read them, by HTML’s rules for non-negative integers: the attribute’s leading digits (1.5is 1,2abcis 2), 1 when it has none or is negative, a colspan of 0 is 1, and a rowspan of 0 covers the rest of its row group (<thead>,<tbody>,<tfoot>, or the run of rows directly in the table), then capped at 1000 and 65534 as before. Before, they were kept as written: a fractional rowspan never ended and filled its column in every later row, so a 2,547-byte page withcolspan="1000" rowspan="1.5"over 165 empty rows gave 166,498,389 characters oftablesJSON and was not omitted (1,498,389 now), and a negative span counted negative characters, which gave budget back to the page’s 5,000,000. A rowspan of 0 was 1. Tables whose spans are absent or positive integers are written as before.EXTRACTOR_VERSIONis nowextract-tf/9; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
A
<head>tag inside the body is ignored, as a browser ignores it, instead of becoming an element. A slash closes only void elements, so<head/>is an open tag, and linkedom, the DOM layer, put everything after it up to its parent’s end inside it, where the Markdown skips a head:htmlToMarkdown('<head/><p>Some text</p>')gave"", and a page whose article had a second<head/>after its<h1>gave the heading alone, inhtmlToMarkdownand in main-content selection. Each such element is now replaced by its children wherever@w2l/extract-tfparses a page (Markdown,tables, main content, links, images, attributes, metadata); the document’s own<head>is unchanged.EXTRACTOR_VERSIONis nowextract-tf/8; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
A whole HTML document that leaves out
<html>(and often<head>and<body>), as HTML allows, is now read as a browser reads it: the head elements before the first content go into<head>, everything else into<body>. linkedom, the DOM layer, has no implied elements: it made the first top-level element the document’s root and left the rest beside it, so<!doctype html><table>…</table>gave each cell as a paragraph and notablesentry,<!doctype html><title>T</title><h1>H</h1><p>x</p>gaveTalone, and<!doctype html><head>…</head><body>…</body>gave empty Markdown; main-content selection gave""for<!doctype html><body><p>x</p></body>and<body></body>for a lone table, and a longer body-only page’s region carried an inserted empty<head></head><body></body>. Every module that parses a page (Markdown,tables, main content, links, images, attributes, metadata) reads the same rebuilt document. A document with<html>is parsed as before.htmlToMarkdownandhtmlToTablesalso take HTML that starts with<headfollowed by whitespace or>(after leading whitespace, a BOM included, and comments closed by-->; the empty comments<!-->and<!--->and a self-closing<head/>are not recognised) as a whole document, where before it was a fragment whose own<base href>was ignored; HTML that starts with<body>stays a fragment, sincemainHtmlcan be the body itself, and is resolved against the base the caller passes.EXTRACTOR_VERSIONis nowextract-tf/7; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
w2l --out <dir>also writesresults.csv: one row per page with its evidence, failed and blocked pages included, in the columns of the Python client’sto_pandas(include_markdown=False)(url,status,reason,final_url,fetched_at,http_status,lane,robots_decision,raw_sha256,markdown_sha256,extractor,source_commit,cache_state,cached_at), thenmarkdown_file, the page’s Markdown file beside it. A value W2L did not observe is empty, never 0. New guides for researchers: docs/guides/url-list-to-csv.md and docs/guides/citing-web-data.md. -
Reliability and throughput:
W2L_WORKER_COUNT(an integer from 1 to 64, default 4) sets how many pages the API’s engine (w2l-api,w2l serve) runs at once, under the per-host limits; a crawl’smaxConcurrencymay go up to it.npm run verify:batch-crash-1000posts a batch of 1,000 URLs over 20 loopback hosts, kills the API with SIGKILL once 400 are done and restarts it on the same task root: 8 checks, among them one item per URL, every URL whose step was on disk at the kill fetched once, and at most one extra fetch per worker for the pages in flight. It runs in the Linux CI job.npm run bench:throughputmeasures both lanes through the API: on 2026-10-03, on loopback, the HTTP lane did 3,460 and 1,880 pages/min (p50 460 and 715 ms) and the browser lane 523 and 525 (p95 987 and 983 ms) (report).npm run verify:serve-smokestartsw2l serve, scrapes a page and a PDF, stops the server mid-batch and checks that the restarted server resumes the batch to one item per URL; a newwindows-latestCI job runs it. -
Python client (
python/, PyPI namew2l, MIT;pip install 'w2l[pandas]'once published):w2l.batch(urls, **options)starts a batch on a W2L API (W2L_API_URL, default http://127.0.0.1:8787;W2L_API_TOKEN), waits for it, pages through every item and returnsJobResult(task_id, report, items);.to_pandas()gives one row per page, failed pages included, with the evidence columns (url,status,reason,final_url,fetched_at,http_statusasInt64,lane,robots_decision,raw_sha256,markdown_sha256,extractor,source_commit,cache_state,cached_at,markdown), an unobserved value missing rather than 0.scrape,mapandcrawltoo; options in snake_case or camelCase; API errors asW2LErrorwith the API’s code. A poll or page read survives a network error, a 5xx or a 429 (5 retries); a listing with more items but no cursor is an error. Tested withhttpx.MockTransport(python/tests, run in a virtualenv; not in CI yet). -
npm packaging:
npm run pack:packagesbuilds@w2l/sdk(MIT; ESM and CJS with@w2l/contractsbundled in, its declarations for import and require, no dependencies),@w2l/cliand@w2l/mcp(AGPL-3.0-only; one bundled file each, the third-party dependencies declared) into.w2l/pack/and packs them.npm run pack:checkinstalls the tarballs into a new project and uses each package against a loopback site. Each package declares only the dependencies its bundle imports (from esbuild’s metafile), and every entry guard but the bin’s is turned off in a bundle, sow2lrun by its real path starts nothing else.@w2l/contractsis MIT (it ships inside the MIT SDK);@w2l/mcpis AGPL-3.0-only in the repository too. Each package has a README. The SDK now exportsEvidenceRecord,PageTable,PdfPageMarkdownandPdfParser.w2l-mcprun throughnode_modules/.bin(npx) did nothing, since its entry guard compared a symlink with its target; it now compares the real path. Nothing is published yet. -
@w2l/cli(binw2l, AGPL-3.0-only):w2l scrape,crawl,batch,mapandserve. It runs the API’s engine in its own process, so every REST option is a flag under its kebab-case name (--max-age,--no-only-main-content,--formats markdown,tables,--header name=value, …), and the REST parser checks the request. Output is JSON,--markdownfor a scrape’s Markdown alone, and--out <dir>writes each page’s Markdown and each table’s CSV besideresults.jsonl. Ctrl-C leaves a crawl or batch paused (w2l crawl --resume <taskId>, refused for a crawl recorded as pending or running). Its task root defaults to.w2l/cli, apart from the API’s, and--webhookis refused, since a command runs no delivery worker.w2l serveis the API server (runApiServer, exported from@w2l/api), and the engine takesresumeOnStart: false, so a one-off command does not run earlier jobs.npm run scrape/crawlnow use it; the earlier in-process ladder CLI isw2l-ladderin@w2l/bench, which no longer claims thew2lbin. -
Map: a link first found over
http:is returned as itshttps:variant when the start page or a sitemap also lists that, instead of whichever came first. The https origin’s own robots.txt must allow it, and only an origin whose robots.txt the map read anyway counts, so the switch reads no further robots.txt, takes no slot of the host cap and cannot time the map out; otherwise the http link stays. The start URL stays as given.refused.samples.collapsednames the URL returned asinto. On www.python.org the map had returnedhttp://docs.python.org/3/tutorial/introduction.htmlthough the page also links its https form (MP13). Crawls still fetch the first variant seen. -
Markdown tables cap
colspanat 1000 androwspanat 65534, as browsers do, and bound the empty cells their grid adds for spans and short rows: 500,000 per table and 2,000,000 per page. A table past either is still one GFM table, written as each row’s own cells, unpadded. Before, an 892-byte page with fourcolspan="1000000"cells over 40 rows gave 516 million characters of Markdown (516,123 now), a 380 KB page whose one wide empty row sat over 20,000 one-cell rows gave 60 million (120,012 now), and a table of 200,000 rows threw. Tables within those limits are written as before.EXTRACTOR_VERSIONis nowextract-tf/6; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
tablesformat (scrape, batch, crawl; REST, SDK, MCP): every data table of the content the Markdown was written from, one entry per GFM table of that Markdown in its order, as{ tableIndex, caption, sourceUrl, headerRows, columns, rows, csv, csvSha256 }. Cells are plain text, a spanned cell is repeated in every slot it covers so no row is shifted, andcsvis RFC 4180 with CRLF line ends. Spans are capped as browsers cap them, and a table that would hold more than 2,000,000 characters (its spans repeated and every row padded to the widest), or more than what is left of 5,000,000 for the page’s tables, isomitted: "too_large"with no rows. The HTML-to-Markdown walk is shared, so the GFM output is unchanged (EXTRACTOR_VERSIONstaysextract-tf/5)./fcrefuses the format, which Firecrawl does not have. -
PDF options: scrape, batch and crawl take Firecrawl’s
parsers([]reads no PDF; onepdfentry,"pdf"or{ type: "pdf", mode, maxPages, pages, pageMarkers }).modeisfastorauto;ocrand theimageparser are refused by name.maxPages(1 to 10,000) cuts as asked and stayssuccesswith apage_capwarning.pages: trueaddspages: [{ pageNumber, markdown }], andpageMarkers: falseleaves out the<!-- page N -->lines./fcmaps them, plus v1parsePDF, and now writes no page markers unless asked, as Firecrawl does.data.metadata.numPagesgives a PDF’s page count.pdfToMarkdowntakespageMarkers, and its default output is unchanged (PDF_TEXT_VERSIONstayspdf-text/1). -
Sitemap reader behind an environment proxy: each redirect hop now carries its own
Host. undici’sProxyAgentwriteshostinto the headers object it is given, and the reader reused one object across a file’s hops, so after a cross-host redirect every later hop still named the first host. A CDN that routes byHostthen redirected again untilredirect_limit: https://python.org/sitemap.xml → https://www.python.org/sitemap.xml looped this way (the 2026-10-03 map runs’ MP13/MP14sitemap_unreadable). Direct connections and the http lane’s page requests were not affected. -
Cache: scrape, batch and crawl (REST, SDK, MCP
scrape/crawl/batch_scrape,/fc) takemaxAgeandminAge(milliseconds, 0 to ten years,minAgeat mostmaxAge),storeInCache(defaulttrue) andlockdown(cache only). W2L stores the most recently fetchedsuccessof each page under one set of options (the URL without its fragment, the mode, every lane option buttimeout, the rungs the request may use,fastMode, a recorded robots override, the extractor versions andW2L_SOURCE_COMMIT) in<taskRoot>/page-cache.sqlite; a request with customheadersstores only withstoreInCache: true, and reuses it only when a request asks: the defaultmaxAgeis 0, so nothing is reused unless asked, on/fctoo. A reuse is the original fetch unchanged, itsevidenceRecordincluded, withmetadata.cacheState: "hit"andmetadata.cachedAt(itsfetchedAt),channelsTried: []and no request, attempt, byte or browser time inusage; a lookup that found nothing iscacheState: "miss"and fetches; a request that looked nothing up carries nocacheState. Batch items and crawl pages carrycacheState/cachedAt, and a reused page iscached: trueand counts incachedPageswithout touching its host’s Crawl-delay.lockdownnever fetches: a page without a stored result isfailedwith the new failure reasoncache_miss(anagentHintsentry says why),/fc/v1/scrapeanswers it with HTTP 404SCRAPE_LOCKDOWN_CACHE_MISS, and a crawl in lockdown needssitemap: "skip". Modeauthedneither stores nor reuses (a lookup there is HTTP 400), a URL the request’s allowlist refuses is never answered from the cache, and Monitor captures neither read nor fill it.useCachedkeeps its meaning: a resumed crawl reuses its own pages. -
Evidence Record v1 gains
identity.device(desktop/mobileas the answering lane declared it; null in research mode and when no request was sent) andidentity.requestHeaders(the customheadersthat lane sent, sorted by name, each with the SHA-256 of its value and never the value;[]when none, null when no request was sent), and the reasoncache_miss. Both fields are optional in the schema file, so records written before them stay valid; W2L writes them on every record. -
HTTP lane: a body is decoded by its
Content-Encodingwhether or not W2L asked for it (its identity sends noAccept-Encoding; www.python.org answers gzip anyway, which left its pagefailed/empty_unverifiedwith 0 links and gave a map a misleadingstart_page_client_renderedwarning).gzip,x-gzip,deflate(zlib or raw) andbrare decoded, under the 50 MiB decompressed cap (decompressed_too_largepast it); scrape, batch, crawl and map all read through it, and so does the crawl’s sitemap reader, which before inflated only gzip found by its magic number.usage.bytesWireis now the size received, not the decoded size. The coding is onevidence.contentEncodingand, additively, on the Evidence Record ascontentEncoding(identityfor none; null in the browser and provider lanes). An unknown coding fails with the new reasonunsupported_content_encoding, a body that does not decode withparse_error; neither is read as the page. The declared identity headers are unchanged. -
Map:
POST /v1/map(SDKmap(url, opts),getMap(id);GET /v1/maps/:id) lists a site’s URLs from the sitemaps it declares and its start page’s links without fetching each page: at most one page body (the start URL, on the http rung alone), robots.txt of the start host and at most 20 others, at most 50 sitemap files, all inside one deadline (timeout, default 60,000 ms, 1,000 to 300,000). It takesurl,mode(standardorresearch),limit(1 to 100,000, default 5,000),originandintegration; anything else is refused by name (useIndexwith a hint;mode: "authed"refused). Each link carriesvia(start,link,sitemap), the sitemap file andlastmodthat listed it, its robots.txt verdict, and a title only from evidence in hand (the start page’s own metadata, an anchor’s text, a sitemap’s<news:title>). Robots-disallowed URLs are left out and counted inrefused; a host whose robots.txt could not be read is left out too, as a complete disallow, and said so apart (arobots_unreachablewarning with the reason; the start page’srobots: "unreachable"androbotsUnreachable), never worded as a rule the publisher wrote. At the deadline the map answers 200 with what it found aspartial(orfailedwhen it found nothing),stoppedBy: "timeout"and amap_timeoutwarning. When the start page’s links alone filllimitthe map answers at once,completedwithstoppedBy: "limit", and reads no sitemap (sooverLimitcounts only what was seen before it stopped). Each map is recorded under<taskRoot>/maps/. A hosted server capslimitat 5,000 andtimeoutat 60,000, refuses a browser-only URL, and counts/v1/mapin the rate limit. -
Map options:
POST /v1/maptakessearch(1 to 200 characters, at most 10 words; a URL is kept when every word appears, case-insensitively, in its percent-decoded URL or its title in hand; a filter in discovery order, counted beforelimit, with what it left out inrefused.searchFiltered),sitemap(include,skip: no sitemap requested,only: no page read, the start URL only when a sitemap lists it, and no sitemap read isfailedwithsitemap_unreadable),includeSubdomains(the crawl’sallowSubdomainsrule; robots.txt read once per new host, at most 20) andignoreQueryParameters(query variants folded into the first one seen, each fold counted with{ url, into }samples), plus the crawl’sincludePaths,excludePaths,regexOnFullURL,crawlEntireDomainanddeduplicateSimilarURLswith their crawl messages;allowSubdomainsandallowExternalLinksstay refused by name. Defaults stay as before (includeSubdomainsandignoreQueryParametersfalse, where Firecrawl v2 documentstrue), so a map without them answers as it did. Withsitemap: "only"a browser-only URL is not refused, since no page is read. A collapsed sample now names the variant as offered, with its query. -
/fc/v1/map(Firecrawl’s map, v1ignoreSitemap/sitemapOnlyand the v2 names):{ success: true, id, links: [url strings], warning?, agent_hints? }, or{ success: false, id, error, links: [] }for a map that found nothing; both v1 flagstrueis HTTP 400;useIndex,location,ignoreCache,threatProtectionandauditMetadataare refused by name.FIRECRAWL_SHIM_SNAPSHOTlists/mapunderpathsand no longer undernotCovered. -
MCP: the local
maptool (annotationsread-only, idempotent, open-world; anoutputSchema) answers a compact map ({ id, status, stoppedBy, links: [{ url, title?, description? }], warning?, agentHints?, counts }, or the native response withdebug: true) as text and asstructuredContent; a tool that declares anoutputSchemanow always answersstructuredContenttoo. The hosted MCP host does not offermap. -
The crawl’s sitemap reader takes two options a crawl does not pass, so a crawl is unchanged:
accept(judge each entry before it is collected; refused entries do not count towardmaxUrls) andsoftDeadlineAt(return what was read, the file in flight recorded asunreadable/timeout,truncated: "time").parseSitemapXml(text, { details: true })reads each entry’s<lastmod>and<news:title>in linear time also when the locs stand bare, and only a load that passesacceptorsoftDeadlineAtasks for them, so a crawl’s parse is the one it was;collectLinkDetails(html, base)gives each link once with its first non-empty anchor text.useIndexrefused on any route now carries the no-URL-index hint. -
Job streams:
GET /v1/crawl/:id/events(new) andGET /v1/batches/:id/events(re-shaped) stream a crawl or batch as server-sent events namedcatchup(the report as the stream opens),document(one per page recorded, every outcome, the compact page the items routes list;id:is the step cursor),snapshot(the report after each page),done(the terminal report, after which the stream closes) anderror({ code, message });?after=<cursor>orLast-Event-IDresumes after a document. The batch route’s oldprogress,pausedandcompleteevents are gone.GET /v1/crawl/:id/wsandGET /v1/batches/:id/wsupgrade to a WebSocket (@hono/node-ws, a new dependency of@w2l/api) carrying the same frames as JSON; a server with tokens takes the bearer token in the Authorization header or as the subprotocolw2l.token.<token>, echoed back, never in a query string; a missing job closes with 4404, a bad cursor with 4400,donewith 1000. The server subscribes to the in-processJobEventHubonce per client and reads the checkpoint once as a stream opens and once per page for every client of a job together (no more per-client 500 ms re-reads).AppOptions.jobStreams/W2L_JOB_STREAMS=offturns all four routes into 404s. NewApiEngine.listJobPagesandjobEvents. SDK:watcher(jobId, { kind, transport, pollIntervalMs, timeoutMs, after, signal, WebSocket })returns aJobWatcher(anEventTargetwithdocument,snapshot,doneanderrorevents, an async iterator,data,status,transport,close()), trying the WebSocket, then SSE, then polling the status and listing routes (pollIntervalMsdefault 2000, at least 250), switching once per level from the last document’s cursor so each document is emitted once per step id; a 401/403 is final,timeoutMsends the watch withwatcher_timeoutwhile the job runs on;crawlAndWatchandbatchScrapeAndWatchstart and watch in one call. MCP keeps request/response only; the frozen v1 shim adds no socket path. -
Batch:
GET /v1/batches/:idcountsempty_verifieditems assucceeded(a page read, with or without content), so with one step per URLcompletedissucceeded + failed.POST /v1/batchestakes Firecrawl’s extract scope flagsallowExternalLinksandincludeSubdomainsasfalseonly, which already holds;trueis HTTP 400 naming the crawl option that does it (allowExternalLinks: true is not offered on a batch: a batch fetches only the URLs given; a crawl takes allowExternalLinks, and extraction across links is the M5 multi-URL extract); MCPbatch_scrapedeclares them asconst: false. Docs name the extract mapping: a batch with a json format, no mergeddataorsourcesuntil the M5 multi-URL extract,creditsUsedandexpiresAtnull on the shim. -
agentHintsgains rows, derived from what the lanes recorded and never a check that was not made: an egress-policy refusal (ssrf_denied/governance_refusal) names the policy and the recorded reason; a page the http lane gotblockedor an HTTP error for and the browser lane then served names what the http lane got (read from the ladder summary) and says to expect the browser lane for the host (the ordinary hop after a thin or empty http answer stays silent);tls_error,timeoutandpartialname the honest option (skipTlsVerificationlocally, a longertimeout);empty_unverifiedwithout a shell caveat points atonlyMainContent: falseandincludeTags(a PDF without a text layer: no OCR); an incompletejsonnames the required fields not found and the model fallback’s environment, and amodel_unavailableissue says the fallback did not run; a 404 addscheck the link. At most five hints stay, in the table’s order. Batch items and crawl pages read the same table from the stored audit. -
Crawl and batch starts take
webhook(REST, SDK, MCPcrawlandbatch_scrape;/fc/v1/crawlmaps Firecrawl’swebhook): a URL string or{ url, headers, metadata, events, secretEnv }naming a receiver the job posts its events to as durable, retried deliveries (the Monitors’ delivery stack, now with akind: jobdestinationjob:<taskId>andevents_json,headers_json,metadata_json,payload_formatcolumns):started(sequence 0), onepageper page recorded, whatever its outcome, with the page as the items routes list it, thencompleted,failedorcancelledwith the job’s status report;eventsnarrows them (default all five; a filtered event is never enqueued),metadata(at most 32 strings of 1000 characters, 8 KiB) is echoed in every payload,headers(at most 32, 8 KiB, token names lower-cased,content-type,content-length,host,connection,transfer-encodingandx-w2l-*refused by name) go with every attempt and are stored in the control database alone (destinations exposeheaderNames), andsecretEnvsigns each delivery as for a Monitor. Payload{ schemaVersion: "w2l.job-event/v1", eventId, sequence, jobId, jobKind, event, at, metadata, page?, report?, error? }with deterministic event ids (<taskId>:started,<taskId>:page:<stepId>,<taskId>:<status>, the attempt id appended for a later attempt’s terminal event), so a resume or restart offers every persisted step again and sends none twice, and a finished job whose events a crash cut off is completed at the next start; headersx-w2l-event-id,x-w2l-event-version,x-w2l-delivery-id(and the signature pair).GET /v1/crawl/:idandGET /v1/batches/:idreportwebhook: { destinationId, url, events, pending, delivered, deadLetter };GET /v1/deliveries,/v1/deliveries/pageand/v1/delivery/destinationstakejobId(SDK and MCP too). The batch status gainssucceededandfailedcounts. A hosted server takes public https receivers only (webhook.url must be https (...),webhook.url must be a public address); a local one also takes plain http to a loopback receiver, sent direct over node:http (DeliveryWorkerOptions.allowHttpLoopback, localw2l-apiand the local MCP service), everything else staying HTTPS with verification.w2l-apinow runs a delivery worker of its own under a delivery policy it prints at start (TLS verified, the shell’s proxy never used,W2L_DELIVERY_PROXY_URL,W2L_DELIVERY_CA_FILE,W2L_DELIVERY_PRIVATE_ALLOWLIST). NewOrchestratorOptions.onStephook andJobEventHubin@w2l/apicarry a job’s events; the Firecrawl shim’s receivers get{ success, type: crawl.started | crawl.page | crawl.completed | crawl.failed, id, data, metadata, error? }(wrapJobWebhook). The hosted MCP host refuseswebhookonbatch_scrape; the lanes and the Firecrawl shim’s other mappings are unchanged. -
Batch and crawl starts take
idempotencyKey(1 to 200 characters; also thex-idempotency-key/Idempotency-Keyheader onPOST /v1/batches,POST /v1/crawland/fc/v1/crawl, which the Firecrawl v1 SDK sends): a retried submission with the same key and body answers the first one’staskIdwithreplayed: trueand starts nothing; the same key with another body is HTTP 409conflict; a body key that differs from the header is HTTP 400; keys live 24 hours in<task root>/idempotency.sqlite(newIdempotencyStorein@w2l/runtime), per the one API process that runs the task root. Batch takesappendToId(SDKappendToBatch(id, urls, options), MCPbatch_scrape): the URLs join an existing batch at the end of its list, the job keeps its options (one in the body is refused by name), a running batch fetches them in the same attempt (no job is added, somaxActiveBatchesdoes not count the append) and a completed one runs again for them in a new attempt (active again, it counts againstmaxActiveBatchesas a new batch does and is refused withactive batch limit reachedwhile the limit is reached); the 202 carriesrequestedandappended; a cancelled or failed batch is HTTP 409, a total over 1000 or a URL already in the batch HTTP 400, an id that is not a batch 404. On the way: a frontier seed passes the host scope (discovered links keep it), a batch’s governance lists no hosts (its own hosts added nothing and would have refused a URL appended on a new host), the orchestrator re-reads a batch’s row while it runs and writes the row as it then is at the end instead of the object it opened with. SDK:chunkUrls(urls, chunkSize = 100)andbatchScrapeChunked(urls, options, { chunkSize, itemLimit, pollIntervalMs, timeoutMs, maxRetries })run a list of any length as batches in sequence (a caller’s key becomes<key>:<chunk index>per job) and merge the items in submission order. The hosted MCP normalizer refuses both options onbatch_scrape; the lanes are unchanged. -
Batch (
POST /v1/batches, SDKbatchScrape, MCPbatch_scrape) takesmaxConcurrency(an integer from 1 to 4: the batch’s pages in flight at once, lowering the service’s worker count and never raising it; stored with the task, kept on resume, and reported as the cap in force onGET /v1/batches/:id) andignoreInvalidURLs(start with the entries that are http(s) URLs and report the rest asinvalidURLson the 202 and the status; a non-string entry or a duplicate is still refused). A batch entry that is not a URL is now refused by its index (urls[2] must be http(s),urls[2] is required) instead ofurl must be http(s). NewGET /v1/batches/:id/errors?cursor=&limit=(SDKgetBatchErrors, MCPget_batch_errors, also on the hosted MCP host as a read): the items that did not succeed across every attempt,{ id, timestamp, url, status, code, error, httpStatus }in pages of up to 1000, withrobotsBlocked, the URLs arobots_disallowedtrace event refused and no recorded override set aside. The hosted MCP normalizer keeps refusing the two options onbatch_scrape; the lanes and the Firecrawl shim are unchanged. -
formatstakesscreenshot(also Firecrawl v1’sscreenshot@fullPage, or one{ "type": "screenshot", "fullPage", "quality", "viewport" }entry per request), on scrape, batch, crawl, MCP and/fc(data.screenshot, a data URI string). Such a request runs on the local browser rung alone (channelsTried: ["browser_local"], the dropped rungs in the ladder audit); a server without one, orfastModebeside it, is refused with HTTP 400 naming it. The capture is taken after load, stability andwaitFor, before the DOM is read, CSS-pixel sized (scale: "css") at the declared 1280x800 viewport or theviewportasked for (integers 320…1920 by 240…1080, within the declared screen: a window size, the identity unchanged, recorded asscreenshot_viewport), the document’s whole height withfullPage(no scrolling first), a JPEG withquality. It is returned as{ contentType, width, height, fullPage, viewport, deviceScaleFactor, quality, bytes, sha256, path, base64 }with ascreenshot_capturedtrace event, saved as<sha256>.png/.jpgunderW2L_CAPTURE_RAW_DIR(the Evidence Record lists it askind: "screenshot"with size and type), attached to a success, an error page or a gate alike,nullwhen the browser could not capture it (screenshot_failed, ascreenshot_unavailablewarning and hint, the page kept) or no page rendered, and never repeated insummary.attemptsor stored audits. -
A thin or shell-like HTTP answer that stays the answer carries a
low_content_yieldwarning (The http lane extracted N tokens at confidence C; the browser lane did not improve it., or… was not available to this request.underfastMode, on an HTTP-only server or without a browser rung;found no main contenton afailed/empty_unverifiedshell), after its other warnings. Every response withwarningsalso carrieswarning, their messages joined with a space (full and compact scrape responses, batch items, crawl pages,/fcdata.warning, which so passes the native warnings through for the first time).agentHintsgains two table entries: alow_content_yieldwarning suggestswaitForor a longertimeoutwith the browser lane available (left out underfastMode, whose own sentence stands), and a file result says what itsmarkdownis. -
formatstakesimages(every image URL of the whole document as received:imgsrcandsrcsetcandidates,<picture>sources, lazydata-src/data-srcset/data-lazy-src/data-original, video posters,image_srclinks,og:imageandtwitter:image, resolved, absolute http(s), fragment stripped, deduplicated, in document order,data:URIs counted and left out;images_collectedin the trace) and one{ type: "attributes", selectors: [{ selector, attribute }] }entry (1 to 50 selectors under theincludeTagsrules; per selector the attribute’s values as written, in document order;attributes_extractedin the trace), on scrape, batch, crawl, MCP and/fc(data.images,data.attributes); both only when asked for and only on a page read as content. The page optionremoveBase64Images(defaulttrue) names what Markdown always did with an<img>whosesrcis adata:URI (left out, alt text kept);falsekeeps the image as, counted bycontentTokens;/fcnow maps the option with its value instead of acceptingtruealone. A json schema request is told apart from other object formats by itstype, so an attributes entry never switches on JSON extraction. -
metadata(scrape responses, batch items, crawl pages,/fcdata.metadata) carries the Open Graph tags a page states (ogTitle,ogDescription,ogUrl,ogImage,ogAudio,ogVideo,ogDeterminer,ogLocale,ogLocaleAlternate,ogSiteName), its Dublin Core tags (dcTermsCreated,dcDateCreated,dcDate,dcTermsType,dcType,dcTermsAudience,dcTermsSubject,dcSubject,dcDescription,dcTermsKeywords) and its article tags (publishedTime,modifiedTime,articleTag,articleSection), under Firecrawl’s names, each present only when the page states it and as written: no date normalisation, no fallback fromtwitter:*,govuk:*orcitation_*tags, JSON-LD or<time>. The seven existing fields keep their always-present, nullable shape. -
Roadmap v2 (weeks 1–16) makes P1 core correctness the current phase and adds the Firecrawl parity audit (
research/parity/, frozen at firecrawl-js v4.42.0) and a real-site test set with a runner (node research/parity/run-sites.mjs) whose runs are recorded with command and commit. -
JSON extraction no longer reports
completewhile a required field has no source: such fields are omitted with amissing_requiredissue, nested fields are not filled from page-level values, and JSON from a non-successpage isincompletewithpage_unsuccessful. -
Markdown keeps block boundaries, inline spacing and emphasis, numbers ordered lists (with
start), keeps code blocks and tables inside list items, drops empty emphasis from icon elements, and resolves link and image targets against the document base (<base href>included). Same-page#fragmentlinks stay as written. -
Scrape responses, batch items and crawl pages carry
metadataread from the page’s own markup: the<title>,metadescription, keywords and robots, the language, the favicon and the canonical URL, each null when the page declares none (nothing is inferred)./fcmaps them intodata.metadata.document.titlestays the content title. -
The API accepts several bearer tokens (repeated
--token,W2L_API_TOKENplus comma-separatedW2L_API_TOKENS) and compares them in constant time; the SDK sendsW2L_API_TOKENwhen no token is passed. -
A page whose blocks sit directly in
<body>(example.com today) is extracted instead of reportedempty_unverified, so the README’s firstnpm run scrape -- https://example.comworks again. -
In the browser lane, Markdown follows the page’s CSS where it differs from the tags: an inline element laid out as a block in the text flow starts its own paragraph (quotes.toscrape.com/js no longer reads
thinking.”by), and inline text hidden withdisplay: noneis left out (GitHub Docs’ platform names readOpen Terminal.). Hidden blocks, such as footnote popups and accordion panels, are kept. Evidence is unchanged (rawBodySha256and raw artifacts are the rendered page without W2L’s markers); a page over 100 000 elements, one still changing after capture, or one short of time converts by its tags, and the trace’slayoutevent records which. Monitors on pages captured in the browser lane may report a one-time change. -
A non-2xx response keeps its page as evidence (Markdown, links and
snapshot.httpStatus, also on the default REST response and/fc) while its status staysfailedorblocked; other 2xx statuses are judged like 200, and 204/205 areempty_verified. A robots.txt that never answers is recorded as unreachable instead of failing the scrape with HTTP 500, and a URL whose scrape throws becomes its own failed item instead of failing the whole batch or crawl. -
Main-content selection keeps a single long block such as a
<pre>news release; a form that wraps page content (a table viewer, an ASP.NET page) is unwrapped instead of removed; an HTTP page whose tables are empty script-filled shells is offered to the browser lane. -
JSON extraction also fills top-level keys from the page’s own labels (two-cell
th/tdrows anddt/ddpairs in the main content), each with its location as evidence; labels that state different values leave the field out with afield_ambiguousissue. A page whose one heading is followed by its one visible price is routed as a product, so that price is read (books.toscrape.com). -
JSON extraction accepts the JSON Schema subset Pydantic and Zod write: annotations (
title,description,$schema,$id,default,examples,format, …), checks (minimum,maxLength,pattern,const, …),definitions, andanyOf/oneOfof a schema and null or of primitive types. Any other keyword is refused withunsupported_parameter, naming it where it was sent (formats[0].schema.properties.author.allOf); README and the docs reference list the subset. -
JSON model fallback keeps every value read from the page with its evidence and merges the model’s answer key by key at every level: the model fills only what is missing, or a page value that breaks the schema, and each value it wrote has
modelevidence. -
JSON model fallback sends OpenAI-compatible strict mode a strict-safe copy of the schema (closed objects, every property required, optional ones nullable), so an ordinary schema no longer fails as
model_provider_error; a schema strict mode cannot express is sent without strict mode, andjson.modelUsage.strict/strictReasonsay which. -
JSON values filled from the content title, the URL or the page type carry evidence (
domh1[0]ortitle,fetchfinalUrl/requestedUrl,inferreddocument.pageType); Evidence Record v1 adds the field-evidence sourcefetch. -
JSON extraction no longer reads a number out of text that is not one amount (a SKU
HL-1was -1, a URL a fraction), no longer overflows the stack on a recursive schema, and names a value that breaks a schema check even when a required field is missing too. -
A listing of cards in which the article cascade finds no text block (a sparse category page, data.gov.uk’s home page) is kept instead of being reported
empty_unverified. An HTTP page that declares data its scripts will fetch (<link rel="preload" as="fetch">) is offered to the browser lane. Markdown leaves an empty first header cell empty instead of writing(header). -
Batch items and crawl pages carry
linkswhen requested.formatshas no count cap; an unsupported format or an unknown request field is rejected with HTTP 400 naming it, on the native API and on/fc. Crawl acceptsformats,includeLinks,includePathsandexcludePaths, and keeps the path filters on resume. -
SDK:
waitCrawl, andpollIntervalMs/timeoutMsforwaitBatchandwaitCrawlwith aWaitTimeoutErrorthat carries the last status. -
Scrape, batch and crawl (per page), MCP and
/fcacceptonlyMainContent,waitForandtimeout.onlyMainContent: falsereturns the whole page’s Markdown with the same evidence.waitForstarts at the browser rung and waits after load before capture; with no browser rung the result says so (policy_denied,wait_for_unavailable).timeout(default 300 000 ms) is the whole scrape’s deadline: when it fires the answer is HTTP 200 withpartial(the best content so far) orfailed/timeout, never HTTP 500, markedusage.deadlineExceeded. A lane timeout no longer setsbudgetExceeded: 'time', which the contract keeps for statusbudget_exceeded. JSON from apartialpage isincompletewith apage_partialissue./fc/v1/scrapenow cancels when the client disconnects. -
formatstakeshtml(the cleaned HTML the Markdown is written from: the main content, the whole page foronlyMainContent: false, or anincludeTagsselection) andrawHtml(the page as the answering rung received it, hashing tosnapshot.rawBodySha256), on scrape, batch, crawl, MCP and/fc; both are returned only when asked for and arenullfor a file or a result that is notsuccessorpartial. Two page options shape the content:includeTagskeeps only the elements its CSS selectors name, in document order, a named navigation included, andexcludeTagsremoves elements from the main content, the whole page, anincludeTagsselection and the evidence page of a failed result. An emptyincludeTagsselection issuccesswith empty Markdown; a block a later rung finds on that page is the answer instead, and a rung that repeats the empty answer ends the ladder before any vendor rung. They take tag, class, id and attribute selectors, descendant and child combinators,:not(),:is(),:where(),:rootand:empty, at most 100 selector parts per list, read with the DOM library’s own parser (an escaped name such as#\31 23is what it is to the library) and matched in time proportional to the page; a selector that does not parse isinvalid_request, and one that uses a sibling combinator, a positional pseudo-class,:has()or another pseudo-class isunsupported_parameter, because matching those can take unbounded time that notimeoutstops. -
Request errors carry one code set across the REST API,
/fc, the SDK and MCP:{ error, code, details? }withinvalid_json,invalid_request,unsupported_parameter,unsupported_format,unauthorized,not_found,conflictorinternal_error; statuses and messages are unchanged. The SDK throwsW2LError(status,code,method,path,body). A 500 hides its internal message unless the server runs in local mode; hosted mode logs the cause to stderr. The docs reference lists the codes. -
Local mode sends outbound requests (HTTP lane, robots.txt, Monitor fetches, the browser lane) through
HTTPS_PROXY/HTTP_PROXY, honouringNO_PROXY; loopback stays direct,W2L_PROXY=offignores the variables, and hosted mode never reads them. Proxied requests recordevidence.envProxy(host:port, never credentials). Response headers up to 64 KiB are accepted, and a name that does not resolve isdns_errorinstead ofpolicy_denied. -
Crawls follow links on the start URL’s
www.twin and on the host it redirects to, and no longer fetch image, font, style, script, media or program links. A crawl task stores every option: a crawl paused by shutdown or left running by a crash resumes when the API starts,POST /v1/crawl/:id/resume(SDKresumeCrawl, MCPresume_crawl) restarts a paused or failed one, andw2l crawl --resumeruns with the stored limits and refuses a flag that differs.maxPagescounts the task’s pages across resumes.GET /v1/crawl/:idcounts pages while the crawl runs, and crawl pages carry their audit and trace only withdebug=true. The HTTP lane reports robots.txtCrawl-delay, a host starts one page at a time until its robots.txt is known, and each crawl page records the delay it waited in acrawl_delaytrace event./fc/v1/crawl/:idreportscreditsUsedandexpiresAtasnullinstead of 0 and an invented expiry. -
Monitor observations record the raw body’s SHA-256 and the extractor version (
EXTRACTOR_VERSION, nowextract-tf/1; unknown on older rows). Fields that change over a byte-identical raw body (a 304 counts as its reused body) areextraction_reprocessed: the run recordschangeReason(MCPget_monitor:latestRun.changeReason) and the baseline moves, but no event or webhook delivery is created. So the first run after this upgrade no longer reportssource_changedwhere only W2L’s Markdown changed (for the Firecrawl introduction preset: restored spaces in three descriptions). When the raw body changed too, the event stayssource_changedand carriesextractorChangeif the extractor differs from the baseline’s. DeliveredeventVersionvalues can now skip a version. -
A robots.txt that cannot be fetched (5xx, network error or lookup timeout) is a complete disallow in local and hosted mode alike (RFC 9309 §2.3.1.4): the page is
failed/policy_denied, and the trace and the compliance record’s robots decision carryunreachablewith the reason, which the record’s hash covers. It is fetched again after five minutes (robotsUnreachableTtlMs). Before, only the public preview failed closed; the API, MCP and CLIs fetched the page anyway. The provider lane follows the same rule and no longer reads a 5xx robots.txt as no robots.txt. -
The hosted public preview names OctoCrawl to the sites it reads: its User-Agent, on robots.txt, the page and the Amazon.sg browser’s requests, is the standard Chrome User-Agent followed by
OctoCrawl-Preview/1.0 (+https://octocrawl.dev), with the same client hints, so a site can address it in robots.txt withUser-agent: octocrawl-preview(oroctocrawl). The token wasW2L-Preview/1.0 (+https://github.com/77777R7/w2l)before the rename, so a robots.txt group namingw2l-previeworw2lno longer applies to the preview. The local API, MCP and CLIs are unchanged: standard mode still sends the plain Chrome User-Agent. -
Without the environment proxy (no variables,
W2L_PROXY=off, hosted mode), the browser lane launches Chromium with--proxy-server=direct://instead of silently using the operating system’s proxy; with it, the launch names the environment proxy. -
W2L_CONTACTadds the operator’s contact to the research-mode User-Agent (; contact: …). On the HTTP lane, a 403 from an SEC host to a request that declared no contact carries adeclared_contact_hinttrace event. -
Research mode with
W2L_CONTACTdeclares SEC’s own User-Agent format,W2L Research <contact>, to sec.gov and its subdomains in the HTTP and browser lanes (robots.txt included), which SEC.gov serves where it answered 403 to the research format; a robots.txt group forw2l-researchstill applies there, the Evidence Record’sidentity.contactreads either format, and other hosts keep the research User-Agent. -
Crawl
includePaths/excludePathsthat can backtrack catastrophically, such as^/(a+)+$or.*a.*b, are refused withinvalid_request(unsafeRegexReasonin@w2l/contracts); the rest run on V8’s linear-time regular expression engine where it can run them, and otherwise (lookaround, backreferences, counted repetitions above 16) on paths of up to 2,048 characters with a 100 ms limit per link, so a crafted link path can no longer stall the server. A link a filter cannot decide is not followed. -
JSON Schema
patterngets the same check: a catastrophic pattern is refused withinvalid_request, and page or model text is matched in linear time where V8 can, otherwise only up to 2,048 characters (longer text counts as breaking the pattern). -
Hosted mode enforces its crawl limit (100 pages; 10 on the hosted MCP): an omitted or
nullmaxPagestakes it and a larger one is refused withinvalid_request, wherenullor a large number used to skip it. -
w2l-apirefuses to start when a--tokenhas no value (last, followed by another flag, or blank) instead of taking the next flag as the token or ignoring it; the error never repeats a token. -
Evidence Record v1: scrape responses (full and compact, so MCP too), batch items and crawl pages carry
evidenceRecord(schemaVersion: "w2l.evidence/1"), one shape for every lane, published as the JSON Schemapackages/contracts/schemas/evidence-record.v1.json: final URL (null when no request was sent), a redirect chain that says whether every hop was observed,fetchedAt, HTTP status, status and reason, lane, the robots.txt decision with its hash, raw and output hashes (delivered Markdown;json.dataas RFC 8785 canonical JSON), extractor name, version andW2L_SOURCE_COMMIT, field evidence, saved artifacts, the environment proxy and the User-Agent sent. Existing fields are unchanged; lanes now also recordevidence.fetchedAt, and the HTTP, browser and provider lanes’robots_checkedtrace events carry the robots.txt hash (the provider lane emits one too)./fcdoes not carry the record. -
PDF text, as a library function only (scrape, batch, crawl, the API and MCP do not reach it yet):
pdfToMarkdownin@w2l/extract-tfturns PDF bytes into Markdown with a<!-- page N -->line before each page and each page’s text offsets, read with the pinnedpdfjs-dist6.3.289 (Apache-2.0). No OCR (no_text_layer), tables unverified (tables_unverified), a page cap and a time budget, and an error result for encrypted, malformed or non-PDF input. Checked on 10 public reports inresearch/pdf-corpus/. -
A scrape’s
timeoutalso sets how long the lanes wait for a slow server: with one, the HTTP lane waits for headers and body and the browser for navigation until the deadline, instead of stopping at 10 s, 30 s and 20 s; without one those defaults stay. The browser’snavigatetrace event records the wait it allowed (timeoutMs). -
MCP: a client that cancels a tool call (
notifications/cancelled) aborts the API requests the call made, so a cancelledscrapestops on the server and a cancelledwait_batchstops waiting. -
MCP over HTTP: the local and hosted Streamable HTTP services stop a call when its client cancels it (
notifications/cancelled) or closes the call’s request before the result. Before, every POST got a new server, so neither reached the call, which ran until its deadline. The services still keep no session, but give each client anMcp-Session-Idat initialize, and a cancellation reaches only a call made with the same session id and, on the hosted service, the same bearer token (a hosted client that sends no session id is matched by its token; a local one can stop a call only by closing its request). The call’s request is answered with the JSON-RPC errorRequest cancelled(code 0). -
SDK:
scrapewaits for the API’s answer until itstimeout(default 300 000 ms) plus 30 s. On Node, fetch stops waiting for the response headers after 300 to 301 s, and the API answers a scrape with the default or largesttimeoutat its 300 000 ms deadline: an answer that arrived later than fetch’s wait threwTypeError: fetch failedinstead of returning the API’sfailed/timeout. -
Batch and crawl: a page’s
timeoutalso bounds JSON extraction’s model fallback, which ran until the task’s own deadline; a fallback it cuts short leaves the JSONincompletewith amodel_timeoutissue. -
SDK:
waitCrawlandwaitBatchretry a status request that fails with a network error, 408, 429 or 5xx (after 1, 2, 4, 8, then 10 s, or aRetry-Afterof 60 s or less;maxRetries, default 5) and throw other errors at once.timeoutMsalso ends a request in flight,WaitTimeoutError.lastis null when no status was read andcauseholds the last failure, invalid wait options throwRangeError, andW2LErrorcarriesretryAfterMs.crawlAndWaitandbatchAndWaitstart a task, wait, and return every page and error, or every item. -
/fc/v1/crawl/:idreports a cancelled crawl ascancelledinstead offailed;completedcounts the latest attempt’s successful pages andtotalall its pages (failed, blocked and duplicate ones included) plus, while this API process runs the crawl, those in flight and queued (nullfor a paused crawl), instead of both being the step count.dataholds at most 100 pages (limit, up to 1 000) with anextURL carrying a native cursor, left out after the last page;skipis rejected. -
Markdown leaves out
data:image URIs and keeps their alt text, as Firecrawl’sremoveBase64Imagesdoes by default;/fcacceptsremoveBase64Images: trueand rejectsfalse. -
The Markdown of an error page kept as evidence resolves link and image targets against the page URL, as on a success (MDN’s 404 page no longer has site-relative links).
-
A page on which the extractor finds no main content keeps the whole page’s Markdown and links as evidence on its
failed/empty_unverifiedresult, in the HTTP, browser and provider lanes. When the HTTP rung found no main content and the browser rung then fails without a page, the answer is the HTTP result with its page (ladder_evidence_keptin the ladder audit) instead of the browser’s failure; when the deadline ends the browser rung,failed/timeoutkeeps that page. -
onlyMainContent: falsereturns the whole page assuccesson a page where no main block is found, instead offailed/empty_unverified; the browser rung is still tried, and answers when it renders more. -
A table whose nested tables hold at least half of its text no longer counts as the page’s data table: on Hacker News the main content is the story list, not the whole layout table with its header and footer. In Markdown, such a table, or one with a single row, becomes paragraphs, and only the data tables inside it become GFM tables. A data table with a small table in one cell stays the page’s table and one GFM grid, with that small table as the cell’s text.
-
The browser lane reports the response’s
content-typeheader instead oftext/html; rendered, and every redirect hop Chromium followed instead of[requested, final]; its Evidence Record saysredirectChain.complete: truewhen every hop is known (evidence.redirectChainComplete). -
The browser lane reports the final URL, status and
content-typeof the document the page shows, and judges the page by them, instead of the status and type of the navigation W2L started next to the URL and content of another document. A page that answered 200 and then went to a 404 page bylocation.replaceor a meta refresh isfailed/http_errorwith the 404 page’s URL, status and type on REST, in the Evidence Record, in the compact snapshot and on/fc(it wassuccesswith 200); a 403 challenge reached that way isblocked. The redirect chain lists each document a script or a meta refresh loaded after the server’s redirects, withcomplete: true, also for a file the browser displays or downloads (which kept[requested, final]without the flag). A URL the page sets with the history API (pushState,replaceState) is not a redirect: the final URL stays the one the document was loaded from, with its status, the page’s URL stays the base of its links, and asame_document_navigationtrace event names both. A page that navigates during the wait for stability orwaitFor(a zero-second meta refresh) is captured on its new document instead of failing withconnection_error. A page that loads a new document during every read (a refresh or script redirect loop) isfailed/redirect_loopwith no content and no status, and the lane no longer waits without end for Chromium to close it. Server redirects Chromium stops following (a loop, or more than 20) areredirect_limiton the browser lane instead ofconnection_error. On the HTTP lane, a redirect loop’s final URL is the URL whose redirect closes the loop, whose status it reports, instead of the URL that redirect points back to. -
The compact scrape response, and so MCP
scrape, carriessnapshot.contentTypenext tosnapshot.httpStatus;/fcmetadata carriesurl(the final URL) andcontentType. -
EXTRACTOR_VERSIONis nowextract-tf/2for the Markdown changes above; a Monitor whose Markdown changes only through them recordsextraction_reprocessedonce. -
File download: scrape, batch, crawl, MCP and
/fctake PDF, CSV, JSON, plain-text, XLSX, XLS and ZIP responses (byContent-Type, or by the bytes underapplication/octet-streamor none) as files. The bytes are saved as received to<W2L_TASK_ROOT>/files/<sha256>.<ext>(stored once per content), described in a newfileblock (kind, detection, content type, declared and received size, SHA-256, path, PDF pages) and listed in the Evidence Record’sartifactsaskind: "file";rawSha256is the SHA-256 of those bytes. A file never escalates to the browser, and the browser lane catches the download a file starts (waitFor) instead of failing withDownload is starting. A PDF’s Markdown is its text with a<!-- page N -->line per page and the extractorpdf-text/1; a PDF with no text layer isfailed/empty_unverified, one that cannot be openedfailed/parse_error; CSV, JSON and text give their text (file-text/1), XLSX, XLS and ZIP none. JSON extraction reads a PDF’sLabel: valuelines withfieldEvidence{ source: "pdf", locator: "page N \"label\"" }and never sends PDF text to a model.W2L_MAX_FILE_BYTEScaps a file (default 50 MiB, at most 500 MiB) and a request’smaxFileBytescan lower it; a file over the cap isfailed/body_too_largewith its declared size and nothing saved. Images, audio, video, fonts and office documents other than spreadsheets arefailed/unsupported_content_typeinstead of being parsed as pages and sent to the browser. A response body that stalls or breaks off after the headers isfailed/timeoutorconnection_errorinstead of HTTP 500. The Evidence Record’s artifacts gainbytesandcontentType, optional in the schema under the new additive versioning rule ofw2l.evidence/1. -
Node.js 22.13 or later is required (
engines), as pdf.js needs it; CI runs Node 22. -
JSON extraction reads a number as the page writes it: a decimal comma or point, thousands grouped by
.,,, any space or an apostrophe (India’s lakh groups too), a currency before or after,,-for a whole amount. On a product page’s visible price,12,99 €gave 1299 and now gives 12.99;1.299,00 €gave 1.299,1 299,00 €29900 andCHF 1'299.001, and each now gives 1299; a JSON-LD priceCall for pricegave 0. Each of those was reportedcomplete. A single.or,before three digits (1.299 €,$1,299) is read only when a review count, a currency without minor units or a JSON-LD / OpenGraph price settles it; otherwise the field is left out (nullwhen nullable) with afield_unavailableissue quoting the text, instead of a number that may be 1000 times off. The page’s language, currency and domain are not used to guess. Page labels and PDFLabel: valuelines follow the same rule (a label reading$1,299gave 1299 and is now left out with that issue;1.299,00 €, refused before, gives 1299), andjson.evidencequotes the text of each number read from text (text); the Evidence Record is unchanged. The visible-price reading in@w2l/extract-tftakes an amount grouped by spaces or apostrophes whole (1 299,00 €, not299,00 €); which elements count as prices is unchanged, so the Markdown is too andEXTRACTOR_VERSIONstaysextract-tf/2. -
A product list or map the extractor reported empty (
images,prices,variants,specifications) has a JSON evidence entry,inferredwithdocument.product.<name>, where it had none. -
README and the docs reference state the JSON Schema
patternlimits the code has: a pattern of at most 2,000 characters, and the 2,048-character text limit also for counted repetitions above 16,\p{…}escapes and text with characters outside the Basic Multilingual Plane, not only for lookaround and backreferences. -
A page written as
<div>s around its tables, such as an SEC EDGAR inline XBRL filing, keeps its text. The table strategy, which keeps one table, is used only when the page’s tables hold at least half its text; on a page whose<p>elements hold under 750 characters,<div>s with no block inside count as paragraphs when they hold at least that much prose; and when the blocks the body lays out itself (inline wrappers such asix:nonNumericlooked through) hold most of the text, the body is the main content. A filing’s hidden<ix:header>(XBRL facts and contexts) is left out of the main content and of every Markdown. Before, IREN’s 10-Q (real-site case A36) wassuccesswith only its statement of operations table (6,843 characters of Markdown).EXTRACTOR_VERSIONis nowextract-tf/3; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
A page W2L extracts by its tables (
document.strategy: "table") keeps all its data tables: its main content is the lowest element that holds them, with the headings, captions and text between them, instead of the largest table alone, and JSON extraction reads page labels from all of them. Before, a statistics page with three captioned tables under<h2>headings and its<h1>in a banner wassuccesswith the first table only. A menu laid out as a table (mostly links, and no figures outside them) counts only when it is the page’s largest table, and a table with no text is not a data table. Table cells and captions keep their links and images with absolute targets, as paragraphs do, with every|written\|; before, Hacker News’ story list had no link target. Adata:link target is dropped as adata:image is, and the link keeps its text.EXTRACTOR_VERSIONis nowextract-tf/4; a Monitor whose Markdown changes only through this recordsextraction_reprocessedonce. -
A recorded robots override for one URL: scrape takes
robotsOverride({ reason, recordedBy? }) and a batch takesrobotsOverrides([{ url, reason, recordedBy? }], eachurlone of the batch’s own, stored with the task), on REST, the SDK and MCP. robots.txt is still read and its verdict recorded; when a rule disallows the URL, the fetch goes ahead with arobots_overriddentrace event afterrobots_disallowed, arobots_overriddenwarning first in the result’s newwarnings(full and compact scrape responses, batch items), the override in the browser lane’s compliance record (skippedFetch: false, covered by its hash) androbotsDecision.userOverride: truein the Evidence Record. Other URLs are unaffected, an unreachable robots.txt is not set aside, and a blanketignoreRobotsTxtstays an HTTP 400 that names it. The warning and the events also stay on afailed/timeoutwhen the deadline passes while the request is out, and on the answer of a later rung that fetched nothing. The override is for local fetches: the provider lane takes none, a scrape that set a rule aside does not go on to a vendor rung, and a hosted server (--hosted) refuses both fields with HTTP 400unsupported_parameter. -
Five execution options on scrape, batch and crawl (per page), MCP and
/fc, each stored with a batch or crawl task.headers(at most 32, values of at most 4,096 characters) are sent to the requested origin after the declared identity and never override it:User-Agent,sec-ch-*,sec-fetch-*,Authorization,Proxy-Authorization,Cookieand the transport headers are refused with HTTP 400 naming them; a redirect hop to another origin and robots.txt get the identity alone on both rungs (custom_headers_withheld; the browser rung adds the headers per request through Chromium’s request interception, which carries no override to a redirect hop, instead of a Playwright route, whose overrides ride every hop), and the headers are on the record (request_headers_added, the HTTP rung’sidentity_sent, the browser rung’s compliancesentHeaders).mobilefetches as a second declared identity, Android Chrome with aligned hints and a 412x915 touch viewport, which passes the same coherence and honesty checks, is evaluated against robots.txt and is recorded (identity_sent.device,identity_declared); it is refused with research mode.skipTlsVerificationrelaxes certificate verification for one local fetch and its robots.txt lookup, through routes of its own closed after it, with atls_verification_skippedtrace event and atls_unverifiedwarning; a hosted server refuses it. Without it a certificate that does not verify is nowfailed/tls_errorwith the error code in the trace (it wasconnection_error, orpolicy_deniedwhen robots.txt failed on the same certificate).fastModekeeps the HTTP rung alone (ladder_channels_filtered): a script-written page isfailed/empty_unverified, never rendered, and the response carriesagentHintswhen the HTTP rung asked for the browser rung; a URL the server binds to the browser lane refuses it.blockAds(default true) aborts requests to a bundled list of about fifty ad-serving hosts on the local browser rung (ads_blocked) and switches the extractor’s ad and cookie-banner pruning, which was always on;falsekeeps them inmarkdownandhtml. A request withheaders,mobileorskipTlsVerificationnever goes on to a vendor rung.ApiEngineOptions.hostedtells the engine it serves a hosted API (set by--hostedand the hosted MCP host). -
The browser lane sets Chromium’s user-agent metadata to the declared identity, so the client hints Chromium generates itself (
sec-ch-ua,sec-ch-ua-mobile,sec-ch-ua-platformon a server redirect hop and on the page’s own requests, where a context’s extra headers do not reach) carry the declared brands instead of the headless shell’sHeadlessChrome, andnavigator.userAgentDatasays the same. Before, a redirected page’s second hop went out withHeadlessChromehints while the compliance record repeated the declared value, since Playwright reported the headers it had set rather than the wire’s. The as-sentsentHeadersnever carry a credential header (cookie,authorization): the access fact keeps the session as a hash. -
The HTTP rung says when the page it read looks like a shell for data its scripts fill in: a
<table>with no cells beside scripts, an empty application root, a page of scripts with little text, an element hidden once scripts run, a hydration blob on a thin page oraria-busy; anoscriptnotice alone counts only on a thin page or beside hydration state. The result keeps its status, thefailed/empty_unverifiedresult that keeps the page as evidence when no main region was found included, and carries aclient_rendered_suspectedwarning naming the rule, its trace aquality_client_renderedevent, and the ladder offers the page to the browser rung as it does a thin result, keeping whichever answer holds more; the browser rung never raises it. The extractor exposes the signals asrender. -
C2 adds Monitor sample preview, paused-by-default MCP creation, durable queued manual runs, run detail, delivery paging and dead-letter controls. A local SDK Streamable HTTP check completed the public Firecrawl document → HTTPS webhook flow with the same event ID at sender and receiver.
-
C3 adds a single-process API/scheduler/delivery-worker runtime, protected-resource metadata, WorkOS JWT validation and a restricted Streamable HTTP MCP endpoint. A two-service Render Blueprint and walkthrough are ready; permanent deployment, browser OAuth and actual hosted restart acceptance remain open.
-
Scrape requests now accept explicit Markdown, links and JSON Schema formats. Deterministic extraction runs directly on subject-bound product HTML/metadata; nullable missing fields include evidence-backed issues, and an OpenAI-compatible model fallback is explicit and disabled when unconfigured.
-
Amazon
/dp/{ASIN}pages use a subject adapter for identity, purchase offers, seller, availability, delivery context, variants, images and specifications. Recommendation shelves are removed before output; Blink subscription pages remain products withkind: subscription. -
HTTP usage includes monotonic queue, robots, cooldown, transport, retry, extract, format, model and total timings.
request_completeis emitted after the response body resolves, while laddertotalMsrecords user-visible elapsed time separately from summed attempt time. -
MCP scrape responses are compact by default and omit trace, ladder audit and nested duplicate bodies;
debug: truerestores the full audit. The fixed three-round Amazon baseline writes ignored local evidence with region pinning, field checks, response bytes and p50/p95 timing. -
Every scrape response carries the facts of the call in
metadata, under Firecrawl’s names, beside the page’s own fields:scrapeId(a UUID perPOST /v1/scrapeor/fc/v1/scrapecall, also at the top of the full response),sourceURL,url,statusCode,contentType,proxyUsed(operatorfor the server’s environment proxy,userfor the caller’s own egress, else null),timezone(the browser lane’s declared zone; null on the HTTP lane) and the concurrency pair below;/fcaddscreditsUsed: null. A scrape response’smetadatais now always present, with the page fields null on a page that was not read as content; batch items and crawl pages are unchanged. The call’s record, without any page body, is written to<task root>/scrapes/<scrapeId>.jsonand served byGET /v1/scrapes/:id(SDKgetScrape, MCPget_scrape); an unknown id is 404not_found, and records have no retention yet. -
metadata.concurrencyLimitedandmetadata.concurrencyQueueDurationMssay whether the per-origin concurrency ceiling (W2L_PER_HOST_CONCURRENCY) held a scrape’s attempts back and for how long, cooldown and pacing excluded; each lane result carries its own hold asusage.timings.concurrencyWaitMswhen there was one, and the origin scheduler’s permits reportlimitedByConcurrencyandconcurrencyWaitMs.usage.timings.queueMskeeps its meaning. -
integrationandoriginon scrape, crawl and batch (REST,/fc, SDK, MCP), 1 to 100 printable characters without spaces, go into W2L’s own records only: the scrape record and the task, whichGET /v1/crawl/:idandGET /v1/batches/:idreport asattribution. The SDK sendsorigin: js-sdk@<SDK_VERSION>unless the call sets one; the MCP server recordsmcp-<client name>@<client version>from the client’sinitializeand refuses anorigina tool call names;/fcmaps theoriginFirecrawl’s SDKs send instead of dropping it. -
agentHintson scrape responses, batch items, crawl pages and the scrape record (agent_hintson/fcpages) say what to change about the request next time, one sentence each: a robots.txt rule and the recorded override, a login wall andmode: authed, a challenge or bot gate and the lanes tried, a Retry-After time, a cut, a script-filled shell and whether the browser lane had its turn, an error page kept as evidence. Refusals ofstealth, a stealth or enhancedproxy,ignoreRobotsTxtand hostedskipTlsVerificationcarry a hint naming the supported route in the error body;W2LError.agentHintscarries them in the SDK. UnderfastModethe existing fastMode hint stands alone. -
W2L_RATE_LIMIT_PER_MINUTE/--rate-limit-per-minute(1 to 100000) cap the requests that start work (POST /v1/scrape,/v1/crawl,/v1/batches,/fc/v1/scrape,/fc/v1/crawl) per bearer token in a sliding minute, in memory and per process: over it, HTTP 429 withRetry-Afterand{ error, code: "rate_limited", retryAfterSeconds, agentHints }(/fc:{ success: false, error, code, agent_hints }). The SDK’sW2LErrorcarriesstatus429,coderate_limited,retryAfterMsand the hints and retries nothing; MCP tool calls fail withrate limited: retry after <s> s (rate_limited). -
Crawl sitemaps, on REST,
/fc, the SDK and MCP:sitemap(include, the default,skiporonly). A crawl now reads the site’s sitemap once per attempt before its first page, the files the start URL’s robots.txt names or/sitemap.xml, so every crawl makes one more request to its host (a sitemap read, or a probe of/sitemap.xmlthat may answer 404), and a bounded crawl of a site with a sitemap returns different pages than before: the entries are queued at depth 1 after the start URL and ahead of its links, under the same host, subtree, path and depth rules.onlyfollows no page link;skiprestores the old behaviour. A<sitemapindex>is followed one level, a gzip file inflated under the 50 MiB cap, at most 20 files read, and the load stops once it holdsmaxPagesentries. Sitemap files are fetched with the crawl mode’s http identity, the SSRF checks, the operator proxy, the origin scheduler’s pacing, the policy’s redirect limit and 10 MiB wire cap, each after its own URL’s robots.txt verdict (a disallowed or unreachable robots.txt refuses it), never through the ladder; they have no signed compliance record. The report’sdiscovery.sitemaplists every file read, refused or unreadable with its status, bytes, SHA-256, kind, entry count, robots verdict and proxy use, and a page found through the sitemap carriesdiscovered { via: "sitemap", from: <file> }. The/fcshim maps v1ignoreSitemap(trueisskip,falseisinclude, wherefalseused to be refused) andsitemapOnly, and passes v2sitemapthrough. A task stored before the option resumes without a sitemap; thew2l crawlCLI reads none yet. -
maxConcurrencyon a crawl (REST,/fc, SDK, MCP): an integer from 1 to the service’s worker count (4 locally, 2 on the hosted MCP host; above it HTTP 400maxConcurrency must be at most N on this service) that lowers the pages the crawl fetches at once and never raises the per-host ceiling; stored with the task and kept on resume. -
GET /v1/crawl/active(SDKgetActiveCrawls, MCPlist_active_crawls): the crawls this API process is running, each with its id, start URL, status, start time, pages so far and stored options; always 200, empty when nothing runs, never a batch. -
SDK pagination caps:
listCrawlPagesandlistBatchItemstakemaxPages,maxResultsandmaxWaitMsand say where they stopped (nextCursor,stoppedBy);getCrawlDocumentsandgetBatchDocumentsreturn the status and the documents in one answer,collectCrawlPagesandcollectBatchItemsthe documents alone. MCPget_crawl_pagesandget_batch_itemstakemaxResults(1 to 200) and follow the cursors themselves. Server page sizes are unchanged. -
Crawl URL scope, on REST,
/fc, the SDK and MCP:crawlEntireDomain(defaultfalse: a crawl now follows links on the start URL’s host only inside the start URL’s path subtree, as Firecrawl’s default does, so a crawl seeded below the root returns fewer pages than before;truerestores the whole host),allowSubdomains,allowExternalLinks(refused besideallowlistedDomains),regexOnFullURL,ignoreQueryParametersanddeduplicateSimilarURLs(defaulttrue:/aand/a/,/and/index.html,www.and apex,httpandhttpsare one page, fetched once, so a crawl of a site that links both variants fetches one page fewer than before; the content-hash duplicate check stays as the backstop).allowlistedDomainsnow adds hosts to the start URL’s own instead of replacing them, and the crawl’s governance allowlist, when there is one, is derived from the same rule. Every crawl reports its link discovery:discoverycounters onGET /v1/crawl/:id(offered,enqueued,duplicate,collapsed,hostDenied,subtreeDenied,pathDenied,depthDenied,duplicateContent;nullfor a batch), written with the attempt after every page, and on each page’s trace adiscoveredevent (via,from) and alinks_offeredevent with the page’s counters and samples of its collapsed and host-refused links.GET /v1/crawl/:id/pagesleaves out pages whose body repeated an earlier page’s unlessincludeDuplicates=true(MCPget_crawl_pages, SDKgetCrawlPages), and/fc/v1/crawl/:idleaves them out ofdatawhiletotalstill counts them. The/fcshim maps the six options, with v1’sallowBackwardLinksascrawlEntireDomain. A crawl task stores the options; one stored before this change resumes with its original whole-host, exact-URL rule. Thew2l crawlCLI has no flags for these options yet: it keeps following the whole host and takes the other defaults.
0.4.0-rc.1 — 2026-09-22
-
Gate 2–4 implementation freeze
99894bd636ecafd254a7c7bc79d26e9a97fa9199was merged by PR #50 intomainat1c1481722ade26b717d18a34fa4b46362f53acf8and published as thev0.4.0-rc.1source prerelease. Workspace packages remain private; no npm package or permanent service deployment is included. -
Gate 2: explicit captureMode, shared cancellation/deadlines, full Retry-After waits and persisted Monitor cooldown; actual process recovery/claim races, baseline/fencing, A/B/A/B, conditional-cache body and multi-Monitor isolation checks passed.
-
Gate 3: durable HTTPS destinations/delivery worker, lease/fencing, retry/dead-letter, same-event replay and a transactional deduplicating receiver. Real HTTPS ACK-loss/restart experiment passed; its temporary endpoint is stopped.
-
Gate 4: Monitor/Delivery SDKs, examples, install/restart documentation and a sanitized agent clean-install record. Independent human acceptance remains pending; Gate 5 external two-week and repeat-use validation has not started.
-
Native Crawl results expose paginated
/v1/crawl/:id/pagesand/v1/crawl/:id/errors, persistent cancellation via/v1/crawl/:id/cancel, and matching SDK/MCP operations. -
Roadmap calibration: B1/B2 and C1 remain in_progress; C2 Monitor/Delivery MCP and guided first use, plus C3 unified process management and remote HTTPS URL MCP are the next unimplemented slices. Existing B3/B4/C4 gaps remain open.
-
API listen is loopback by default (
hostname: 127.0.0.1).--hosted --tokenis the public mode: bearer auth, private/metadata SSRF deny on seed + redirects, 10 MB body cap, crawlmaxPagesdefault 100. -
Local mode still allowlists loopback/RFC1918 so fixture servers work. API callers cannot widen that list.
-
Crawl scrape or store errors write task/attempt
failedinstead of leavingrunning. Resume with no contentful checkpoint reseeds the seed URL. -
HTTP and local browser fills
rawBodySha256. Same body on a later URL isduplicate, not a crawl-stoppingloop_detected. -
A 200 challenge page with extractable prose is
blocked, not success. Decisive challenge evidence (vendor header / Cloudflare plumbing / interstitial copy pair) is consulted after extract; an embedded widget on a real article is not. -
Ladder runs now expose task-level execution accounting: every attempted channel remains available alongside
channelsTriedandladderTrace; unknown cost, token, or wire-byte measurements staynullinstead of being treated as zero. -
Browser-rendered DOM size is not reported as
bytesWire; browser paths usenullwhen actual network transfer bytes cannot be proven.artifacts: []means this run produced no screenshot or DOM artifact. -
Multi-page crawls use bounded workers, enforce Frontier host concurrency and robots crawl delays, reuse channels and routing history within an API engine, and reuse browser processes while keeping fetch contexts isolated.
-
Earlier foundation review baseline:
main@6965168after PR #14 and PR #15 were merged; current freeze evidence is linked in Gate 2–4 acceptance.
0.3.0 — 2026-09-18
Programmable scrape and crawl. Same runner as the CLI.
- REST:
POST /v1/scrape,POST /v1/crawl(202 + taskId),GET /v1/crawl/:id. Responses areFetchResult/CrawlReport. - TypeScript SDK (
@w2l/sdk, MIT):W2L.scrape/W2L.crawl/W2L.getCrawl. Server remains AGPL. - MCP stdio server (
@w2l/mcp, MIT): toolsscrape,crawl,get_crawlover the REST contract. No OAuth, no resources. - Firecrawl v1 shim (
/fc/v1/scrape,/fc/v1/crawl): snapshot 2026-09-18, maps onto the native contract. Challenge pages are not success; no fire-engine; resume defaults to refetch. Not a compatibility layer. - 1000-page kill/resume probe: SIGKILL at 27 pages, resume to 1001 unique URLs, 0 lost,
cached=0. - Workspace packages versioned
0.3.0.
0.2.0 — 2026-09-18
Crawl composes scrape. Checkpoint from day one.
w2l crawl <url>with--resume,--use-cached,--max-pages,--headed(browser arm only; CI stays headless).@w2l/runtime:TaskStore(memory + SQLite next to the task dir), frontier, orchestrator.- Checkpoint is
task → attempt → stepat URL granularity. Caller-generated UUIDs; repeat writes are idempotent. - Resume restores the queue. Default is refetch;
--use-cachedis the only skip-fetch path and marks cached pages. - Link harvest from the full document after extract, before HTML is dropped. Nav links are kept; markdown stays chrome-free.
- HTTP arm honours robots.txt with the same semantics as the browser arm.
- Loop stop: two distinct canonical URLs with the same
rawBodySha256→loop_detected. - Page / time / cost / token budgets can set
budget_exceeded. - Workspace packages versioned
0.2.0.
0.1.0 — 2026-09-18
First product-shaped cut of the identity ladder.
- Anti-bot is a coverage ladder (ADR 0004), not an in-tree circumvention engine.
- L0 identity bundle: UA, Client Hints, locale, timezone, viewport must agree; fail closed.
- Product HTTP arms send that bundle (
standarddefault;researchis a declared bot). - Ladder refuses channels with a missing or contradictory identity before
fetch. w2l scrape <url>is the user entry (w2l-fetchremains an alias).- Extract-tf emits Markdown after extraction, not raw HTML.
- Provider lane measures vendor identity and does not inject ours; HeadlessChrome / research-as-Chrome / UA-hint mismatch are not success.
- Changing IP or session does not change identity (
identityForRoute). - Workspace packages versioned
0.1.0.