Browse docs

Octocrawl / Changelog

Changelog

What changed in Octocrawl, newest first. Unreleased is on main and on this site, and not yet in a published package; each numbered version is on npm and PyPI.

Unreleased

  • octocrawl.dev has one top navigation on every page (apps/public-web/scripts/siteNav.mjs): Product (what Octocrawl does, the ways to use it, and connecting an agent over MCP in one line), Docs (the guide and reference pages), Free tiers and Changelog, with GitHub and Try it beside them. The menus are <details>, so they open without a script; docs-assets/nav.js adds hover, one menu at a time, Escape and, on narrow screens, a Menu button. A menu prints open in the site’s glyph style: the card appears a text line at a time under a dotted print head, its entries in turn, and its monospaced words decode from noise; with reduced motion it opens at once. The new Changelog page (/changelog/) is built from CHANGELOG.md, its links to repository files pointing at GitHub, and is in the sitemap and llms.txt; .gcloudignore lets CHANGELOG.md into the Cloud Build upload.
  • A page with little text (at most 4,000 characters) that still shows its data is on the way (a visible element marked aria-busy, a visible short “Loading…” or “Fetching results…” text, not on a button or a link) is waited for up to 8 s in all on the browser rung and on a provider’s, instead of being read after the usual settle of at most 1.5 s; a request with waitFor or actions keeps its own wait instead. The trace records the wait (loading_wait), and a page read as content while it still showed the sign carries a page_still_loading warning. A provider page read that way in the PA 4 Steel run had answered success with “Inventory Search Results Fetching…” for its content.
  • The published packages point at the site: npm octocrawl, @octocrawl/cli, @octocrawl/sdk and @octocrawl/mcp get https://octocrawl.dev as their homepage, the GitHub issues as bugs and search keywords; PyPI octocrawl-client gets the site, the docs and the issues as project URLs and keywords; each package README names the site, and @octocrawl/mcp’s README offers the hosted server (https://mcp.octocrawl.dev/mcp) for trying it without installing. They take effect with the next release.
  • octocrawl.dev, from the 2026-10-09 site check: the home page carries its FAQ as FAQPage data, built from the same list the FAQ shows; every sitemap address has a lastmod (the day its content last changed, kept in apps/public-web/scripts/docsPages.mjs and checked against git by a test); the hashed scripts and stylesheets, and files asked for with their ?v= content version, are cached for a year (immutable) instead of an hour; the Codex and Cursor MCP links on Connect MCP point at those docs’ current addresses; the docs introduction has its own heading instead of the home page’s; the header’s Docs link no longer carries the outside-link arrow; and the home page description fits in 155 characters.
  • The ladder goes on to its next rung for more failures a stronger rung may answer (ROADMAP PA item 4): a 403 or 405 answered without a gate it recognises, a page a browser rendered with no main content it could verify, and the HTTP rung’s refused connection (its next rung is the browser, whose own network failure ends the run); ladder_step with escalate says which. When the stronger rungs fail without a page, the page stepped past is the answer (ladder_evidence_kept). On the PA 9 gap tasks such failures had ended the run before the browser or the provider was asked. A timeout, a rate limit (429), an error status that is the page’s own answer (404, 410, 5xx), a failure of the saved login’s rung and a failure after content an earlier rung found still end it.
  • The home page shows what Octocrawl does beyond the one-page preview, how to start and what is free, in three sections between How it works and the FAQ. What it does: six cells say what runs on hosted Octocrawl (scrape, map, the Evidence Record) and what runs only on your computer (crawl and batch, Monitors, your own Chrome), each in two or three short points, with a filter; a cell pointed at fills with orange from the left and its words turn white. Get started: a Use it from window with CLI, MCP, TypeScript, Python and REST tabs (the CLI, TypeScript and Python lines were run as written on 2026-10-09; the MCP command is the one Connect MCP marks verified on 2026-10-06), under a heading where a glyph octopus drawn after the brand mark lives in a small sea: it swims, rests, sleeps, changes colour, waves, peeks at the heading, blows bubbles, reads a page that drifts away as Markdown, chases or greets a passing fish, squirts ink and flees from a pointer that comes close, comes to a click, and keeps clear of the words. Free tiers: the four allowances the services enforce (five previews a visitor a day, 20 pages a day per address without a key, 1,000 to start with one, no daily limit on your own computer) pass one at a time over the glyph Earth artwork (scene-earth.webp), the scroll settling on a tier when it comes to rest; each tier lights that many of the planet’s own painted marks (5, 20, 1,000, or every mark on the planet), found by a worker the way the hero finds its glyphs. The glyph band under the hero rolls with two slow waves. Motion runs only while on screen; with reduced motion everything shows still, and without a script every code line and every tier still shows.
  • Every paid provider call is on the page’s Evidence Record (ROADMAP PA item 4): access.paidCalls lists, in order, the provider, the rung, the ADR 0005 capabilities its session was created with, the price ceiling reserved, what the spend ledger charged, the price the provider stated (null when none), and what Octocrawl made of the page the call returned by its own checks (outcome, reason; null when it returned none), never the provider’s word, with answer marking the call whose page is the record’s. access.grant names the access grant they were made under by the SHA-256 of its text, its tier and its attestation time. A provider called with no budget is settled in an uncapped ledger of the run’s own, so it is recorded too, and a page keeps the calls of a read that did not become its answer: one given up for another egress, the run whose stopped page the person read in their Chrome, and a batch or crawl run that threw after a paid call. Both fields are optional in the v1 schema, so earlier records stay valid.
  • Provider spend under a ledger (ROADMAP PA item 4): an access grant’s tariffs give each provider’s prices per call and per hour and what bounds a call’s time (maxSessionMs, minBilledMs, billingIncrementMs), from which its price ceiling is computed; a provider without one is not called, and a per-GB price is refused (bandwidth cannot be bounded from here). A task’s pages, retries and providers reserve each call’s ceiling from one ledger before the call and settle it after (at the reported price, or at the ceiling when none was reported, a call that threw or was cut included), so concurrent workers cannot pass perRunUsd, and a resumed or appended run opens the ledger with what the task was already charged (stored on each attempt as chargedUsd); perRequestUsd now caps one page’s calls and a single scrape’s. Under a tariff a provider call opens a connection and a session of its own, created with the provider’s own timeout and its proxies off, and released when the call ends. A run stops at its cap (cost) instead of at the first unpriced provider page (cost_unknown). Answers add usage.externalCostChargedUsd (the run’s total) and a spend_settled trace event per call; externalCostUsd stays the exact cost or null. The ladder CLI takes the grant’s tariffs too.
  • The my-browser lane says more exactly why a page was not read: a page still on a site’s check when the wait ends is blocked with that check (a Cloudflare block page answered 403 had been failed/timeout), and a Chrome that refuses a command (a tab it will not open) is failed/connection_error, not timeout. A page the site leads elsewhere on it each time Octocrawl takes the tab back (www.linkedin.com/mynetwork/ in the 2026-10-09 acceptance run) is failed/redirect_limit, not timeout. Its warning says the person got through only when the page showed a check and they acted in its tab, not when the check cleared by itself. A page read in the person’s Chrome counts as handed_to_person only when it showed a check and they acted in its tab; a check that cleared by itself is user_browser.
  • The main content of a Wikipedia article leaves out its hidden categories (.mw-hidden-catlinks, the maintenance categories the site’s stylesheet hides from every reader), which ended the Markdown after the visible categories; those stay, and so does the whole page’s copy with onlyMainContent: false (#289, the other half of the Wikipedia main-content gap).
  • The main content of a page leaves out the navigation a site lays out inside its content column: an element with role="navigation", and MediaWiki’s portlets and menus (.mw-portlet, .vector-menu, the language menu #p-lang-btn), which Wikipedia’s Vector 2022 skin puts in <main> beside the heading. A scrape of a Wikipedia article with onlyMainContent (the default) now opens with the article, where it opened with the interlanguage list (“22 languages” and a link per language) and the page tabs; the “See also” portal box goes too. onlyMainContent: false and includeTags keep them, as they keep <nav> (the Wikipedia main-content gap seen on parity case A16).
  • A list continued in the person’s Chrome counts its items again: actions.lists[].items is the last page’s count and itemsRead the kept pages’ count plus each page the person showed (the reader’s tab counts the step’s itemSelector on every page), where before both were null after a continuation. itemsRead is a sum only when the addresses show no page can be counted twice (the kept pages each at its own, the check’s page and the pages shown at none of them, the check not at the list’s own address) and no page shows again what another shows (its items’ whole text, as the list merge tells it, since a result set tied to the session that made it comes back at new addresses), and stays null otherwise (also when the step’s itemSelector uses a form the extractor does not take, such as :has(), so pages cannot be compared), as for a pager that reloads its items in place at a redirected address. The counts are not added to actions.scrapes. Found by the real-site run on indeed.com (2026-10-09): pages 2 to 5 went through, and the record could not count them.
  • A page handed to the person whose own script rewrites its address once it has come is still the page asked for: the handoff reader hears the tab’s navigations (Page.frameNavigated, Page.navigatedWithinDocument) and takes an address the page set in place by itself, on the document that came at the page asked for, before the person clicked or typed on that document (its user activation, asked of Chrome in Octocrawl’s own world when the address changes), when it keeps the path and the value of every parameter both name, and drops no parameter whose value is a number (a page or an offset). Found on indeed.com (ROADMAP PA item 3’s real-site run, 2026-10-09): its list page drops its paging token pp and adds the job shown as vjk, so the reader took the real page for one elsewhere, took the tab back twice and gave up. A page parameter that changed (another page of a list), or an address the person’s own click moved in place (Next, a sort), is still not the page: the tab is taken back.
  • Where a pool egress leaves from (ROADMAP PA item 3, “every result records … the location”): with W2L_EGRESS_ECHO_URL beside W2L_EGRESS_PROXIES, each proxy is asked the echo URL through itself (once, again after it failed or after 10 minutes), and the Evidence Record’s access.egress gains exit: { ip, country, observedAt } on every page read through it (an egress_exit trace event; the country only when the service gives a two-letter one). Without the echo URL, when it did not answer, or for the environment proxy, exit is null. Optional in the v1 schema file. The server refuses the echo URL without egress proxies or when it is not http(s).
  • A list that stopped at a check goes on in the person’s own Chrome (ROADMAP PA item 3): a batch whose only step is paginate with an itemSelector is handed over for the items whose list stopped at a check; the page the check was on opens in their Chrome, they get through it and page on by clicking Next themselves, Octocrawl only reads that tab (a read-only script; it clicks nothing), and the pages its own browser read before the check are merged with those into one list, each page once (actions.scrapes[].by says who read a page; the list_continued trace counts the kept pages and names the lane that read them). The item becomes the whole list with actions.lists[].continued = { from, pages, by: "user_browser" } and the list_continued trace event; stoppedBy says how the reading ended (end once Next has been unusable on the last page read for 5 s, max, or deadline after 60 s without a new page or at the handoff’s waitMs, with list_not_exhausted); a person who got through but showed no page after the check’s within that time leaves the item stopped, for a later handoff to go on from. A page off the site, on a login path, with a password field or not answered 2xx is not read, as for any page read in their Chrome. The CLI prompt and the MCP hand_off_batch text say so. Any other batch with steps is still not handed over.
  • A paginate step stops at a check the site puts up where the next page should be (ROADMAP PA item 3): each page is put to the gate classifier’s decisive marks as it is read, and an interstitial (Cloudflare, a PerimeterX press-and-hold, a verification form) ends the step as challenge without being read, told or counted; a page of nothing but a CAPTCHA widget, which only the extractor tells from a page with little on it, is marked the same way by the page’s own verdict after the steps (a page with records on it, or with other content, is a page whatever widget it carries). The pages before it keep their records in list, the result is blocked with the check’s reason, actions.lists[].challenge names the page and the URL, the list_challenge trace event says what the gate saw, and list_not_exhausted says the list stopped at a check. In a batch the pages before it stay in the task’s checkpoint (going on from them once the check is handled is still to come); a page with no record on it is not kept there. Before, the check’s page was read as an empty page of the list and the checkpoint was cleared.
  • With egress proxies, a page whose request never went out (robots.txt or a policy refused it, a lockdown found no cached copy) no longer has its egress probed: only a page that failed on the network with no HTTP answer from any rung does. The probe reaches no site, but it was a CONNECT to the proxy for nothing (PR #247’s follow-up).
  • A task’s cookie session file and egress binding are removed before its terminal event goes out, so a webhook receiver or an events stream told of the end finds them gone; before, they went after the event (PR #242’s follow-up).
  • A list task cut at page N resumes there (ROADMAP PA item 3): a batch with a paginate step keeps each page the step reads in the task’s checkpoint the moment it is read, and when the batch resumes on the next start it passes over those pages along the site’s own Next links (pages have no address of their own in a click-driven list), reads the pages after them and merges every page once; a kept page is known again by its address, its items, or their links alone when a price or date changed meanwhile (rows with no links whose text changed are read again, repeat, and count twice toward maxPages); actions.lists[].resumed and the list_resumed trace event say how many pages came from the checkpoint. The lane tells each page through ExecutionContext.onListPage and takes them back through listResume. A paginate step that fails at page N (a click that never lands) now keeps the records of the pages it read in list, with the failure in actions.failed.
  • The Evidence Record’s access block names the egress a page left through and the task session it was read with (ROADMAP PA item 3): egress is { proxy, source, switchedFrom } (the proxy’s host:port, never its credentials; pool, environment or direct on a request that got a page response; the pool egress the task last moved off before the page was read, else null), null when nothing says: a vendor’s service, the person’s own browser, or a lane that stopped before any page request; session is { id }, never the cookies, null when none. Both are optional in the v1 schema file, so earlier records stay valid.
  • The browser-compatible transport has a default host list (ROADMAP PA item 2): a local server whose access grant names compatible_transport and that sets no W2L_COMPAT_HOSTS uses it for the five hosts G1’s two-window acceptance showed it helps (research/access/benefit-hosts.v1.json: fred.stlouisfed.org, www.idealo.de, www.investing.com, www.ironmountain.com, www.wayfair.com). W2L_COMPAT_HOSTS=none turns it off; naming hosts replaces the list. The list is a constant in the code, kept equal to the JSON record by a test, so every build has it.
  • The README, the Chrome connection hint, the CLI prompt and the MCP scrape tool say that while Chrome’s remote debugging is on (which the handoff, octocrawl login import and the my-browser lane need), every page sees navigator.webdriver as true, Octocrawl connected or not (seen on Chrome 153 with the chrome://inspect switch and on 154 with --remote-debugging-port), so a bot check may refuse the person’s Chrome; turn it off when done.
  • The my-browser lane waits 10 minutes, not 120 s, for the person to allow the sites in Chrome; a refusal says what Octocrawl’s page last answered; a batch waiting for the approval reports waitingForApproval: true; and the server logs each step of the approval (my_browser_approval on stderr). Three real-page runs had timed out at this step with nothing to tell a slow click from one not recognised.
  • One plain choice of how pages are reached (ROADMAP PA item 7): "access": "standard" (no rung that costs a third party), "enhanced" (what the server’s access grant of tier enhanced approves, its providers in mode standard too; refused by name without one) or "my-browser" (the person’s own Chrome, as lane: "my-browser") on scrape and batch, standard or enhanced on crawl; in the MCP scrape, batch_scrape and crawl tools and --access on the CLI. Batches and crawls keep the choice. The /fc shim maps Firecrawl’s proxy onto it (basic is standard; stealth and auto are enhanced) instead of refusing it.
  • The my-browser lane for batches (ROADMAP PA item 8): "lane": "my-browser" on POST /v1/batches, the MCP batch_scrape tool and octocrawl batch --lane my-browser read every page in the person’s own Chrome, one at a time. The person allows every site of the batch (host and port) once for the run, in the page Octocrawl opens there; a page on another site, or after Revoke, is not read. Never cached. Refused with a webhook and with maxConcurrency above 1, besides what a scrape on the lane refuses.
  • The my-browser lane (ROADMAP PA item 8), for one page: "lane": "my-browser" on POST /v1/scrape, the SDK, the MCP scrape tool and octocrawl scrape --lane my-browser reads the page in the person’s own Chrome on a server on their machine. After Chrome’s Allow, Octocrawl opens a page of its own there listing the site and the task; only the person’s click on Allow reading these sites lets it read that site without a further click, and closing that page or clicking Revoke stops it. Recorded as lane my_browser (a new value of lane), never cached. Refused by name on other servers and with actions, a screenshot, lockdown or a mode other than standard. Every Evidence Record’s access gains completion (unattended, authorized_session, user_browser, handed_to_person, or null when no page was read). Batches and MCP batch_scrape follow.
  • Managed sessions (/v1/sessions/*) no longer answer with the profile’s path on the server (profileDir) or a CDP endpoint (cdpEndpoint); every route returns the public fields only. A hosted engine refuses them all (409), in the engine itself as well as at the hosted gate, and makes no browser profile (ROADMAP PA, G4).

0.3.1 — 2026-10-06

The published packages (octocrawl, @octocrawl/cli, @octocrawl/sdk, @octocrawl/mcp, octocrawl-client) at 0.3.1: everything below since 0.3.0 on 2026-10-05, and @octocrawl/mcp now carries mcpName for the official MCP Registry.

  • The published @octocrawl/mcp carries mcpName: io.github.77777R7/octocrawl and the repository has server.json for the official MCP Registry (registry.modelcontextprotocol.io): the npm package over stdio, with W2L_API_URL and W2L_API_TOKEN, and the hosted remote https://mcp.octocrawl.dev/mcp. scripts/release-version.mjs keeps server.json’s versions equal to the packages’. Publishing to the registry takes a release that carries mcpName (0.3.1 or later), then mcp-publisher login github and mcp-publisher publish.

  • docs/hosted-api.md: the key and waitlist scripts need NODE_USE_ENV_PROXY=1 behind a proxy, since Node’s fetch ignores HTTPS_PROXY; and where to keep a freshly issued key.

  • Egress proxies (ADR 0005 egress_sessions, ROADMAP PA item 3): W2L_EGRESS_PROXIES names the operator’s own http(s) proxies. A batch or crawl keeps one for its run (kept in egress.json in its directory, so a resumed task goes on through it) and moves to the next healthy one only when the proxy itself fails (after a page got no HTTP answer, a probe finds the proxy does not answer or refuses its credentials), at most twice a run, with a new cookie session; the page is read again there and its trace says so (egress_switched). A block, a challenge, a 429 or a connection the site reset never moves it. Cookie session files are kept per route, so a task resumed on another route starts a new session; mode authed never uses the pool. A failed proxy cools down for 10 minutes; scrapes and maps take the next healthy one. Each proxy reads robots.txt and sitemaps for its own pages; per-host pacing stays shared. Needs the grant; refused on a hosted server; credentials are never recorded.

  • The public site and the README say hosted Octocrawl is live (PH phase 1, step 5). Connect MCP leads with https://mcp.octocrawl.dev/mcp: one picker for the hosted URL (Claude Code, Cursor and OpenCode verified on 2026-10-06; Codex from its docs), then how to add a key (--header in Claude Code, headers in Cursor and OpenCode, --bearer-token-env-var in Codex), the first task, what the hosted service does not do, then “Run it on your computer” with a second picker for the stdio server, and self-hosting. The home page’s Get code panel is “Use it in your code or agent”: the cURL listing calls api.octocrawl.dev, the MCP and prompt steps name the hosted URL first; the FAQ, the quota advice and the waitlist band (now “Ask for a hosted Octocrawl key”) follow. Limits gains a hosted allowance table, Privacy what the hosted service records, Terms and the acceptable-use policy cover it. The README’s MCP section is hosted, on your computer, self-hosted. llms.txt and every docs footer say the same.

  • Hosted Octocrawl: the health route is GET /health. /healthz on a Cloud Run URL is answered by Google’s front end with its own 404 page and never reaches the container (seen on the first deploy, 2026-10-06); the runbook’s check uses the new path.

  • Hosted Octocrawl, phase 1, the deployment (ROADMAP PH): Dockerfile.hosted-api and cloudbuild.hosted-api.yaml build the octocrawl-api image (Chromium included, task root in memory and swept); cloudflare/hosted-api-proxy/ answers on api.octocrawl.dev and mcp.octocrawl.dev and forwards to Cloud Run with the shared secret, redirecting only plain http; scripts/hosted/issue-key.mjs issues, lists and revokes keys (a key is printed once, its HMAC is the Firestore document id); docs/hosted-api.md is the runbook: deploy without traffic, move traffic under a tag, the Firestore TTL policy, the Worker, the outside checks, the rollback, and the two-week cost record.

  • Hosted Octocrawl, phase 1 (ROADMAP PH), the service: npm run hosted:api (packages/mcp/src/hostedApiCli.ts) runs the API’s hosted mode and a remote MCP endpoint at /mcp in one process. It serves POST /v1/scrape, POST /v1/map and the MCP tools scrape, map and scrape_product; every other route and tool is refused by name with a hint to run Octocrawl locally. A caller with Authorization: Bearer <key> is looked up by the key’s HMAC in Firestore (hostedApiKeys/{digest}: enabled, plan, dailyLimit, browser); a caller without one is keyless, counted by the address the Cloudflare Worker reports (proven by the shared secret), within W2L_KEYLESS_DAILY pages a day (20) on the HTTP lane alone (fastMode is pinned; a screenshot is refused by name). Every start consumes the caller’s and the service’s (W2L_SITE_DAILY, 1500) daily counters in one Firestore commit, the public preview’s pattern; over the allowance the answer is 429 quota_exhausted with Retry-After to 00:00 UTC; per minute, 10 starts keyless and 60 with a key. Every request is also limited per address (120 a minute) before any key lookup, the limiter and key cache remember at most 20,000 callers, a file a call reads is capped at 5 MiB, nothing is stored for reuse (storeInCache is pinned off) and scrape and map records and saved files are swept from the task root after 10 minutes, since Cloud Run keeps it in memory. W2L_HOSTED_STORE=memory runs it without Firestore for a local check. The deployment (Dockerfile, Cloud Build, the Worker routes, the key script) is the next change.

  • A batch’s or crawl’s cookie session (egress_sessions) now survives a restart: it is written to cookie-session.json in the task’s directory after every change (created 0600, replaced whole) and read back when the task resumes, with the same session id; the file is deleted when the task completes, fails or is cancelled. A run paused by shutdown keeps it.

  • The public site no longer contradicts the published packages (PH phase 1, step 1): the Get code panel, the Extract page guide and the reference say the API and MCP server come from npx octocrawl serve and npx -y @octocrawl/mcp, not from a repository checkout; the reference names @octocrawl/sdk and octocrawl-client instead of calling the SDK a private workspace package; the introduction drops its source commits and “local machine” narration and names the four MCP clients; the footer of every docs page, the client picker, llms.txt and the Connect MCP page say a hosted URL is coming (with the early-access link) instead of “hosted MCP is paused”; the Monitor guide says it runs from a checkout and that a hosted Monitor is a later phase. The early-access form takes connect-mcp as a trigger (/?from=connect-mcp#waitlist).

  • ROADMAP: hosted Octocrawl restarted as phase PH (decided 2026-10-06). The hosted API and remote MCP leave the Paused table: phase 1 serves scrape and map at api.octocrawl.dev and mcp.octocrawl.dev in hosted mode, keyless within a per-IP allowance and metered by credits with a key, in Firecrawl’s shape; batch, crawl, Monitor and PA’s access routes are phase 2. The local path stays free and complete. P4 prices hosted credit packs first; “Free core and Pro” says so. The Paused row becomes “A hosted browser cluster beyond Cloud Run’s instance cap” (hosted_browser_cluster in @w2l/http-core and ADR 0005 follow it). The research behind it is docs/launch/2026-10-06-hosted-feasibility.md.

  • Cookie sessions for batches and crawls (ADR 0005 egress_sessions, ROADMAP PA item 3): under a grant that names it, the cookies a task’s pages set are sent again to their site on its later pages, by the HTTP rung (each redirect hop included), the compatible transport and the browser, which starts from the session’s cookies and leaves its own there. Matching follows RFC 6265 (tough-cookie 6.0.2, a new dependency). The session lives in memory for one run of the task; values are never recorded (session_cookies names a random session id and counts), and no page read with a session is cached. Without the grant nothing changes.

  • The README links to octocrawl.dev at the top (the preview, the documentation and Connect MCP, each with utm_source=github so the site’s page events can tell these visits apart), its Quick Start no longer says the preview has no permanent URL, and “For local MCP use” starts from the published packages (npx octocrawl serve, then npx -y @octocrawl/mcp in the client); the checkout’s managed local service stays as the Monitor → HTTPS delivery path.

  • The public site’s Connect MCP page now starts from the published packages: step 1 npx octocrawl serve, step 2 add npx -y @octocrawl/mcp to the client, with the Claude Code command, the Cursor and OpenCode configs (each run here and connected) and the Codex command (its documented syntax; not run here). The first task is a scrape of the first-use page. The repository checkout’s managed local service (127.0.0.1:8791/mcp, w2l-local) moves to a short section at the end. The home page’s Get code panel says the same: the cURL and MCP listings start with npx octocrawl serve, and the prompt asks for “Octocrawl’s scrape tool” instead of w2l-local.

  • The Evidence Record states how each page was reached: a new access field gives the route (http, http_compat, browser, enhanced_browser, authed_browser, user_browser, vendor), the client that sent the requests and its version when known (undici, impit 0.14.5, playwright, patchright, the person’s browser, the vendor’s id), the HTTP transport’s browser profile, and the run’s third-party spend (0 when none was paid, null when a vendor stated no price). It is read from the page’s own trace, so batch items and crawl pages carry it too; a page served from the cache states the route and cost of the fetch it reuses. A result no lane produced (a rung the deadline cut, that failed, or whose identity was refused; a cache-only miss) has a null route and client rather than a guess. It is optional in the v1 schema file, so records written before it stay valid. Each ladder attempt in summary.attempts (full responses) also carries its place in the run (ordinal), when its rung was asked and answered (startedAt, endedAt), and a provider rung’s vendorId.

  • The public site’s automated flag on page events and preview outcomes now also covers monitors that do not call themselves bots (Dataprovider.com, DomainMonitor) and browser strings no one runs any more (iOS before 15, Chrome before 110). In the first week of logging, most page views came from such clients: they opened the page and never touched it, and the funnel counted them as visitors. Current browsers, Chrome on iOS and Edge included, are unchanged.

  • robots.txt by who chose the URL (decided 2026-10-05). On a local server a URL the request names (a scrape, a batch entry, the CLI’s URL list, MCP scrape and batch_scrape, /fc/v1/scrape) is fetched when robots.txt disallows it or cannot be read; robots.txt is still read and recorded, with a robots_overridden warning and the new Evidence Record field robotsDecision.overrideBasis (user_named_url, robots_override, ignore_robots_txt; optional in the v1 schema file). The links a crawl or map discovers and a Monitor’s re-reads still obey it, and a hosted server obeys it for every URL. A crawl or map takes ignoreRobotsTxt on a local server (CLI --ignore-robots-txt, MCP, and Firecrawl v2’s name on /fc/v1/crawl): a crawl fetches the pages and sitemap files robots.txt disallows, a map returns those URLs with robots: "disallowed" or "unreachable". A recorded robotsOverride now sets an unreachable robots.txt aside as well. robots.txt is matched with the product token Octocrawl added to every User-Agent: a group for Octocrawl (or w2l-research) is the site owner’s targeted opt-out, which a named URL and ignoreRobotsTxt do not set aside; only a recorded robotsOverride does. Crawl-delay pacing and the 429 cooldown are unchanged. A hosted engine refuses robotsOverride, robotsOverrides and ignoreRobotsTxt by name, the hosted MCP engine included.

  • The HTTP lane no longer reports a SHA-256 of an empty body when no response arrived (a refused connection): rawSha256 is null.

  • Repository cleanup (ROADMAP P0): the earlier product documents (PHASE1_ENGINEERING_NOTES.md, PRODUCT_PLAN_V2.md, PRODUCT_STRATEGY.md) and the Render Blueprint move from the root to docs/archive/; the hosted MCP pilot’s Render and WorkOS setup moves from docs/mcp-first-use.md to docs/archive/hosted-mcp-pilot.md, and its code (packages/mcp/src/host.ts, npm run hosted:mcp) is marked experimental, saying so when it starts. The README gains a table of the local ports (8787 API, 8791 local MCP, 8788 first-use webhook receiver, 8798 site preview).

  • Releases are published from a tag vX.Y.Z by the Release workflow, through npm and PyPI trusted publishing (no stored token, npm provenance), after the type check, the tests and the package install check. scripts/release-version.mjs sets and checks the one version the published packages share; scripts/check-packages.mjs reads the expected CLI version from its manifest instead of a fixed 0.3.0.

  • The README starts from the published packages (npx octocrawl, pip install octocrawl-client, @octocrawl/sdk, @octocrawl/mcp). A new Install check workflow runs npx -y octocrawl@latest scrape https://example.com from an empty npm cache on macOS, Windows and Linux (scripts/install-check.mjs, limit 5 minutes) and pip install octocrawl-client, each week and on demand.

  • Security, before the first publish:

    • A local server on an address other than 127.0.0.1, localhost or ::1 (--host 0.0.0.0, a LAN address, W2L_API_HOST) refuses to start without a token. Before, it answered anyone, other machines and rebinding web pages included, and fetched the person’s localhost and network for them.
    • Mode authed refuses executeJavascript (a script could read the session’s cookies and storage) and, on a batch, a webhook (pages read with the session are not sent to another address). A batch with a webhook is no longer handed to the person (its stopped items carry no handoff, and handOffBatch refuses it): a page read in their own Chrome is read signed in as them, and was sent to the webhook as a handoff page event. Click, write, press, scroll and the list steps stay available. The MCP tools say so, and tell the model to ask the person before using a saved login.
    • A job id that is not one the server issues (..%2F…) names nothing: crawl and batch routes answer 404 instead of opening or creating a task store outside the task root.
    • The loopback server refuses a request body that is not JSON (content-type: application/json), so a page on another local port cannot send one as a form or text; a POST without a body (curl -X POST …/cancel) still needs no type.
    • The one-off commands (octocrawl scrape, batch, crawl, map) ignore W2L_API_HOST: they listen nowhere.
    • Paid browser services are used only when named in W2L_VENDORS (browserbase, steel) with their key; a BROWSERBASE_API_KEY or STEEL_API_KEY in the shell alone no longer sends pages to them.
    • A task root the engine creates is readable by the person alone (0700) and holds a .gitignore, so a repository it sits in does not take in pages and job databases; the folder of the saved logins is 0700 too.
    • The CLI’s README no longer says every fetch declares its identity: the default mode sends the Chrome user agent and client hints, and --mode research names itself. The research user agent links to the Octocrawl repository.
  • The packages to be published carry the Octocrawl name: @octocrawl/cli (also as the unscoped octocrawl, so npx octocrawl scrape <url> runs it), @octocrawl/sdk, @octocrawl/mcp and the Python client octocrawl-client (import octocrawl_client; the name octocrawl on PyPI belongs to another project). The command is octocrawl and the MCP server octocrawl-mcp, and the CLI’s messages, the API’s hints and the MCP tool descriptions name octocrawl commands; the MCP server introduces itself as octocrawl. The workspace keeps its @w2l/* package names, the W2L_* environment variables and the .w2l/ directories. Nothing is published yet.

  • A page that is a list of items is now read as its content with the default onlyMainContent: true, not failed as empty_unverified (EXTRACTOR_VERSION is now extract-tf/14). Before, two kinds of list page failed. A list whose items carry no link, such as quotes with their authors, was not a listing of cards to the last-resort fallback, which also wanted a page with exactly one h1. A grid of cards that each declare a schema.org Product in microdata was routed as one product page, whose recommendation pruning then cut the cards. The extractor now takes the list the list format’s detection finds when no other region was found: at least three of its items must hold 40 characters of text each, all different, and the region widens to the last h1 before them. Such a page is still checked for a wall as one with nothing found (lastResort on the extractor’s output), so a login form with a list beside it stays blocked, and a handoff in the person’s Chrome keeps waiting at a captcha with a list beside it. Three or more outermost Product scopes of one tag and the same classes now route as a collection, unless the page declares a Product in JSON-LD, has a buy box or a lone h1 with a price shown outside the cards, or shows the cards as recommendations (an h1 inside one, a recommendation heading before them, or an element around them named for recommendations). Found by the real-site cases AC01, AC05, AC07, AC08 and AC09, which ran their steps and then failed on the read.

  • A rendered answer (the browser or a provider lane’s) that its own extraction found thin and unsure (300 main-content tokens or fewer, confidence 0.3 or less) now carries the low_content_yield warning, with an agent hint on what to pass (waitFor, actions, onlyMainContent: false); its status stands. The warning was the http lane’s alone, so IMF’s datamapper, whose figures are drawn by script, came back as success with 226 tokens of social and navigation links and no caveat. The 300 is set from the browser lane’s yield on the real-site set (record): of its 73 rendered pages, only a tag page whose extraction took its sidebar falls under it.

  • A step cursor this API did not issue (made up or cut short) on GET /v1/batches/:id/errors, /v1/batches/:id/items, /v1/crawl/:id/pages or /v1/crawl/:id/errors is now HTTP 400 invalid_request “cursor is not one this API issued”, as on the crawl status routes; it was a 500 internal_error.

  • A crawl’s or a map’s sitemap reader now keeps the Crawl-delay a host’s robots.txt declares for the request’s identity between its own requests to that host: a sitemap index’s files came the policy’s 250 ms apart. The delay counts from the reader’s last request to the host and is waited before it takes its turn on the origin scheduler, so it neither slows another job’s requests to the host nor waits for a gap in them; the pages after the files are paced by the crawl’s frontier as before. On www.cbs.nl (Crawl-delay: 1) its index and first file are now 1.0 s apart, 0.25 to 0.29 s before (record).

  • Lists and quotes are indented 32 levels deep at most, so the Markdown of a list nested thousands deep grows with its text, not with the square of its depth: each level indented every line inside it, and 2,000 nested <ol><li>a gave 6 million characters. The blocks of a list item or quote deeper than 32 levels are written as blocks of the 32nd level’s item or quote, with no marker or > of their own and a blank line between each two, so none runs into another (a line into a table below it, a line before --- into a heading); their text is kept, in order. The same 2,000 levels give 194,000 characters, and twice the depth gives twice the Markdown. Rendered with markdown-it, random pages nested 33 to 70 levels deep show every word of the page in order, with no code block, heading, table or cell the page does not have (2,122 of 2,122). Up to 32 levels of the Markdown’s own nesting nothing changes: the same Markdown and tables for 30,000 random inputs, 3,860 random pages nested up to 32 levels and 120 locally captured pages (a list written directly in a list is one level deeper in the Markdown, inside the item before it). EXTRACTOR_VERSION is now extract-tf/13.

  • Deep lists and quotes are written with less memory: instead of a prefix kept for each line at each level (the entry below), a block is kept as the tree of its items and quotes and written out once, each line getting the prefixes of the items and quotes it is in as it is written. Peak memory for 2,000 nested levels, less the 110 MB of an idle process: <ol><li>a 103 MB (276 MB as a string rewritten at each level, 203 MB with a prefix per line), a quote of a paragraph 170 MB (359 and 385 MB), a quote of a list item 315 MB (446 and 697 MB), list items of three paragraphs 241 MB (362 and 520 MB); the time stays as it was with the prefixes. The Markdown and tables do not change: the same for 120,000 random inputs and 120 locally captured pages, so EXTRACTOR_VERSION stays.

  • Lists and quotes nested deep are written in time that grows with their Markdown, not faster: each list item and quote wrote the text of everything inside it again to indent it, so the time grew with the cube of the depth while the Markdown (indented at each level) grows with its square. A block now keeps its lines, and a list item or quote adds its indent, marker or > to each line, written out once at the end. 2,000 nested <ol><li>a took 1.8 s and take 0.1 s; 2,000 quotes of a list item, 6.5 s and 0.4 s. The Markdown and tables do not change: the same for 120,000 random inputs of nested lists, quotes, code, tables and line-start characters, and for 120 locally captured pages, so EXTRACTOR_VERSION stays. Converting the 120 pages takes as long as before (4,310 and 4,302 ms, median of five).

  • Blocks nested thousands deep (<div>, lists, quotes, layout tables, an emphasis or <span> around blocks, <pre> content) no longer run the Markdown converter out of stack: 2,000 nested lists or layout tables, 4,000 quotes or <b><div>, or 8,000 <div>s threw RangeError: Maximum call stack size exceeded, which failed the page. The block walk keeps its place on a stack of its own, each level with what it writes once its children are (a list item’s marker, a quote’s >, a closing paragraph break), and the text of a <pre> is read the same way. A table’s own rows and cells are now found by looking through it once, not by searching all of its descendants and dropping a nested table’s: a table nested 5,000 deep took 3.1 s and takes 33 ms. The Markdown and tables do not change: the same for 60,000 random inputs of nested blocks, tables, templates, svg and math, and for 120 locally captured pages, so EXTRACTOR_VERSION stays. Converting the 120 pages takes about as long (3,678 and 3,734 ms, median of five).

  • Inline elements nested thousands deep (<sup>, <sub>, <span>, <i><b>, <code> and the like) no longer run the Markdown converter out of stack: 8,000 nested <sup> or <span>, or 4,000 <i><b>, threw RangeError: Maximum call stack size exceeded, which failed the page. The inline walk, the search for a block inside an element and the text of an element with hidden parts now keep their place on a stack of their own; 20,000 levels convert as 3 do. The Markdown does not change: the same for 20,000 random pages, their fragments and 120 locally captured pages, so EXTRACTOR_VERSION stays. Converting the 120 pages takes about 1% longer. Blocks nested that deep (lists, quotes, tables) still run out of stack.

  • Nested emphasis that ends in punctuation right before a letter (<i><b>"y"</b></i>z) is now written so CommonMark reads it as emphasis: the punctuation goes after the closing markers of both runs (***"y***"z); it was written ***"y"***z, which shows its markers as text. An emphasis run’s content that ends with another run passes on how that run is written before a letter, so the outer run moves the same punctuation after its own marker. Checked by rendering the Markdown with markdown-it and comparing each page’s visible text with Chromium’s: on random pages of two nested runs with punctuation at their edges and a letter after them, 0 of four runs of 1,500 differ (319, 117, 309 and 318 before); on random pages of emphasis, punctuation, code and links, 85, 72, 76 and 76 differ (86, 73, 76 and 76 before); no page differs that matched before. The Markdown of 120 locally captured pages does not change. EXTRACTOR_VERSION is now extract-tf/12.

  • Nested emphasis that starts with punctuation right after a letter (x<i><b>"y"</b></i>) is now written so CommonMark reads it as emphasis: the markers of both runs make one delimiter run, read by the letter before it, so the punctuation goes before them all (x"***y"***); it was written x***"y"***, which shows its markers as text. After a space, or at the start of a link’s text, nothing moves. Checked by rendering the Markdown with markdown-it and comparing each page’s visible text with Chromium’s on random pages of nested emphasis, punctuation, code and links: 103, 96, 95 and 106 of four runs of 1,500 differ (116, 109, 108 and 119 before); no page differs that matched before. The Markdown of 120 locally captured pages does not change. EXTRACTOR_VERSION is now extract-tf/11.

  • The Markdown converter writes emphasis, escapes and adjacent runs so CommonMark reads them as the page shows them (EXTRACTOR_VERSION is now extract-tf/10):

    • A <b>, <strong>, <em> or <i> that holds blocks keeps its emphasis on each paragraph in it, as a browser shows it; it was dropped.
    • White space at the edges of emphasis goes outside the markers, a full-width space included, and a backslash at the end of emphasised text is escaped, so the markers are read as emphasis ( **indent**, not ** indent**).
    • Emphasis next to punctuation keeps the punctuation outside the markers where CommonMark would not otherwise read them as emphasis (a"**x**"b); a <, & or &# so moved next to what follows is escaped, as it would make a tag or an entity with it.
    • A backslash that would escape the ] or ) closing a link or image is escaped ([C:\\](…)).
    • Text is escaped only where CommonMark would read it as Markdown: *, _ and ` that could open emphasis or code, [ and ] that could make a link, < that could start a tag, a backslash before punctuation, and what would start a heading, list, quote, rule or table at a line’s start (1. alone on a line gives 1\., which CommonMark would read as an empty list item).
    • Adjacent runs of one emphasis that are plain text are joined (<b>a</b><b>b</b> gives **ab**), in linear time, as is code; runs holding their own Markdown or a character that pairs across the join (<b><i>x</i></b><b><i>y</i></b>, <b>&lt;</b><b>span&gt;</b>) are written side by side, the second with underscores.

    These six fixes were written for the htmlparser2 parser and are carried over onto parse5 unchanged. Checked by rendering the Markdown with markdown-it and comparing, for every character that is not white space, its bold, italic, code and link with Chromium’s on random pages of emphasis, code, links, blocks and characters Markdown reads (with raw HTML read as HTML): 257, 279, 262 and 262 of four runs of 1,500 differ (779, 780, 782 and 784 before); no page differs that matched before. Of 120 locally captured pages, the Markdown of 55 changes (49 for their main content); their tables do not.

  • EXTRACTOR_VERSION is extract-tf/9: the entries down to the table span parsing below landed on main together, after extract-tf/8, so each version they name is that one step; a Monitor whose Markdown changes only through them records extraction_reprocessed once.

  • In a template’s content, a table’s text still pending when the input ends or a </template> closes the template is written after that token is done, as in Chromium, which queues the write: an option the token closes is copied into its select’s <selectedcontent> without the text, which then goes where the table’s rules put it (fostered out of the table, or kept in it when all whitespace). <template><select><button><selectedcontent></selectedcontent></button><option selected>w8<table>w10</template> gives a <selectedcontent> of w8 and the table, where it also held w10. At the end of the input the text is written once no template is left open, so an option around the templates, at a page’s or a fragment’s own level, still has it when it closes; with any other token the copy has the text, as before. Template content is not read by the Markdown, links, images, attributes or main-content selection, so only the html format shows this and EXTRACTOR_VERSION stays extract-tf/9. Against Chromium, with each input given its own time limit, random inputs of selects, tables, templates and formatting elements now match wherever Chromium gives a result (the one input per 5,000 that differed now matches), with no regressions; the 120 captured pages build the same trees.

  • Every option a copy into a <selectedcontent> holds is handled in turn as inserted, as Chromium handles each node of an insertion: one the copy of an earlier one already took out of the tree is still selected and copied, once, if it has a selected attribute, but is then not the select’s (never selected by default, and not copied into a <selectedcontent> inserted later); while it is copied the select has no selection of its own, so an option its copy holds is selected by default if it is the first enabled one left. Before, such an option was skipped, so a list box whose selected option holds selected options and <selectedcontent> elements kept the copy of the first of them where Chromium goes on to the last: <select size=3><button><selectedcontent></selectedcontent></button><option selected>A<span><option selected>B<button><selectedcontent></selectedcontent></button></option><option selected>Q</option></span></option></select> now shows Q. Against Chromium, with each input given its own time limit (some hang Chromium), random select inputs now differ only where they did before (a <selectedcontent> holding a table’s fostered text, 1 in 5,000 in one run), and hand-written cases of taken-out copies with and without selected, disabled first options and later <selectedcontent> elements match; the 120 captured pages build the same trees. EXTRACTOR_VERSION is now extract-tf/9.

  • An option’s or <selectedcontent>'s select is found up its ancestors, not the stack of open elements, so one fostered out of a table, or moved by the adoption agency, is read where it is, as in Chromium. An option fostered out of a table in a <selectedcontent> is the select’s even once a copy took the table out of the tree: <select><selectedcontent>w<table><option><option><option> leaves the <selectedcontent> empty, each option copied in turn taking out what was there, where the later options were left in it. Against Chromium, random inputs of the selects sets (options, optgroups, datalists, tables, templates and <selectedcontent> written freely, and realistic customizable selects) now differ in at most 1 of 5,000 per run, where 0 or 1 did before, with no input that matched before differing now; what still differs differed before too (a <selectedcontent> holding a table’s fostered text, a list box with several <selectedcontent>, where Chromium’s own results disagree). The 120 captured pages build the same trees. EXTRACTOR_VERSION is now extract-tf/9.

  • An option that holds a select, or a <selectedcontent> with options, is copied into its select’s <selectedcontent> elements as any other, where it was left alone; the options the copy holds are then the select’s, by the rules a parsed option follows. In the document, a copied option with a selected attribute is selected and copied in turn, which takes it out again, as in Chromium (<select size=3><button><selectedcontent></selectedcontent></button><option selected>A<selectedcontent><option selected></option></selectedcontent></option></select> leaves the first <selectedcontent> empty); in a template’s content and in a fragment the option is only copied. Each such copy is of an option nested deeper, so this always ends, at the deepest selected copy; a chain deeper than 100 is parsed by linkedom instead, and each step up the tree it takes counts against the parser’s budget. Chromium itself gives no result for some of these: its DOMParser hangs on such an option in a select that is not a list box, and its renderer crashes on some list boxes with several <selectedcontent>. Against Chromium, random inputs that write options and <selectedcontent> freely into selects now differ in 0 or 1 of 5,000 per run (0 to 6 before), with no input that matched before differing now; realistic selects still match in every case, and the 120 captured pages build the same trees. EXTRACTOR_VERSION is now extract-tf/9.

  • An option written in a select’s <selectedcontent> is read as Chromium reads it, where it was left alone. It is one of the select’s options, and a copy into the <selectedcontent> that holds it takes it out of the select. In a page, Chromium copies an option as soon as it is selected, while still empty, so <select><selectedcontent><option>A</option>z</selectedcontent></select> gives <selectedcontent>z</selectedcontent>; when the selected option is taken out, the first enabled option left is selected without a copy (a <selectedcontent> written later takes it), and parts taken out of the tree are no longer the select’s. In a template’s content and in a fragment the copy is made only when the option closes, so the same input gives Az, and the next option is then selected. An option in an <optgroup> in another <optgroup> is not one of the select’s, a select in another select (reached through a table) copies nothing, and a <selectedcontent> in a <datalist> is still the select’s, as in Chromium. Against Chromium, random inputs that write options, optgroups and <selectedcontent> freely into selects now differ in 0 to 7 of 5,000 per run (40 to 45 before), with no input that matched before differing now; the realistic set matches in every case (one <selectedcontent> in a <datalist> per 5,000 differed before in one run), and the 120 captured pages build the same trees. Each visit to a <selectedcontent> and each element walked in an option counts against the parser’s budget, so a page that would make a select re-copy many times falls back to linkedom. EXTRACTOR_VERSION is now extract-tf/9, as includeTags selections and microdata see the changed copies.

  • A <select>'s <selectedcontent> elements hold a copy of its selected option, as in Chromium (the standard’s customizable select): the option with the selected attribute (the last one), or else the first that is not disabled (none in a list box, size above 1; no copies in a multiple select). The copy replaces what the <selectedcontent> held and is made when the option closes (also at the end of the page); in a page, a <selectedcontent> written after the selected option takes it when it is inserted, in a template’s content and in a fragment it does not. Every copied node counts against the parser’s element budget (comments now count too), so a page that would copy a large option into many <selectedcontent> elements is parsed by linkedom instead. An option written in a <selectedcontent> or in another option, and an option that holds a select or a <selectedcontent> holding an option, are left as they were: Chromium reads them through a chain of removals and resets this does not follow. The main content’s Markdown is unchanged (a select is left out of it), and links and images drop the repeats, but includeTags matches the copies as a browser’s DOM would (includeTags: ['.price'] over a select’s options now also gets the selected option’s price from its <selectedcontent>), and so do product facts read from microdata, the attributes and the html formats. Against Chromium, random pages and fragments of selects with <button><selectedcontent></selectedcontent></button>, selected and disabled options, optgroups, datalists, templates, links, images, multiple and size now match in every case (28 to 38 of 5,000 per run differed before), and inputs that write options into <selectedcontent> differ no more than before; the 120 captured pages give the same output. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • <select> is read by the HTML Standard’s rules since July 2025 (whatwg/html#10548), as Chromium has read it since version 134. The “in select” modes are gone: a select and its options hold what is written in them (<div>, <b>, <p>, tables, svg, links, images), where they kept only the text and <option>/<optgroup>/<hr>. A <select> bounds every scope, so a <p> or <li> written in it does not close one outside; </select> closes what is open in it; an <input> or a second <select> still ends it, a <textarea> or <keygen> no longer does. Markdown skips the select itself, but what follows it can change (<p>a <select><b>x</select>b</p> is a **b**, not a b), and links and images in a select’s options are now collected. Against Chromium, random pages and fragments of select, option, form-control, table, template and formatting tags now match in every case (1,480 to 1,530 of 5,000 per run differed before); the 120 captured pages give the same output. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • An end tag in svg or math is matched to an open element by its exact name, as in Chromium. In svg it first takes svg’s spelling (</foreignObject>, </clipPath>, </linearGradient>, </textPath>…); in math it keeps its own. An end tag that meets an HTML element first is read by the HTML rules, where a name svg respelled matches nothing, so it closes nothing. The standard (and parse5) compares names lowercased, so </foreignObject> in an svg closed an HTML <foreignobject> (or <clippath>…) or math’s around it, and what followed left the svg or math, where Chromium keeps it: in <p>a<lineargradient><svg></lineargradient>b</p> the b is in the svg, which shows no text, so the Markdown is a, not ab. Two more places where parse5 also took an svg or math element for an HTML one now read HTML elements only, as the standard and Chromium do: an end tag written in HTML inside svg or math (a </desc> or </mi> closed the svg <desc> or math <mi> around it, and what followed left it), and the insertion mode chosen after a table, cell or select closes (an svg <tfoot> made the rest be read as a row group’s, so a second <table> was dropped). Against Chromium, random pages of table, template, svg and math tags (with foreignObject, clipPath and linearGradient in either case) now match in every case (36 to 50 of 5,001 per run differed before), and random fragments in every case but one written as <body>…</body>, which is read as a selection; the 120 captured pages give the same output. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • A <title>, <base>, <basefont>, <bgsound> or <noframes> at a template’s top level switches it to the body’s rules, as in Chromium, which reads every start tag there by the head’s rules only for <link>, <meta>, <script>, <style> and <template>. The standard reads these five by the head’s rules too, which left the template’s mode as it was, so rows, cells and columns after them were kept where Chromium drops them and keeps their text. A fragment (main content, a selection) is read as a template’s content, so its Markdown changes when its top level has such a tag before rows: <title>Rows</title><tr><td>a</td><td>b</td></tr> is ab, not a and b as paragraphs. Against Chromium, random fragments and pages of table, template and head tags now match in every case (788 to 850 of 5,000 fragments and 236 to 303 of 5,001 pages per run differed before); the 120 captured pages give the same output. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • A fragment (main content, a selection) that starts with a <col> ends as in Chromium’s template.innerHTML: a table’s text still pending at its end, in a <template> in it, is dropped, where the standard writes it. <col><template><table>abc is <col><template><table></table></template>; the formatting elements the text re-opens before the table are still re-opened. Such text is always in a nested template, which no output reads, so no output changes and EXTRACTOR_VERSION stays extract-tf/9. Against Chromium’s template.innerHTML, random fragments of table and template tags now match in all 50,000 of ten runs (1 to 6 of 5,000 per run differed before).

  • A <form> in a <template> is read as Chromium reads it, where its parser differs from the standard (and parse5). A <form> written in a table inside a template is kept where it is written and closed at once: Chromium drops such a form only when a form is open outside any template, the standard whenever a template is open. A </form> in a template closes its form as any other end tag closes its element, so not past a <p>, <div>, <li> or other special element still open in it, where the standard closes those first: <template><form><p></form>x keeps x in the paragraph. Neither becomes the page’s form. Against Chromium, random pages of template, table and form tags now match in every case (11 to 18 of 1,500 per run differed before); with svg and math as well, what still differs is Chromium’s </foreignObject> past an HTML <foreignObject>. The Markdown, links, images and attributes leave template content out, so no output changes and EXTRACTOR_VERSION stays extract-tf/9.

  • A table’s end tags in a <template> that is in a table stay in the template, as in a browser. parse5’s table scope stopped only at <table> and <html>, where the standard’s stops at <template> too, so a </table>, </tr> or row group end tag in such a template closed the cells, rows, groups and table outside it, and the template’s text became page text: <table><tr><td>a</td><template><td>hidden</td></table>x</template><td>b</td></tr><tr><td>c</td><td>d</td></tr></table> gave a one-cell table followed by xbcd instead of the table a | b, c | d. Against Chromium, random pages of table and template tags now match in every case but two kinds parse5 and the standard read alike: a <form> in a table in a template, which Chromium keeps, and Chromium’s newer <select>. The 120 captured pages have no <template> and give the same output. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • The trees of the last two pages parse5 built are kept until the next task, not one. The browser lane reads two versions of each page by turns, the page as rendered (main-content selection, the Markdown, tables) and its body as received (links, images, attributes), and each read pushed the other’s tree out: with onlyMainContent: false it built the rendered page three times and the body twice. Run as the browser lane runs them, the 120 locally captured pages, with onlyMainContent: false, images and tables, now build 234 trees instead of 576 (about two per page, one for each version) and take 15.5 to 16.6 s (23.6 to 25.8 s before), with the same output; with the main content only, each version was already built once (234 trees both times). A loop over pages keeps two trees at most.

  • The tree parse5 builds for a page is kept until the next task, so a page read several times in one request (main-content selection, the Markdown, tables, links, images) is parsed once and copied into a new linkedom document each time. Main-content selection, both Markdowns, tables, links and images of the 120 locally captured pages together take 17.1 to 17.5 s again (25.6 to 30.3 s with a parse each, 17.2 s before parse5), with the same output. Only one tree is held, and only until the next task. One page’s own parse still takes about 1.7 times what linkedom’s took: that is parse5’s tokenizer and the garbage it makes; building the linkedom nodes straight from parse5 (a tree adapter) instead of copying was not faster (3.7 s against 3.6 s for the 120 pages), as creating linkedom nodes costs the same either way.

  • Pages are parsed by the HTML standard’s tree construction: parse5 builds the tree, which is copied node by node into linkedom (dom.ts parse, used by main-content selection, links, images, attributes, adapters, the Markdown and tables). htmlparser2, linkedom’s own parser, built misnested markup otherwise than a browser, which the table tag pass and rebuild of the entries above patched for tables only; both are removed. Now formatting elements are reopened as a browser reopens them (<b>1<p>2</b>3</p> gives **1** and **2**3), a page written without <html> or <body>, or with content after </body>, keeps all of it, and a <tbody> opens where a browser opens one. A fragment, such as the main content, is read as a <template>'s content, so one row or cell of a layout table stays one; a selection given as <body>…</body> is read without its <body> tag, whose attributes are kept, so includeTags: ["tr.athing"] keeps its rows. A <noscript>'s content, text to a browser running scripts, is kept as its elements, so the images and links in it are still the page’s. A page whose reopened formatting elements would outgrow its tags (four elements, or twenty element-name reads, per <: the standard reopens every <b> a block closed in each block after it, so 3,000 differently attributed <b> before 3,000 paragraphs, 56 KB, would make 9 million elements and run out of memory) is parsed by linkedom instead; the 120 locally captured pages use at most 12% of that budget. Three changes to parse5 8.0.1 (pinned): in a row it closed the row at a </tbody>, </tfoot> or </thead> whose group is not open, where the standard and Chromium ignore the tag; it moved a node’s children one by one with a linear search for each (a page of 200,000 lines without a doctype, or 80,000 under a misnested <b>, took seconds), and now moves them together; and a tag of more than 256 attributes (pages have at most 47; parse5 and linkedom check each against those before it, so 100,000 took half a minute) sends the page to linkedom. Checked against Chromium’s tree on random whole pages of misnested tags: every word has the ancestors Chromium gives it in all of 10,500 pages of 7 runs (formatting elements, tables, svg and math, templates, raw-text elements, pages without a doctype, a bare doctype, or no <html>), where 365 to 1,130 per run of 1,500 differed before; the table oracles of the entries above match as before. Still different: Chromium’s newer <select> parsing (content in a select, which the Markdown skips), </foreignObject> in svg past an HTML <foreignObject> (Chromium ignores it, the standard closes it), and misnested markup in a <template> in a table (parse5 differs from the standard there; 12 of 3,000 random sequences). On the 114 HTML pages of 120 locally captured ones the Markdown, tables, main content’s Markdown, links and images are unchanged (the other 6 are gzip data saved as .html, whose NUL bytes are now dropped); the html format gains the <tbody> a browser opens (42 pages) and keeps what follows </body> in the body (109 pages). Parsing takes about 1.7 times as long (3.1 s against 1.8 s for the 120 pages, 102 MB), selection plus both Markdowns about 1.35 to 1.5 times, and extraction (whole and with a selection), the three Markdowns, tables, links, images and the html formats together 1.7 times (46.1 s against 27.1 s). parse5 8.0.1 is a dependency of @w2l/extract-tf, and htmlparser2 no longer is. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • svg and math in tables are read as a browser reads them where the tag pass still differed. Integration points are now per namespace: svg’s <foreignObject>, <desc> and <title> and math’s <mi>, <mo>, <mn>, <ms> and <mtext>, so a <td> in math’s <foreignObject> or svg’s <mi> is theirs, not a cell. Closing an svg or math writes out the end tag of every svg and math element open in it, innermost first, so htmlparser2 does not close an outer <math> for an inner one, nor a cell for an svg’s own <td>; such end tags never carry the pass’s internal names (it wrote </^foreignobject>, which htmlparser2 kept as a comment). htmlparser2’s implied closes inside svg and math (a <tr> there closes a <tr> before it) and its view of a self-closing slash (it ignores one under an element named mi, title, … in any namespace) are followed, writing out the end tag where it would keep the element open. An end tag in svg that closes no svg element, of a name svg writes in camel case (</foreignObject>, </clipPath>, …), is ignored, as Chromium ignores it. And an end tag of an element other than a block, list item, heading, <p>, <form> or formatting element no longer passes a special element (<div>, <p>, <li>, …) in a table: <td><span><div>x</span>y</div> keeps xy together, as a browser does. Checked against Chromium on random tag sequences in a table, written as whole pages: every table matches in 2 runs of 2,000 with svg, math, <foreignObject> and <mi> (13 and 12 differed before), 4 runs of 2,000 adding <g>, <mo>, <mtext> and <desc> (7, 6, 6 and 4 before), 3 runs of 1,000 of loose svg and math tags (26, 28 and 25 before) and 3 runs of 2,000 with self-closing svg and math elements (0, 1 and 1 still differ, as they reopen <b>; 51, 52 and 48 before). Still different: an svg <title>, <style> or <script> holding tags (htmlparser2’s tokenizer reads its content as text, as for HTML’s; 1 of 2,000 such sequences), formatting elements a browser reopens after a misnested end tag (its adoption agency), and an HTML <title/> in a body (a browser reads the rest of the page as its text). 120 locally captured pages need no edit and give the same output; 20,000 random sloppy tables give the same output as before. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • A whole page (one with an <html> tag or a doctype) is parsed with its table tags outside any table ignored, as a browser ignores them: a stray <tr>, <td>, <th>, row group, caption or column after a table ended, or on a page without tables, is dropped and its content stays where it is written. Before, htmlparser2 kept them, so <table>…</table><tr><td>x</td><td>y</td></tr> gave the paragraphs x and y where a browser shows xy. Inside a <template>, svg or math they stay, and a fragment keeps them (given to the converter, or the main content the html format returns), as the main content of a layout table can be one of its rows or cells. A </template> for a template opened before a table it left open now closes the template and the table; before, it was dropped and the template held the rest of the page (<template><table><tr><td>x</template> followed by a table gave an empty Markdown). Checked against Chromium’s text of whole pages with a table followed by random stray table tags, text, <span>, <b>, <p> and tables: the text and its breaks match in all of 8,000 pages (183, 191, 195 and 324 of four runs of 2,000 differed before). 120 locally captured pages need no edit and give the same output; 20,000 random sloppy tables (fragments) give the same output as before. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • Pages are parsed with their table tags read as a browser reads them (dom.ts parse, so main-content selection, links and images as well as the Markdown and tables). htmlparser2 applies an end tag to the nearest open element of its name wherever it is, so a cell’s own </td> after a <td> written in it, or a </div> or </body> written in a table, closed cells, rows and whole tables further out: <table><tr><td>a</td><td>b</td></tr></body><tr><td>c</td><td>d</td></tr></table><p>after</p> lost the second row and the paragraph after the table. Before linkedom parses, htmlparser2’s own tokenizer now reads the tags and follows the stack of elements a browser has open from the outermost table in (with htmlparser2’s implied closes such as <p> closing a <p>, and svg and math content read as a browser reads it: a self-closing <svg/> closes, an HTML element such as <div> or an end tag that closes the cell ends the svg), and the source is edited only where the two differ: an end tag a browser ignores there is dropped, and so is an end tag with space after </, a comment to a browser; the end tags of the cells, rows, row groups and captions a browser closes at a table tag are written out, with the <tr> it opens for a cell written directly in a row group (htmlparser2 closed the <thead> there) and a <tbody> where rows directly in the table follow a closed implied one, so they stay two row groups; a <table> where rows belong ends the table there (inside a <template> a browser ignores it), and the table tags left of it outside any table are dropped; a column group ends at anything but a column. Outside tables nothing changes, and a <td> or <tr> outside any table stays, as main content can be one cell of a layout table. The converter no longer takes an svg’s <tr> or <td> for a row or cell, and moves an svg written at a table’s own level before the table. Checked against Chromium on random tag sequences inside a table, every table matches: 2 runs of 3,000 with cells, rows, row groups, captions, <div>, <span>, <p> and stray </body> (1,267 and 1,257 differed before), 3 runs of 3,000 adding column groups, <b>, <li>, <form> and </html> (643, 654 and 462 before), 2 runs of 3,000 adding <template> (509 and 511 before), and runs of <p>/<li>/<dd> implied closes (2,000; 308 before), end tags with space after </ (2,000; 758 before), self-closing <svg/> in cells (400; 71 before), svg icons with <title> and MathML <mi> in cells (400; 79 before) and icons with <title>, <desc>, <use/>, <foreignObject> and MathML (1,000; 275 before) and an svg in an svg’s <foreignObject> (400; 143 before). Still different: 4 and 26 of 3,000 sequences that also close tables (409 and 754 before), as their pages start <!doctype html><div> with no <html> or <body>, so the converter reads only that first element once the table and the <div> have ended (an older converter issue, open separately; written with <html><body>, none differ), and svg and math holding table tags and stray end tags (13 of 2,000, 394 before; 26 of 1,000 sequences of loose svg, <foreignObject>, math and <mi> tags, 244 before; see the svg and math entry above). Random tables with structure written in cells match in all of 1,939 and 1,765 (545 and 824 before). 120 locally captured pages need no edit and give the same output; 20,000 random sloppy tables give the same output as before. The tokenizer pass skips a page without <table>, keeps indexes into its stack so no step walks it, and takes 93 ms on 1.4 MB of 100,000 unclosed cells. htmlparser2 (already installed with linkedom) is now a declared dependency of @w2l/extract-tf. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • Tables (Markdown and the tables format) are rebuilt as a browser’s parser builds them where linkedom kept tags where they are written. A <thead>, <tbody>, <tfoot>, <tr>, <td>, <th>, <caption> or <col> inside a cell closes that cell there, also inside a <span> or <div> (htmlparser2 closed a cell only at a <tr> or <td> that is its direct child, so <td>a<th>b nested the <th> in the <td>). An element where a row group, row or cell belongs (a <div> around rows, a caption in a <tbody>, a cell directly in a <tbody>) and text there move as a browser moves them, the text before the table. A <table> there ends the table, and it and what follows come after the table. A <template> stays whole and its rows are not the table’s. Before, such a cell’s text held the rows written in it, and those rows were also rows of the table: <td>a<thead><tr><td>x</td><td>y</td></tr></thead></td> gave the cell a x y. Only tables built otherwise are rebuilt: 120 locally captured pages with 298 tables give the same Markdown and tables as before. Checked against Chromium on random tables: with structure written in cells (inside a <span> or <div> where it is a <td> or <tr>), the cells, caption and text before the table match in all of 1,905, 1,939 and 1,765 tables of three runs (1,782, 1,818 and 1,765 differed before); of 20,000 tables written with unclosed cells, comments and <form> wrappers, all 4,718 whose output changed now match. A <td> or <tr> written directly in a cell whose own end tag follows still differed (545 of 1,939 random tables): see the next entry. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • Table rows (Markdown and the tables format) are in the order browsers lay them out: the first <thead> first and the first <tfoot> last (an empty one included), wherever they are written, and a later <thead> or <tfoot> where it is written, as CSS lays out only the first as the header or footer. Before, rows kept the HTML order, so a <tfoot> written before the <tbody> became the GFM header row and a <thead> written after it was a body row (headerRows 0). A row inside a <div> in a <thead> belongs to that <thead>, and each run of rows directly in the table is its own row group, as the browser’s parser wraps each in a <tbody>. On 8,306 random tables with <thead>, <tbody>, <tfoot> and bare rows, the tables grid now matches the cell positions Chromium lays out in every one (2,169 differed before). EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • Table rowspans (Markdown and the tables format) cover the rows browsers give them. A rowspan now counts every row it spans, also one whose cells end before its column, where it used to wait for the next row that reached the column: <tr><td>a</td><td>a2</td><td rowspan="3">b</td></tr><tr><td>c</td></tr> followed by two full rows put b in the fourth row and shifted its last cell one column right. A rowspan no longer runs past the end of its row group into the next <tbody>, and two spans over one slot both count every row. On 4,047 random tables of <tbody> groups and bare rows, the tables grid now matches the cell positions Chromium lays out in every one (1,897 differed before). EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • Table spans (Markdown and the tables format) are read as browsers read them, by HTML’s rules for non-negative integers: the attribute’s leading digits (1.5 is 1, 2abc is 2), 1 when it has none or is negative, a colspan of 0 is 1, and a rowspan of 0 covers the rest of its row group (<thead>, <tbody>, <tfoot>, or the run of rows directly in the table), then capped at 1000 and 65534 as before. Before, they were kept as written: a fractional rowspan never ended and filled its column in every later row, so a 2,547-byte page with colspan="1000" rowspan="1.5" over 165 empty rows gave 166,498,389 characters of tables JSON and was not omitted (1,498,389 now), and a negative span counted negative characters, which gave budget back to the page’s 5,000,000. A rowspan of 0 was 1. Tables whose spans are absent or positive integers are written as before. EXTRACTOR_VERSION is now extract-tf/9; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • A <head> tag inside the body is ignored, as a browser ignores it, instead of becoming an element. A slash closes only void elements, so <head/> is an open tag, and linkedom, the DOM layer, put everything after it up to its parent’s end inside it, where the Markdown skips a head: htmlToMarkdown('<head/><p>Some text</p>') gave "", and a page whose article had a second <head/> after its <h1> gave the heading alone, in htmlToMarkdown and in main-content selection. Each such element is now replaced by its children wherever @w2l/extract-tf parses a page (Markdown, tables, main content, links, images, attributes, metadata); the document’s own <head> is unchanged. EXTRACTOR_VERSION is now extract-tf/8; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • A whole HTML document that leaves out <html> (and often <head> and <body>), as HTML allows, is now read as a browser reads it: the head elements before the first content go into <head>, everything else into <body>. linkedom, the DOM layer, has no implied elements: it made the first top-level element the document’s root and left the rest beside it, so <!doctype html><table>…</table> gave each cell as a paragraph and no tables entry, <!doctype html><title>T</title><h1>H</h1><p>x</p> gave T alone, and <!doctype html><head>…</head><body>…</body> gave empty Markdown; main-content selection gave "" for <!doctype html><body><p>x</p></body> and <body></body> for a lone table, and a longer body-only page’s region carried an inserted empty <head></head><body></body>. Every module that parses a page (Markdown, tables, main content, links, images, attributes, metadata) reads the same rebuilt document. A document with <html> is parsed as before. htmlToMarkdown and htmlToTables also take HTML that starts with <head followed by whitespace or > (after leading whitespace, a BOM included, and comments closed by -->; the empty comments <!--> and <!---> and a self-closing <head/> are not recognised) as a whole document, where before it was a fragment whose own <base href> was ignored; HTML that starts with <body> stays a fragment, since mainHtml can be the body itself, and is resolved against the base the caller passes. EXTRACTOR_VERSION is now extract-tf/7; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • w2l --out <dir> also writes results.csv: one row per page with its evidence, failed and blocked pages included, in the columns of the Python client’s to_pandas(include_markdown=False) (url, status, reason, final_url, fetched_at, http_status, lane, robots_decision, raw_sha256, markdown_sha256, extractor, source_commit, cache_state, cached_at), then markdown_file, the page’s Markdown file beside it. A value W2L did not observe is empty, never 0. New guides for researchers: docs/guides/url-list-to-csv.md and docs/guides/citing-web-data.md.

  • Reliability and throughput: W2L_WORKER_COUNT (an integer from 1 to 64, default 4) sets how many pages the API’s engine (w2l-api, w2l serve) runs at once, under the per-host limits; a crawl’s maxConcurrency may go up to it. npm run verify:batch-crash-1000 posts a batch of 1,000 URLs over 20 loopback hosts, kills the API with SIGKILL once 400 are done and restarts it on the same task root: 8 checks, among them one item per URL, every URL whose step was on disk at the kill fetched once, and at most one extra fetch per worker for the pages in flight. It runs in the Linux CI job. npm run bench:throughput measures both lanes through the API: on 2026-10-03, on loopback, the HTTP lane did 3,460 and 1,880 pages/min (p50 460 and 715 ms) and the browser lane 523 and 525 (p95 987 and 983 ms) (report). npm run verify:serve-smoke starts w2l serve, scrapes a page and a PDF, stops the server mid-batch and checks that the restarted server resumes the batch to one item per URL; a new windows-latest CI job runs it.

  • Python client (python/, PyPI name w2l, MIT; pip install 'w2l[pandas]' once published): w2l.batch(urls, **options) starts a batch on a W2L API (W2L_API_URL, default http://127.0.0.1:8787; W2L_API_TOKEN), waits for it, pages through every item and returns JobResult(task_id, report, items); .to_pandas() gives one row per page, failed pages included, with the evidence columns (url, status, reason, final_url, fetched_at, http_status as Int64, lane, robots_decision, raw_sha256, markdown_sha256, extractor, source_commit, cache_state, cached_at, markdown), an unobserved value missing rather than 0. scrape, map and crawl too; options in snake_case or camelCase; API errors as W2LError with the API’s code. A poll or page read survives a network error, a 5xx or a 429 (5 retries); a listing with more items but no cursor is an error. Tested with httpx.MockTransport (python/tests, run in a virtualenv; not in CI yet).

  • npm packaging: npm run pack:packages builds @w2l/sdk (MIT; ESM and CJS with @w2l/contracts bundled in, its declarations for import and require, no dependencies), @w2l/cli and @w2l/mcp (AGPL-3.0-only; one bundled file each, the third-party dependencies declared) into .w2l/pack/ and packs them. npm run pack:check installs the tarballs into a new project and uses each package against a loopback site. Each package declares only the dependencies its bundle imports (from esbuild’s metafile), and every entry guard but the bin’s is turned off in a bundle, so w2l run by its real path starts nothing else. @w2l/contracts is MIT (it ships inside the MIT SDK); @w2l/mcp is AGPL-3.0-only in the repository too. Each package has a README. The SDK now exports EvidenceRecord, PageTable, PdfPageMarkdown and PdfParser. w2l-mcp run through node_modules/.bin (npx) did nothing, since its entry guard compared a symlink with its target; it now compares the real path. Nothing is published yet.

  • @w2l/cli (bin w2l, AGPL-3.0-only): w2l scrape, crawl, batch, map and serve. It runs the API’s engine in its own process, so every REST option is a flag under its kebab-case name (--max-age, --no-only-main-content, --formats markdown,tables, --header name=value, …), and the REST parser checks the request. Output is JSON, --markdown for a scrape’s Markdown alone, and --out <dir> writes each page’s Markdown and each table’s CSV beside results.jsonl. Ctrl-C leaves a crawl or batch paused (w2l crawl --resume <taskId>, refused for a crawl recorded as pending or running). Its task root defaults to .w2l/cli, apart from the API’s, and --webhook is refused, since a command runs no delivery worker. w2l serve is the API server (runApiServer, exported from @w2l/api), and the engine takes resumeOnStart: false, so a one-off command does not run earlier jobs. npm run scrape / crawl now use it; the earlier in-process ladder CLI is w2l-ladder in @w2l/bench, which no longer claims the w2l bin.

  • Map: a link first found over http: is returned as its https: variant when the start page or a sitemap also lists that, instead of whichever came first. The https origin’s own robots.txt must allow it, and only an origin whose robots.txt the map read anyway counts, so the switch reads no further robots.txt, takes no slot of the host cap and cannot time the map out; otherwise the http link stays. The start URL stays as given. refused.samples.collapsed names the URL returned as into. On www.python.org the map had returned http://docs.python.org/3/tutorial/introduction.html though the page also links its https form (MP13). Crawls still fetch the first variant seen.

  • Markdown tables cap colspan at 1000 and rowspan at 65534, as browsers do, and bound the empty cells their grid adds for spans and short rows: 500,000 per table and 2,000,000 per page. A table past either is still one GFM table, written as each row’s own cells, unpadded. Before, an 892-byte page with four colspan="1000000" cells over 40 rows gave 516 million characters of Markdown (516,123 now), a 380 KB page whose one wide empty row sat over 20,000 one-cell rows gave 60 million (120,012 now), and a table of 200,000 rows threw. Tables within those limits are written as before. EXTRACTOR_VERSION is now extract-tf/6; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • tables format (scrape, batch, crawl; REST, SDK, MCP): every data table of the content the Markdown was written from, one entry per GFM table of that Markdown in its order, as { tableIndex, caption, sourceUrl, headerRows, columns, rows, csv, csvSha256 }. Cells are plain text, a spanned cell is repeated in every slot it covers so no row is shifted, and csv is RFC 4180 with CRLF line ends. Spans are capped as browsers cap them, and a table that would hold more than 2,000,000 characters (its spans repeated and every row padded to the widest), or more than what is left of 5,000,000 for the page’s tables, is omitted: "too_large" with no rows. The HTML-to-Markdown walk is shared, so the GFM output is unchanged (EXTRACTOR_VERSION stays extract-tf/5). /fc refuses the format, which Firecrawl does not have.

  • PDF options: scrape, batch and crawl take Firecrawl’s parsers ([] reads no PDF; one pdf entry, "pdf" or { type: "pdf", mode, maxPages, pages, pageMarkers }). mode is fast or auto; ocr and the image parser are refused by name. maxPages (1 to 10,000) cuts as asked and stays success with a page_cap warning. pages: true adds pages: [{ pageNumber, markdown }], and pageMarkers: false leaves out the <!-- page N --> lines. /fc maps them, plus v1 parsePDF, and now writes no page markers unless asked, as Firecrawl does. data.metadata.numPages gives a PDF’s page count. pdfToMarkdown takes pageMarkers, and its default output is unchanged (PDF_TEXT_VERSION stays pdf-text/1).

  • Sitemap reader behind an environment proxy: each redirect hop now carries its own Host. undici’s ProxyAgent writes host into the headers object it is given, and the reader reused one object across a file’s hops, so after a cross-host redirect every later hop still named the first host. A CDN that routes by Host then redirected again until redirect_limit: https://python.org/sitemap.xml → https://www.python.org/sitemap.xml looped this way (the 2026-10-03 map runs’ MP13/MP14 sitemap_unreadable). Direct connections and the http lane’s page requests were not affected.

  • Cache: scrape, batch and crawl (REST, SDK, MCP scrape / crawl / batch_scrape, /fc) take maxAge and minAge (milliseconds, 0 to ten years, minAge at most maxAge), storeInCache (default true) and lockdown (cache only). W2L stores the most recently fetched success of each page under one set of options (the URL without its fragment, the mode, every lane option but timeout, the rungs the request may use, fastMode, a recorded robots override, the extractor versions and W2L_SOURCE_COMMIT) in <taskRoot>/page-cache.sqlite; a request with custom headers stores only with storeInCache: true, and reuses it only when a request asks: the default maxAge is 0, so nothing is reused unless asked, on /fc too. A reuse is the original fetch unchanged, its evidenceRecord included, with metadata.cacheState: "hit" and metadata.cachedAt (its fetchedAt), channelsTried: [] and no request, attempt, byte or browser time in usage; a lookup that found nothing is cacheState: "miss" and fetches; a request that looked nothing up carries no cacheState. Batch items and crawl pages carry cacheState / cachedAt, and a reused page is cached: true and counts in cachedPages without touching its host’s Crawl-delay. lockdown never fetches: a page without a stored result is failed with the new failure reason cache_miss (an agentHints entry says why), /fc/v1/scrape answers it with HTTP 404 SCRAPE_LOCKDOWN_CACHE_MISS, and a crawl in lockdown needs sitemap: "skip". Mode authed neither stores nor reuses (a lookup there is HTTP 400), a URL the request’s allowlist refuses is never answered from the cache, and Monitor captures neither read nor fill it. useCached keeps its meaning: a resumed crawl reuses its own pages.

  • Evidence Record v1 gains identity.device (desktop / mobile as the answering lane declared it; null in research mode and when no request was sent) and identity.requestHeaders (the custom headers that lane sent, sorted by name, each with the SHA-256 of its value and never the value; [] when none, null when no request was sent), and the reason cache_miss. Both fields are optional in the schema file, so records written before them stay valid; W2L writes them on every record.

  • HTTP lane: a body is decoded by its Content-Encoding whether or not W2L asked for it (its identity sends no Accept-Encoding; www.python.org answers gzip anyway, which left its page failed/empty_unverified with 0 links and gave a map a misleading start_page_client_rendered warning). gzip, x-gzip, deflate (zlib or raw) and br are decoded, under the 50 MiB decompressed cap (decompressed_too_large past it); scrape, batch, crawl and map all read through it, and so does the crawl’s sitemap reader, which before inflated only gzip found by its magic number. usage.bytesWire is now the size received, not the decoded size. The coding is on evidence.contentEncoding and, additively, on the Evidence Record as contentEncoding (identity for none; null in the browser and provider lanes). An unknown coding fails with the new reason unsupported_content_encoding, a body that does not decode with parse_error; neither is read as the page. The declared identity headers are unchanged.

  • Map: POST /v1/map (SDK map(url, opts), getMap(id); GET /v1/maps/:id) lists a site’s URLs from the sitemaps it declares and its start page’s links without fetching each page: at most one page body (the start URL, on the http rung alone), robots.txt of the start host and at most 20 others, at most 50 sitemap files, all inside one deadline (timeout, default 60,000 ms, 1,000 to 300,000). It takes url, mode (standard or research), limit (1 to 100,000, default 5,000), origin and integration; anything else is refused by name (useIndex with a hint; mode: "authed" refused). Each link carries via (start, link, sitemap), the sitemap file and lastmod that listed it, its robots.txt verdict, and a title only from evidence in hand (the start page’s own metadata, an anchor’s text, a sitemap’s <news:title>). Robots-disallowed URLs are left out and counted in refused; a host whose robots.txt could not be read is left out too, as a complete disallow, and said so apart (a robots_unreachable warning with the reason; the start page’s robots: "unreachable" and robotsUnreachable), never worded as a rule the publisher wrote. At the deadline the map answers 200 with what it found as partial (or failed when it found nothing), stoppedBy: "timeout" and a map_timeout warning. When the start page’s links alone fill limit the map answers at once, completed with stoppedBy: "limit", and reads no sitemap (so overLimit counts only what was seen before it stopped). Each map is recorded under <taskRoot>/maps/. A hosted server caps limit at 5,000 and timeout at 60,000, refuses a browser-only URL, and counts /v1/map in the rate limit.

  • Map options: POST /v1/map takes search (1 to 200 characters, at most 10 words; a URL is kept when every word appears, case-insensitively, in its percent-decoded URL or its title in hand; a filter in discovery order, counted before limit, with what it left out in refused.searchFiltered), sitemap (include, skip: no sitemap requested, only: no page read, the start URL only when a sitemap lists it, and no sitemap read is failed with sitemap_unreadable), includeSubdomains (the crawl’s allowSubdomains rule; robots.txt read once per new host, at most 20) and ignoreQueryParameters (query variants folded into the first one seen, each fold counted with { url, into } samples), plus the crawl’s includePaths, excludePaths, regexOnFullURL, crawlEntireDomain and deduplicateSimilarURLs with their crawl messages; allowSubdomains and allowExternalLinks stay refused by name. Defaults stay as before (includeSubdomains and ignoreQueryParameters false, where Firecrawl v2 documents true), so a map without them answers as it did. With sitemap: "only" a browser-only URL is not refused, since no page is read. A collapsed sample now names the variant as offered, with its query.

  • /fc/v1/map (Firecrawl’s map, v1 ignoreSitemap / sitemapOnly and the v2 names): { success: true, id, links: [url strings], warning?, agent_hints? }, or { success: false, id, error, links: [] } for a map that found nothing; both v1 flags true is HTTP 400; useIndex, location, ignoreCache, threatProtection and auditMetadata are refused by name. FIRECRAWL_SHIM_SNAPSHOT lists /map under paths and no longer under notCovered.

  • MCP: the local map tool (annotations read-only, idempotent, open-world; an outputSchema) answers a compact map ({ id, status, stoppedBy, links: [{ url, title?, description? }], warning?, agentHints?, counts }, or the native response with debug: true) as text and as structuredContent; a tool that declares an outputSchema now always answers structuredContent too. The hosted MCP host does not offer map.

  • The crawl’s sitemap reader takes two options a crawl does not pass, so a crawl is unchanged: accept (judge each entry before it is collected; refused entries do not count toward maxUrls) and softDeadlineAt (return what was read, the file in flight recorded as unreadable/timeout, truncated: "time"). parseSitemapXml(text, { details: true }) reads each entry’s <lastmod> and <news:title> in linear time also when the locs stand bare, and only a load that passes accept or softDeadlineAt asks for them, so a crawl’s parse is the one it was; collectLinkDetails(html, base) gives each link once with its first non-empty anchor text. useIndex refused on any route now carries the no-URL-index hint.

  • Job streams: GET /v1/crawl/:id/events (new) and GET /v1/batches/:id/events (re-shaped) stream a crawl or batch as server-sent events named catchup (the report as the stream opens), document (one per page recorded, every outcome, the compact page the items routes list; id: is the step cursor), snapshot (the report after each page), done (the terminal report, after which the stream closes) and error ({ code, message }); ?after=<cursor> or Last-Event-ID resumes after a document. The batch route’s old progress, paused and complete events are gone. GET /v1/crawl/:id/ws and GET /v1/batches/:id/ws upgrade to a WebSocket (@hono/node-ws, a new dependency of @w2l/api) carrying the same frames as JSON; a server with tokens takes the bearer token in the Authorization header or as the subprotocol w2l.token.<token>, echoed back, never in a query string; a missing job closes with 4404, a bad cursor with 4400, done with 1000. The server subscribes to the in-process JobEventHub once per client and reads the checkpoint once as a stream opens and once per page for every client of a job together (no more per-client 500 ms re-reads). AppOptions.jobStreams / W2L_JOB_STREAMS=off turns all four routes into 404s. New ApiEngine.listJobPages and jobEvents. SDK: watcher(jobId, { kind, transport, pollIntervalMs, timeoutMs, after, signal, WebSocket }) returns a JobWatcher (an EventTarget with document, snapshot, done and error events, an async iterator, data, status, transport, close()), trying the WebSocket, then SSE, then polling the status and listing routes (pollIntervalMs default 2000, at least 250), switching once per level from the last document’s cursor so each document is emitted once per step id; a 401/403 is final, timeoutMs ends the watch with watcher_timeout while the job runs on; crawlAndWatch and batchScrapeAndWatch start and watch in one call. MCP keeps request/response only; the frozen v1 shim adds no socket path.

  • Batch: GET /v1/batches/:id counts empty_verified items as succeeded (a page read, with or without content), so with one step per URL completed is succeeded + failed. POST /v1/batches takes Firecrawl’s extract scope flags allowExternalLinks and includeSubdomains as false only, which already holds; true is HTTP 400 naming the crawl option that does it (allowExternalLinks: true is not offered on a batch: a batch fetches only the URLs given; a crawl takes allowExternalLinks, and extraction across links is the M5 multi-URL extract); MCP batch_scrape declares them as const: false. Docs name the extract mapping: a batch with a json format, no merged data or sources until the M5 multi-URL extract, creditsUsed and expiresAt null on the shim.

  • agentHints gains rows, derived from what the lanes recorded and never a check that was not made: an egress-policy refusal (ssrf_denied / governance_refusal) names the policy and the recorded reason; a page the http lane got blocked or an HTTP error for and the browser lane then served names what the http lane got (read from the ladder summary) and says to expect the browser lane for the host (the ordinary hop after a thin or empty http answer stays silent); tls_error, timeout and partial name the honest option (skipTlsVerification locally, a longer timeout); empty_unverified without a shell caveat points at onlyMainContent: false and includeTags (a PDF without a text layer: no OCR); an incomplete json names the required fields not found and the model fallback’s environment, and a model_unavailable issue says the fallback did not run; a 404 adds check the link. At most five hints stay, in the table’s order. Batch items and crawl pages read the same table from the stored audit.

  • Crawl and batch starts take webhook (REST, SDK, MCP crawl and batch_scrape; /fc/v1/crawl maps Firecrawl’s webhook): a URL string or { url, headers, metadata, events, secretEnv } naming a receiver the job posts its events to as durable, retried deliveries (the Monitors’ delivery stack, now with a kind: job destination job:<taskId> and events_json, headers_json, metadata_json, payload_format columns): started (sequence 0), one page per page recorded, whatever its outcome, with the page as the items routes list it, then completed, failed or cancelled with the job’s status report; events narrows them (default all five; a filtered event is never enqueued), metadata (at most 32 strings of 1000 characters, 8 KiB) is echoed in every payload, headers (at most 32, 8 KiB, token names lower-cased, content-type, content-length, host, connection, transfer-encoding and x-w2l-* refused by name) go with every attempt and are stored in the control database alone (destinations expose headerNames), and secretEnv signs each delivery as for a Monitor. Payload { schemaVersion: "w2l.job-event/v1", eventId, sequence, jobId, jobKind, event, at, metadata, page?, report?, error? } with deterministic event ids (<taskId>:started, <taskId>:page:<stepId>, <taskId>:<status>, the attempt id appended for a later attempt’s terminal event), so a resume or restart offers every persisted step again and sends none twice, and a finished job whose events a crash cut off is completed at the next start; headers x-w2l-event-id, x-w2l-event-version, x-w2l-delivery-id (and the signature pair). GET /v1/crawl/:id and GET /v1/batches/:id report webhook: { destinationId, url, events, pending, delivered, deadLetter }; GET /v1/deliveries, /v1/deliveries/page and /v1/delivery/destinations take jobId (SDK and MCP too). The batch status gains succeeded and failed counts. A hosted server takes public https receivers only (webhook.url must be https (...), webhook.url must be a public address); a local one also takes plain http to a loopback receiver, sent direct over node:http (DeliveryWorkerOptions.allowHttpLoopback, local w2l-api and the local MCP service), everything else staying HTTPS with verification. w2l-api now runs a delivery worker of its own under a delivery policy it prints at start (TLS verified, the shell’s proxy never used, W2L_DELIVERY_PROXY_URL, W2L_DELIVERY_CA_FILE, W2L_DELIVERY_PRIVATE_ALLOWLIST). New OrchestratorOptions.onStep hook and JobEventHub in @w2l/api carry a job’s events; the Firecrawl shim’s receivers get { success, type: crawl.started | crawl.page | crawl.completed | crawl.failed, id, data, metadata, error? } (wrapJobWebhook). The hosted MCP host refuses webhook on batch_scrape; the lanes and the Firecrawl shim’s other mappings are unchanged.

  • Batch and crawl starts take idempotencyKey (1 to 200 characters; also the x-idempotency-key / Idempotency-Key header on POST /v1/batches, POST /v1/crawl and /fc/v1/crawl, which the Firecrawl v1 SDK sends): a retried submission with the same key and body answers the first one’s taskId with replayed: true and starts nothing; the same key with another body is HTTP 409 conflict; a body key that differs from the header is HTTP 400; keys live 24 hours in <task root>/idempotency.sqlite (new IdempotencyStore in @w2l/runtime), per the one API process that runs the task root. Batch takes appendToId (SDK appendToBatch(id, urls, options), MCP batch_scrape): the URLs join an existing batch at the end of its list, the job keeps its options (one in the body is refused by name), a running batch fetches them in the same attempt (no job is added, so maxActiveBatches does not count the append) and a completed one runs again for them in a new attempt (active again, it counts against maxActiveBatches as a new batch does and is refused with active batch limit reached while the limit is reached); the 202 carries requested and appended; a cancelled or failed batch is HTTP 409, a total over 1000 or a URL already in the batch HTTP 400, an id that is not a batch 404. On the way: a frontier seed passes the host scope (discovered links keep it), a batch’s governance lists no hosts (its own hosts added nothing and would have refused a URL appended on a new host), the orchestrator re-reads a batch’s row while it runs and writes the row as it then is at the end instead of the object it opened with. SDK: chunkUrls(urls, chunkSize = 100) and batchScrapeChunked(urls, options, { chunkSize, itemLimit, pollIntervalMs, timeoutMs, maxRetries }) run a list of any length as batches in sequence (a caller’s key becomes <key>:<chunk index> per job) and merge the items in submission order. The hosted MCP normalizer refuses both options on batch_scrape; the lanes are unchanged.

  • Batch (POST /v1/batches, SDK batchScrape, MCP batch_scrape) takes maxConcurrency (an integer from 1 to 4: the batch’s pages in flight at once, lowering the service’s worker count and never raising it; stored with the task, kept on resume, and reported as the cap in force on GET /v1/batches/:id) and ignoreInvalidURLs (start with the entries that are http(s) URLs and report the rest as invalidURLs on the 202 and the status; a non-string entry or a duplicate is still refused). A batch entry that is not a URL is now refused by its index (urls[2] must be http(s), urls[2] is required) instead of url must be http(s). New GET /v1/batches/:id/errors?cursor=&limit= (SDK getBatchErrors, MCP get_batch_errors, also on the hosted MCP host as a read): the items that did not succeed across every attempt, { id, timestamp, url, status, code, error, httpStatus } in pages of up to 1000, with robotsBlocked, the URLs a robots_disallowed trace event refused and no recorded override set aside. The hosted MCP normalizer keeps refusing the two options on batch_scrape; the lanes and the Firecrawl shim are unchanged.

  • formats takes screenshot (also Firecrawl v1’s screenshot@fullPage, or one { "type": "screenshot", "fullPage", "quality", "viewport" } entry per request), on scrape, batch, crawl, MCP and /fc (data.screenshot, a data URI string). Such a request runs on the local browser rung alone (channelsTried: ["browser_local"], the dropped rungs in the ladder audit); a server without one, or fastMode beside it, is refused with HTTP 400 naming it. The capture is taken after load, stability and waitFor, before the DOM is read, CSS-pixel sized (scale: "css") at the declared 1280x800 viewport or the viewport asked for (integers 320…1920 by 240…1080, within the declared screen: a window size, the identity unchanged, recorded as screenshot_viewport), the document’s whole height with fullPage (no scrolling first), a JPEG with quality. It is returned as { contentType, width, height, fullPage, viewport, deviceScaleFactor, quality, bytes, sha256, path, base64 } with a screenshot_captured trace event, saved as <sha256>.png / .jpg under W2L_CAPTURE_RAW_DIR (the Evidence Record lists it as kind: "screenshot" with size and type), attached to a success, an error page or a gate alike, null when the browser could not capture it (screenshot_failed, a screenshot_unavailable warning and hint, the page kept) or no page rendered, and never repeated in summary.attempts or stored audits.

  • A thin or shell-like HTTP answer that stays the answer carries a low_content_yield warning (The http lane extracted N tokens at confidence C; the browser lane did not improve it., or … was not available to this request. under fastMode, on an HTTP-only server or without a browser rung; found no main content on a failed/empty_unverified shell), after its other warnings. Every response with warnings also carries warning, their messages joined with a space (full and compact scrape responses, batch items, crawl pages, /fc data.warning, which so passes the native warnings through for the first time). agentHints gains two table entries: a low_content_yield warning suggests waitFor or a longer timeout with the browser lane available (left out under fastMode, whose own sentence stands), and a file result says what its markdown is.

  • formats takes images (every image URL of the whole document as received: img src and srcset candidates, <picture> sources, lazy data-src / data-srcset / data-lazy-src / data-original, video posters, image_src links, og:image and twitter:image, resolved, absolute http(s), fragment stripped, deduplicated, in document order, data: URIs counted and left out; images_collected in the trace) and one { type: "attributes", selectors: [{ selector, attribute }] } entry (1 to 50 selectors under the includeTags rules; per selector the attribute’s values as written, in document order; attributes_extracted in the trace), on scrape, batch, crawl, MCP and /fc (data.images, data.attributes); both only when asked for and only on a page read as content. The page option removeBase64Images (default true) names what Markdown always did with an <img> whose src is a data: URI (left out, alt text kept); false keeps the image as ![alt](data:…), counted by contentTokens; /fc now maps the option with its value instead of accepting true alone. A json schema request is told apart from other object formats by its type, so an attributes entry never switches on JSON extraction.

  • metadata (scrape responses, batch items, crawl pages, /fc data.metadata) carries the Open Graph tags a page states (ogTitle, ogDescription, ogUrl, ogImage, ogAudio, ogVideo, ogDeterminer, ogLocale, ogLocaleAlternate, ogSiteName), its Dublin Core tags (dcTermsCreated, dcDateCreated, dcDate, dcTermsType, dcType, dcTermsAudience, dcTermsSubject, dcSubject, dcDescription, dcTermsKeywords) and its article tags (publishedTime, modifiedTime, articleTag, articleSection), under Firecrawl’s names, each present only when the page states it and as written: no date normalisation, no fallback from twitter:*, govuk:* or citation_* tags, JSON-LD or <time>. The seven existing fields keep their always-present, nullable shape.

  • Roadmap v2 (weeks 1–16) makes P1 core correctness the current phase and adds the Firecrawl parity audit (research/parity/, frozen at firecrawl-js v4.42.0) and a real-site test set with a runner (node research/parity/run-sites.mjs) whose runs are recorded with command and commit.

  • JSON extraction no longer reports complete while a required field has no source: such fields are omitted with a missing_required issue, nested fields are not filled from page-level values, and JSON from a non-success page is incomplete with page_unsuccessful.

  • Markdown keeps block boundaries, inline spacing and emphasis, numbers ordered lists (with start), keeps code blocks and tables inside list items, drops empty emphasis from icon elements, and resolves link and image targets against the document base (<base href> included). Same-page #fragment links stay as written.

  • Scrape responses, batch items and crawl pages carry metadata read from the page’s own markup: the <title>, meta description, keywords and robots, the language, the favicon and the canonical URL, each null when the page declares none (nothing is inferred). /fc maps them into data.metadata. document.title stays the content title.

  • The API accepts several bearer tokens (repeated --token, W2L_API_TOKEN plus comma-separated W2L_API_TOKENS) and compares them in constant time; the SDK sends W2L_API_TOKEN when no token is passed.

  • A page whose blocks sit directly in <body> (example.com today) is extracted instead of reported empty_unverified, so the README’s first npm run scrape -- https://example.com works again.

  • In the browser lane, Markdown follows the page’s CSS where it differs from the tags: an inline element laid out as a block in the text flow starts its own paragraph (quotes.toscrape.com/js no longer reads thinking.”by), and inline text hidden with display: none is left out (GitHub Docs’ platform names read Open Terminal.). Hidden blocks, such as footnote popups and accordion panels, are kept. Evidence is unchanged (rawBodySha256 and raw artifacts are the rendered page without W2L’s markers); a page over 100 000 elements, one still changing after capture, or one short of time converts by its tags, and the trace’s layout event records which. Monitors on pages captured in the browser lane may report a one-time change.

  • A non-2xx response keeps its page as evidence (Markdown, links and snapshot.httpStatus, also on the default REST response and /fc) while its status stays failed or blocked; other 2xx statuses are judged like 200, and 204/205 are empty_verified. A robots.txt that never answers is recorded as unreachable instead of failing the scrape with HTTP 500, and a URL whose scrape throws becomes its own failed item instead of failing the whole batch or crawl.

  • Main-content selection keeps a single long block such as a <pre> news release; a form that wraps page content (a table viewer, an ASP.NET page) is unwrapped instead of removed; an HTTP page whose tables are empty script-filled shells is offered to the browser lane.

  • JSON extraction also fills top-level keys from the page’s own labels (two-cell th/td rows and dt/dd pairs in the main content), each with its location as evidence; labels that state different values leave the field out with a field_ambiguous issue. A page whose one heading is followed by its one visible price is routed as a product, so that price is read (books.toscrape.com).

  • JSON extraction accepts the JSON Schema subset Pydantic and Zod write: annotations (title, description, $schema, $id, default, examples, format, …), checks (minimum, maxLength, pattern, const, …), definitions, and anyOf / oneOf of a schema and null or of primitive types. Any other keyword is refused with unsupported_parameter, naming it where it was sent (formats[0].schema.properties.author.allOf); README and the docs reference list the subset.

  • JSON model fallback keeps every value read from the page with its evidence and merges the model’s answer key by key at every level: the model fills only what is missing, or a page value that breaks the schema, and each value it wrote has model evidence.

  • JSON model fallback sends OpenAI-compatible strict mode a strict-safe copy of the schema (closed objects, every property required, optional ones nullable), so an ordinary schema no longer fails as model_provider_error; a schema strict mode cannot express is sent without strict mode, and json.modelUsage.strict / strictReason say which.

  • JSON values filled from the content title, the URL or the page type carry evidence (dom h1[0] or title, fetch finalUrl / requestedUrl, inferred document.pageType); Evidence Record v1 adds the field-evidence source fetch.

  • JSON extraction no longer reads a number out of text that is not one amount (a SKU HL-1 was -1, a URL a fraction), no longer overflows the stack on a recursive schema, and names a value that breaks a schema check even when a required field is missing too.

  • A listing of cards in which the article cascade finds no text block (a sparse category page, data.gov.uk’s home page) is kept instead of being reported empty_unverified. An HTTP page that declares data its scripts will fetch (<link rel="preload" as="fetch">) is offered to the browser lane. Markdown leaves an empty first header cell empty instead of writing (header).

  • Batch items and crawl pages carry links when requested. formats has no count cap; an unsupported format or an unknown request field is rejected with HTTP 400 naming it, on the native API and on /fc. Crawl accepts formats, includeLinks, includePaths and excludePaths, and keeps the path filters on resume.

  • SDK: waitCrawl, and pollIntervalMs / timeoutMs for waitBatch and waitCrawl with a WaitTimeoutError that carries the last status.

  • Scrape, batch and crawl (per page), MCP and /fc accept onlyMainContent, waitFor and timeout. onlyMainContent: false returns the whole page’s Markdown with the same evidence. waitFor starts at the browser rung and waits after load before capture; with no browser rung the result says so (policy_denied, wait_for_unavailable). timeout (default 300 000 ms) is the whole scrape’s deadline: when it fires the answer is HTTP 200 with partial (the best content so far) or failed/timeout, never HTTP 500, marked usage.deadlineExceeded. A lane timeout no longer sets budgetExceeded: 'time', which the contract keeps for status budget_exceeded. JSON from a partial page is incomplete with a page_partial issue. /fc/v1/scrape now cancels when the client disconnects.

  • formats takes html (the cleaned HTML the Markdown is written from: the main content, the whole page for onlyMainContent: false, or an includeTags selection) and rawHtml (the page as the answering rung received it, hashing to snapshot.rawBodySha256), on scrape, batch, crawl, MCP and /fc; both are returned only when asked for and are null for a file or a result that is not success or partial. Two page options shape the content: includeTags keeps only the elements its CSS selectors name, in document order, a named navigation included, and excludeTags removes elements from the main content, the whole page, an includeTags selection and the evidence page of a failed result. An empty includeTags selection is success with empty Markdown; a block a later rung finds on that page is the answer instead, and a rung that repeats the empty answer ends the ladder before any vendor rung. They take tag, class, id and attribute selectors, descendant and child combinators, :not(), :is(), :where(), :root and :empty, at most 100 selector parts per list, read with the DOM library’s own parser (an escaped name such as #\31 23 is what it is to the library) and matched in time proportional to the page; a selector that does not parse is invalid_request, and one that uses a sibling combinator, a positional pseudo-class, :has() or another pseudo-class is unsupported_parameter, because matching those can take unbounded time that no timeout stops.

  • Request errors carry one code set across the REST API, /fc, the SDK and MCP: { error, code, details? } with invalid_json, invalid_request, unsupported_parameter, unsupported_format, unauthorized, not_found, conflict or internal_error; statuses and messages are unchanged. The SDK throws W2LError (status, code, method, path, body). A 500 hides its internal message unless the server runs in local mode; hosted mode logs the cause to stderr. The docs reference lists the codes.

  • Local mode sends outbound requests (HTTP lane, robots.txt, Monitor fetches, the browser lane) through HTTPS_PROXY / HTTP_PROXY, honouring NO_PROXY; loopback stays direct, W2L_PROXY=off ignores the variables, and hosted mode never reads them. Proxied requests record evidence.envProxy (host:port, never credentials). Response headers up to 64 KiB are accepted, and a name that does not resolve is dns_error instead of policy_denied.

  • Crawls follow links on the start URL’s www. twin and on the host it redirects to, and no longer fetch image, font, style, script, media or program links. A crawl task stores every option: a crawl paused by shutdown or left running by a crash resumes when the API starts, POST /v1/crawl/:id/resume (SDK resumeCrawl, MCP resume_crawl) restarts a paused or failed one, and w2l crawl --resume runs with the stored limits and refuses a flag that differs. maxPages counts the task’s pages across resumes. GET /v1/crawl/:id counts pages while the crawl runs, and crawl pages carry their audit and trace only with debug=true. The HTTP lane reports robots.txt Crawl-delay, a host starts one page at a time until its robots.txt is known, and each crawl page records the delay it waited in a crawl_delay trace event. /fc/v1/crawl/:id reports creditsUsed and expiresAt as null instead of 0 and an invented expiry.

  • Monitor observations record the raw body’s SHA-256 and the extractor version (EXTRACTOR_VERSION, now extract-tf/1; unknown on older rows). Fields that change over a byte-identical raw body (a 304 counts as its reused body) are extraction_reprocessed: the run records changeReason (MCP get_monitor: latestRun.changeReason) and the baseline moves, but no event or webhook delivery is created. So the first run after this upgrade no longer reports source_changed where only W2L’s Markdown changed (for the Firecrawl introduction preset: restored spaces in three descriptions). When the raw body changed too, the event stays source_changed and carries extractorChange if the extractor differs from the baseline’s. Delivered eventVersion values can now skip a version.

  • A robots.txt that cannot be fetched (5xx, network error or lookup timeout) is a complete disallow in local and hosted mode alike (RFC 9309 §2.3.1.4): the page is failed/policy_denied, and the trace and the compliance record’s robots decision carry unreachable with the reason, which the record’s hash covers. It is fetched again after five minutes (robotsUnreachableTtlMs). Before, only the public preview failed closed; the API, MCP and CLIs fetched the page anyway. The provider lane follows the same rule and no longer reads a 5xx robots.txt as no robots.txt.

  • The hosted public preview names OctoCrawl to the sites it reads: its User-Agent, on robots.txt, the page and the Amazon.sg browser’s requests, is the standard Chrome User-Agent followed by OctoCrawl-Preview/1.0 (+https://octocrawl.dev), with the same client hints, so a site can address it in robots.txt with User-agent: octocrawl-preview (or octocrawl). The token was W2L-Preview/1.0 (+https://github.com/77777R7/w2l) before the rename, so a robots.txt group naming w2l-preview or w2l no longer applies to the preview. The local API, MCP and CLIs are unchanged: standard mode still sends the plain Chrome User-Agent.

  • Without the environment proxy (no variables, W2L_PROXY=off, hosted mode), the browser lane launches Chromium with --proxy-server=direct:// instead of silently using the operating system’s proxy; with it, the launch names the environment proxy.

  • W2L_CONTACT adds the operator’s contact to the research-mode User-Agent (; contact: …). On the HTTP lane, a 403 from an SEC host to a request that declared no contact carries a declared_contact_hint trace event.

  • Research mode with W2L_CONTACT declares SEC’s own User-Agent format, W2L Research <contact>, to sec.gov and its subdomains in the HTTP and browser lanes (robots.txt included), which SEC.gov serves where it answered 403 to the research format; a robots.txt group for w2l-research still applies there, the Evidence Record’s identity.contact reads either format, and other hosts keep the research User-Agent.

  • Crawl includePaths / excludePaths that can backtrack catastrophically, such as ^/(a+)+$ or .*a.*b, are refused with invalid_request (unsafeRegexReason in @w2l/contracts); the rest run on V8’s linear-time regular expression engine where it can run them, and otherwise (lookaround, backreferences, counted repetitions above 16) on paths of up to 2,048 characters with a 100 ms limit per link, so a crafted link path can no longer stall the server. A link a filter cannot decide is not followed.

  • JSON Schema pattern gets the same check: a catastrophic pattern is refused with invalid_request, and page or model text is matched in linear time where V8 can, otherwise only up to 2,048 characters (longer text counts as breaking the pattern).

  • Hosted mode enforces its crawl limit (100 pages; 10 on the hosted MCP): an omitted or null maxPages takes it and a larger one is refused with invalid_request, where null or a large number used to skip it.

  • w2l-api refuses to start when a --token has no value (last, followed by another flag, or blank) instead of taking the next flag as the token or ignoring it; the error never repeats a token.

  • Evidence Record v1: scrape responses (full and compact, so MCP too), batch items and crawl pages carry evidenceRecord (schemaVersion: "w2l.evidence/1"), one shape for every lane, published as the JSON Schema packages/contracts/schemas/evidence-record.v1.json: final URL (null when no request was sent), a redirect chain that says whether every hop was observed, fetchedAt, HTTP status, status and reason, lane, the robots.txt decision with its hash, raw and output hashes (delivered Markdown; json.data as RFC 8785 canonical JSON), extractor name, version and W2L_SOURCE_COMMIT, field evidence, saved artifacts, the environment proxy and the User-Agent sent. Existing fields are unchanged; lanes now also record evidence.fetchedAt, and the HTTP, browser and provider lanes’ robots_checked trace events carry the robots.txt hash (the provider lane emits one too). /fc does not carry the record.

  • PDF text, as a library function only (scrape, batch, crawl, the API and MCP do not reach it yet): pdfToMarkdown in @w2l/extract-tf turns PDF bytes into Markdown with a <!-- page N --> line before each page and each page’s text offsets, read with the pinned pdfjs-dist 6.3.289 (Apache-2.0). No OCR (no_text_layer), tables unverified (tables_unverified), a page cap and a time budget, and an error result for encrypted, malformed or non-PDF input. Checked on 10 public reports in research/pdf-corpus/.

  • A scrape’s timeout also sets how long the lanes wait for a slow server: with one, the HTTP lane waits for headers and body and the browser for navigation until the deadline, instead of stopping at 10 s, 30 s and 20 s; without one those defaults stay. The browser’s navigate trace event records the wait it allowed (timeoutMs).

  • MCP: a client that cancels a tool call (notifications/cancelled) aborts the API requests the call made, so a cancelled scrape stops on the server and a cancelled wait_batch stops waiting.

  • MCP over HTTP: the local and hosted Streamable HTTP services stop a call when its client cancels it (notifications/cancelled) or closes the call’s request before the result. Before, every POST got a new server, so neither reached the call, which ran until its deadline. The services still keep no session, but give each client an Mcp-Session-Id at initialize, and a cancellation reaches only a call made with the same session id and, on the hosted service, the same bearer token (a hosted client that sends no session id is matched by its token; a local one can stop a call only by closing its request). The call’s request is answered with the JSON-RPC error Request cancelled (code 0).

  • SDK: scrape waits for the API’s answer until its timeout (default 300 000 ms) plus 30 s. On Node, fetch stops waiting for the response headers after 300 to 301 s, and the API answers a scrape with the default or largest timeout at its 300 000 ms deadline: an answer that arrived later than fetch’s wait threw TypeError: fetch failed instead of returning the API’s failed/timeout.

  • Batch and crawl: a page’s timeout also bounds JSON extraction’s model fallback, which ran until the task’s own deadline; a fallback it cuts short leaves the JSON incomplete with a model_timeout issue.

  • SDK: waitCrawl and waitBatch retry a status request that fails with a network error, 408, 429 or 5xx (after 1, 2, 4, 8, then 10 s, or a Retry-After of 60 s or less; maxRetries, default 5) and throw other errors at once. timeoutMs also ends a request in flight, WaitTimeoutError.last is null when no status was read and cause holds the last failure, invalid wait options throw RangeError, and W2LError carries retryAfterMs. crawlAndWait and batchAndWait start a task, wait, and return every page and error, or every item.

  • /fc/v1/crawl/:id reports a cancelled crawl as cancelled instead of failed; completed counts the latest attempt’s successful pages and total all its pages (failed, blocked and duplicate ones included) plus, while this API process runs the crawl, those in flight and queued (null for a paused crawl), instead of both being the step count. data holds at most 100 pages (limit, up to 1 000) with a next URL carrying a native cursor, left out after the last page; skip is rejected.

  • Markdown leaves out data: image URIs and keeps their alt text, as Firecrawl’s removeBase64Images does by default; /fc accepts removeBase64Images: true and rejects false.

  • The Markdown of an error page kept as evidence resolves link and image targets against the page URL, as on a success (MDN’s 404 page no longer has site-relative links).

  • A page on which the extractor finds no main content keeps the whole page’s Markdown and links as evidence on its failed/empty_unverified result, in the HTTP, browser and provider lanes. When the HTTP rung found no main content and the browser rung then fails without a page, the answer is the HTTP result with its page (ladder_evidence_kept in the ladder audit) instead of the browser’s failure; when the deadline ends the browser rung, failed/timeout keeps that page.

  • onlyMainContent: false returns the whole page as success on a page where no main block is found, instead of failed/empty_unverified; the browser rung is still tried, and answers when it renders more.

  • A table whose nested tables hold at least half of its text no longer counts as the page’s data table: on Hacker News the main content is the story list, not the whole layout table with its header and footer. In Markdown, such a table, or one with a single row, becomes paragraphs, and only the data tables inside it become GFM tables. A data table with a small table in one cell stays the page’s table and one GFM grid, with that small table as the cell’s text.

  • The browser lane reports the response’s content-type header instead of text/html; rendered, and every redirect hop Chromium followed instead of [requested, final]; its Evidence Record says redirectChain.complete: true when every hop is known (evidence.redirectChainComplete).

  • The browser lane reports the final URL, status and content-type of the document the page shows, and judges the page by them, instead of the status and type of the navigation W2L started next to the URL and content of another document. A page that answered 200 and then went to a 404 page by location.replace or a meta refresh is failed/http_error with the 404 page’s URL, status and type on REST, in the Evidence Record, in the compact snapshot and on /fc (it was success with 200); a 403 challenge reached that way is blocked. The redirect chain lists each document a script or a meta refresh loaded after the server’s redirects, with complete: true, also for a file the browser displays or downloads (which kept [requested, final] without the flag). A URL the page sets with the history API (pushState, replaceState) is not a redirect: the final URL stays the one the document was loaded from, with its status, the page’s URL stays the base of its links, and a same_document_navigation trace event names both. A page that navigates during the wait for stability or waitFor (a zero-second meta refresh) is captured on its new document instead of failing with connection_error. A page that loads a new document during every read (a refresh or script redirect loop) is failed/redirect_loop with no content and no status, and the lane no longer waits without end for Chromium to close it. Server redirects Chromium stops following (a loop, or more than 20) are redirect_limit on the browser lane instead of connection_error. On the HTTP lane, a redirect loop’s final URL is the URL whose redirect closes the loop, whose status it reports, instead of the URL that redirect points back to.

  • The compact scrape response, and so MCP scrape, carries snapshot.contentType next to snapshot.httpStatus; /fc metadata carries url (the final URL) and contentType.

  • EXTRACTOR_VERSION is now extract-tf/2 for the Markdown changes above; a Monitor whose Markdown changes only through them records extraction_reprocessed once.

  • File download: scrape, batch, crawl, MCP and /fc take PDF, CSV, JSON, plain-text, XLSX, XLS and ZIP responses (by Content-Type, or by the bytes under application/octet-stream or none) as files. The bytes are saved as received to <W2L_TASK_ROOT>/files/<sha256>.<ext> (stored once per content), described in a new file block (kind, detection, content type, declared and received size, SHA-256, path, PDF pages) and listed in the Evidence Record’s artifacts as kind: "file"; rawSha256 is the SHA-256 of those bytes. A file never escalates to the browser, and the browser lane catches the download a file starts (waitFor) instead of failing with Download is starting. A PDF’s Markdown is its text with a <!-- page N --> line per page and the extractor pdf-text/1; a PDF with no text layer is failed/empty_unverified, one that cannot be opened failed/parse_error; CSV, JSON and text give their text (file-text/1), XLSX, XLS and ZIP none. JSON extraction reads a PDF’s Label: value lines with fieldEvidence { source: "pdf", locator: "page N \"label\"" } and never sends PDF text to a model. W2L_MAX_FILE_BYTES caps a file (default 50 MiB, at most 500 MiB) and a request’s maxFileBytes can lower it; a file over the cap is failed/body_too_large with its declared size and nothing saved. Images, audio, video, fonts and office documents other than spreadsheets are failed/unsupported_content_type instead of being parsed as pages and sent to the browser. A response body that stalls or breaks off after the headers is failed/timeout or connection_error instead of HTTP 500. The Evidence Record’s artifacts gain bytes and contentType, optional in the schema under the new additive versioning rule of w2l.evidence/1.

  • Node.js 22.13 or later is required (engines), as pdf.js needs it; CI runs Node 22.

  • JSON extraction reads a number as the page writes it: a decimal comma or point, thousands grouped by ., ,, any space or an apostrophe (India’s lakh groups too), a currency before or after, ,- for a whole amount. On a product page’s visible price, 12,99 € gave 1299 and now gives 12.99; 1.299,00 € gave 1.299, 1 299,00 € 29900 and CHF 1'299.00 1, and each now gives 1299; a JSON-LD price Call for price gave 0. Each of those was reported complete. A single . or , before three digits (1.299 €, $1,299) is read only when a review count, a currency without minor units or a JSON-LD / OpenGraph price settles it; otherwise the field is left out (null when nullable) with a field_unavailable issue quoting the text, instead of a number that may be 1000 times off. The page’s language, currency and domain are not used to guess. Page labels and PDF Label: value lines follow the same rule (a label reading $1,299 gave 1299 and is now left out with that issue; 1.299,00 €, refused before, gives 1299), and json.evidence quotes the text of each number read from text (text); the Evidence Record is unchanged. The visible-price reading in @w2l/extract-tf takes an amount grouped by spaces or apostrophes whole (1 299,00 €, not 299,00 €); which elements count as prices is unchanged, so the Markdown is too and EXTRACTOR_VERSION stays extract-tf/2.

  • A product list or map the extractor reported empty (images, prices, variants, specifications) has a JSON evidence entry, inferred with document.product.<name>, where it had none.

  • README and the docs reference state the JSON Schema pattern limits the code has: a pattern of at most 2,000 characters, and the 2,048-character text limit also for counted repetitions above 16, \p{…} escapes and text with characters outside the Basic Multilingual Plane, not only for lookaround and backreferences.

  • A page written as <div>s around its tables, such as an SEC EDGAR inline XBRL filing, keeps its text. The table strategy, which keeps one table, is used only when the page’s tables hold at least half its text; on a page whose <p> elements hold under 750 characters, <div>s with no block inside count as paragraphs when they hold at least that much prose; and when the blocks the body lays out itself (inline wrappers such as ix:nonNumeric looked through) hold most of the text, the body is the main content. A filing’s hidden <ix:header> (XBRL facts and contexts) is left out of the main content and of every Markdown. Before, IREN’s 10-Q (real-site case A36) was success with only its statement of operations table (6,843 characters of Markdown). EXTRACTOR_VERSION is now extract-tf/3; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • A page W2L extracts by its tables (document.strategy: "table") keeps all its data tables: its main content is the lowest element that holds them, with the headings, captions and text between them, instead of the largest table alone, and JSON extraction reads page labels from all of them. Before, a statistics page with three captioned tables under <h2> headings and its <h1> in a banner was success with the first table only. A menu laid out as a table (mostly links, and no figures outside them) counts only when it is the page’s largest table, and a table with no text is not a data table. Table cells and captions keep their links and images with absolute targets, as paragraphs do, with every | written \|; before, Hacker News’ story list had no link target. A data: link target is dropped as a data: image is, and the link keeps its text. EXTRACTOR_VERSION is now extract-tf/4; a Monitor whose Markdown changes only through this records extraction_reprocessed once.

  • A recorded robots override for one URL: scrape takes robotsOverride ({ reason, recordedBy? }) and a batch takes robotsOverrides ([{ url, reason, recordedBy? }], each url one of the batch’s own, stored with the task), on REST, the SDK and MCP. robots.txt is still read and its verdict recorded; when a rule disallows the URL, the fetch goes ahead with a robots_overridden trace event after robots_disallowed, a robots_overridden warning first in the result’s new warnings (full and compact scrape responses, batch items), the override in the browser lane’s compliance record (skippedFetch: false, covered by its hash) and robotsDecision.userOverride: true in the Evidence Record. Other URLs are unaffected, an unreachable robots.txt is not set aside, and a blanket ignoreRobotsTxt stays an HTTP 400 that names it. The warning and the events also stay on a failed/timeout when the deadline passes while the request is out, and on the answer of a later rung that fetched nothing. The override is for local fetches: the provider lane takes none, a scrape that set a rule aside does not go on to a vendor rung, and a hosted server (--hosted) refuses both fields with HTTP 400 unsupported_parameter.

  • Five execution options on scrape, batch and crawl (per page), MCP and /fc, each stored with a batch or crawl task. headers (at most 32, values of at most 4,096 characters) are sent to the requested origin after the declared identity and never override it: User-Agent, sec-ch-*, sec-fetch-*, Authorization, Proxy-Authorization, Cookie and the transport headers are refused with HTTP 400 naming them; a redirect hop to another origin and robots.txt get the identity alone on both rungs (custom_headers_withheld; the browser rung adds the headers per request through Chromium’s request interception, which carries no override to a redirect hop, instead of a Playwright route, whose overrides ride every hop), and the headers are on the record (request_headers_added, the HTTP rung’s identity_sent, the browser rung’s compliance sentHeaders). mobile fetches as a second declared identity, Android Chrome with aligned hints and a 412x915 touch viewport, which passes the same coherence and honesty checks, is evaluated against robots.txt and is recorded (identity_sent.device, identity_declared); it is refused with research mode. skipTlsVerification relaxes certificate verification for one local fetch and its robots.txt lookup, through routes of its own closed after it, with a tls_verification_skipped trace event and a tls_unverified warning; a hosted server refuses it. Without it a certificate that does not verify is now failed/tls_error with the error code in the trace (it was connection_error, or policy_denied when robots.txt failed on the same certificate). fastMode keeps the HTTP rung alone (ladder_channels_filtered): a script-written page is failed/empty_unverified, never rendered, and the response carries agentHints when the HTTP rung asked for the browser rung; a URL the server binds to the browser lane refuses it. blockAds (default true) aborts requests to a bundled list of about fifty ad-serving hosts on the local browser rung (ads_blocked) and switches the extractor’s ad and cookie-banner pruning, which was always on; false keeps them in markdown and html. A request with headers, mobile or skipTlsVerification never goes on to a vendor rung. ApiEngineOptions.hosted tells the engine it serves a hosted API (set by --hosted and the hosted MCP host).

  • The browser lane sets Chromium’s user-agent metadata to the declared identity, so the client hints Chromium generates itself (sec-ch-ua, sec-ch-ua-mobile, sec-ch-ua-platform on a server redirect hop and on the page’s own requests, where a context’s extra headers do not reach) carry the declared brands instead of the headless shell’s HeadlessChrome, and navigator.userAgentData says the same. Before, a redirected page’s second hop went out with HeadlessChrome hints while the compliance record repeated the declared value, since Playwright reported the headers it had set rather than the wire’s. The as-sent sentHeaders never carry a credential header (cookie, authorization): the access fact keeps the session as a hash.

  • The HTTP rung says when the page it read looks like a shell for data its scripts fill in: a <table> with no cells beside scripts, an empty application root, a page of scripts with little text, an element hidden once scripts run, a hydration blob on a thin page or aria-busy; a noscript notice alone counts only on a thin page or beside hydration state. The result keeps its status, the failed/empty_unverified result that keeps the page as evidence when no main region was found included, and carries a client_rendered_suspected warning naming the rule, its trace a quality_client_rendered event, and the ladder offers the page to the browser rung as it does a thin result, keeping whichever answer holds more; the browser rung never raises it. The extractor exposes the signals as render.

  • C2 adds Monitor sample preview, paused-by-default MCP creation, durable queued manual runs, run detail, delivery paging and dead-letter controls. A local SDK Streamable HTTP check completed the public Firecrawl document → HTTPS webhook flow with the same event ID at sender and receiver.

  • C3 adds a single-process API/scheduler/delivery-worker runtime, protected-resource metadata, WorkOS JWT validation and a restricted Streamable HTTP MCP endpoint. A two-service Render Blueprint and walkthrough are ready; permanent deployment, browser OAuth and actual hosted restart acceptance remain open.

  • Scrape requests now accept explicit Markdown, links and JSON Schema formats. Deterministic extraction runs directly on subject-bound product HTML/metadata; nullable missing fields include evidence-backed issues, and an OpenAI-compatible model fallback is explicit and disabled when unconfigured.

  • Amazon /dp/{ASIN} pages use a subject adapter for identity, purchase offers, seller, availability, delivery context, variants, images and specifications. Recommendation shelves are removed before output; Blink subscription pages remain products with kind: subscription.

  • HTTP usage includes monotonic queue, robots, cooldown, transport, retry, extract, format, model and total timings. request_complete is emitted after the response body resolves, while ladder totalMs records user-visible elapsed time separately from summed attempt time.

  • MCP scrape responses are compact by default and omit trace, ladder audit and nested duplicate bodies; debug: true restores the full audit. The fixed three-round Amazon baseline writes ignored local evidence with region pinning, field checks, response bytes and p50/p95 timing.

  • Every scrape response carries the facts of the call in metadata, under Firecrawl’s names, beside the page’s own fields: scrapeId (a UUID per POST /v1/scrape or /fc/v1/scrape call, also at the top of the full response), sourceURL, url, statusCode, contentType, proxyUsed (operator for the server’s environment proxy, user for the caller’s own egress, else null), timezone (the browser lane’s declared zone; null on the HTTP lane) and the concurrency pair below; /fc adds creditsUsed: null. A scrape response’s metadata is now always present, with the page fields null on a page that was not read as content; batch items and crawl pages are unchanged. The call’s record, without any page body, is written to <task root>/scrapes/<scrapeId>.json and served by GET /v1/scrapes/:id (SDK getScrape, MCP get_scrape); an unknown id is 404 not_found, and records have no retention yet.

  • metadata.concurrencyLimited and metadata.concurrencyQueueDurationMs say whether the per-origin concurrency ceiling (W2L_PER_HOST_CONCURRENCY) held a scrape’s attempts back and for how long, cooldown and pacing excluded; each lane result carries its own hold as usage.timings.concurrencyWaitMs when there was one, and the origin scheduler’s permits report limitedByConcurrency and concurrencyWaitMs. usage.timings.queueMs keeps its meaning.

  • integration and origin on scrape, crawl and batch (REST, /fc, SDK, MCP), 1 to 100 printable characters without spaces, go into W2L’s own records only: the scrape record and the task, which GET /v1/crawl/:id and GET /v1/batches/:id report as attribution. The SDK sends origin: js-sdk@<SDK_VERSION> unless the call sets one; the MCP server records mcp-<client name>@<client version> from the client’s initialize and refuses an origin a tool call names; /fc maps the origin Firecrawl’s SDKs send instead of dropping it.

  • agentHints on scrape responses, batch items, crawl pages and the scrape record (agent_hints on /fc pages) say what to change about the request next time, one sentence each: a robots.txt rule and the recorded override, a login wall and mode: authed, a challenge or bot gate and the lanes tried, a Retry-After time, a cut, a script-filled shell and whether the browser lane had its turn, an error page kept as evidence. Refusals of stealth, a stealth or enhanced proxy, ignoreRobotsTxt and hosted skipTlsVerification carry a hint naming the supported route in the error body; W2LError.agentHints carries them in the SDK. Under fastMode the existing fastMode hint stands alone.

  • W2L_RATE_LIMIT_PER_MINUTE / --rate-limit-per-minute (1 to 100000) cap the requests that start work (POST /v1/scrape, /v1/crawl, /v1/batches, /fc/v1/scrape, /fc/v1/crawl) per bearer token in a sliding minute, in memory and per process: over it, HTTP 429 with Retry-After and { error, code: "rate_limited", retryAfterSeconds, agentHints } (/fc: { success: false, error, code, agent_hints }). The SDK’s W2LError carries status 429, code rate_limited, retryAfterMs and the hints and retries nothing; MCP tool calls fail with rate limited: retry after <s> s (rate_limited).

  • Crawl sitemaps, on REST, /fc, the SDK and MCP: sitemap (include, the default, skip or only). A crawl now reads the site’s sitemap once per attempt before its first page, the files the start URL’s robots.txt names or /sitemap.xml, so every crawl makes one more request to its host (a sitemap read, or a probe of /sitemap.xml that may answer 404), and a bounded crawl of a site with a sitemap returns different pages than before: the entries are queued at depth 1 after the start URL and ahead of its links, under the same host, subtree, path and depth rules. only follows no page link; skip restores the old behaviour. A <sitemapindex> is followed one level, a gzip file inflated under the 50 MiB cap, at most 20 files read, and the load stops once it holds maxPages entries. Sitemap files are fetched with the crawl mode’s http identity, the SSRF checks, the operator proxy, the origin scheduler’s pacing, the policy’s redirect limit and 10 MiB wire cap, each after its own URL’s robots.txt verdict (a disallowed or unreachable robots.txt refuses it), never through the ladder; they have no signed compliance record. The report’s discovery.sitemap lists every file read, refused or unreadable with its status, bytes, SHA-256, kind, entry count, robots verdict and proxy use, and a page found through the sitemap carries discovered { via: "sitemap", from: <file> }. The /fc shim maps v1 ignoreSitemap (true is skip, false is include, where false used to be refused) and sitemapOnly, and passes v2 sitemap through. A task stored before the option resumes without a sitemap; the w2l crawl CLI reads none yet.

  • maxConcurrency on a crawl (REST, /fc, SDK, MCP): an integer from 1 to the service’s worker count (4 locally, 2 on the hosted MCP host; above it HTTP 400 maxConcurrency must be at most N on this service) that lowers the pages the crawl fetches at once and never raises the per-host ceiling; stored with the task and kept on resume.

  • GET /v1/crawl/active (SDK getActiveCrawls, MCP list_active_crawls): the crawls this API process is running, each with its id, start URL, status, start time, pages so far and stored options; always 200, empty when nothing runs, never a batch.

  • SDK pagination caps: listCrawlPages and listBatchItems take maxPages, maxResults and maxWaitMs and say where they stopped (nextCursor, stoppedBy); getCrawlDocuments and getBatchDocuments return the status and the documents in one answer, collectCrawlPages and collectBatchItems the documents alone. MCP get_crawl_pages and get_batch_items take maxResults (1 to 200) and follow the cursors themselves. Server page sizes are unchanged.

  • Crawl URL scope, on REST, /fc, the SDK and MCP: crawlEntireDomain (default false: a crawl now follows links on the start URL’s host only inside the start URL’s path subtree, as Firecrawl’s default does, so a crawl seeded below the root returns fewer pages than before; true restores the whole host), allowSubdomains, allowExternalLinks (refused beside allowlistedDomains), regexOnFullURL, ignoreQueryParameters and deduplicateSimilarURLs (default true: /a and /a/, / and /index.html, www. and apex, http and https are one page, fetched once, so a crawl of a site that links both variants fetches one page fewer than before; the content-hash duplicate check stays as the backstop). allowlistedDomains now adds hosts to the start URL’s own instead of replacing them, and the crawl’s governance allowlist, when there is one, is derived from the same rule. Every crawl reports its link discovery: discovery counters on GET /v1/crawl/:id (offered, enqueued, duplicate, collapsed, hostDenied, subtreeDenied, pathDenied, depthDenied, duplicateContent; null for a batch), written with the attempt after every page, and on each page’s trace a discovered event (via, from) and a links_offered event with the page’s counters and samples of its collapsed and host-refused links. GET /v1/crawl/:id/pages leaves out pages whose body repeated an earlier page’s unless includeDuplicates=true (MCP get_crawl_pages, SDK getCrawlPages), and /fc/v1/crawl/:id leaves them out of data while total still counts them. The /fc shim maps the six options, with v1’s allowBackwardLinks as crawlEntireDomain. A crawl task stores the options; one stored before this change resumes with its original whole-host, exact-URL rule. The w2l crawl CLI has no flags for these options yet: it keeps following the whole host and takes the other defaults.

0.4.0-rc.1 — 2026-09-22

  • Gate 2–4 implementation freeze 99894bd636ecafd254a7c7bc79d26e9a97fa9199 was merged by PR #50 into main at 1c1481722ade26b717d18a34fa4b46362f53acf8 and published as the v0.4.0-rc.1 source prerelease. Workspace packages remain private; no npm package or permanent service deployment is included.

  • Gate 2: explicit captureMode, shared cancellation/deadlines, full Retry-After waits and persisted Monitor cooldown; actual process recovery/claim races, baseline/fencing, A/B/A/B, conditional-cache body and multi-Monitor isolation checks passed.

  • Gate 3: durable HTTPS destinations/delivery worker, lease/fencing, retry/dead-letter, same-event replay and a transactional deduplicating receiver. Real HTTPS ACK-loss/restart experiment passed; its temporary endpoint is stopped.

  • Gate 4: Monitor/Delivery SDKs, examples, install/restart documentation and a sanitized agent clean-install record. Independent human acceptance remains pending; Gate 5 external two-week and repeat-use validation has not started.

  • Native Crawl results expose paginated /v1/crawl/:id/pages and /v1/crawl/:id/errors, persistent cancellation via /v1/crawl/:id/cancel, and matching SDK/MCP operations.

  • Roadmap calibration: B1/B2 and C1 remain in_progress; C2 Monitor/Delivery MCP and guided first use, plus C3 unified process management and remote HTTPS URL MCP are the next unimplemented slices. Existing B3/B4/C4 gaps remain open.

  • API listen is loopback by default (hostname: 127.0.0.1). --hosted --token is the public mode: bearer auth, private/metadata SSRF deny on seed + redirects, 10 MB body cap, crawl maxPages default 100.

  • Local mode still allowlists loopback/RFC1918 so fixture servers work. API callers cannot widen that list.

  • Crawl scrape or store errors write task/attempt failed instead of leaving running. Resume with no contentful checkpoint reseeds the seed URL.

  • HTTP and local browser fills rawBodySha256. Same body on a later URL is duplicate, not a crawl-stopping loop_detected.

  • A 200 challenge page with extractable prose is blocked, not success. Decisive challenge evidence (vendor header / Cloudflare plumbing / interstitial copy pair) is consulted after extract; an embedded widget on a real article is not.

  • Ladder runs now expose task-level execution accounting: every attempted channel remains available alongside channelsTried and ladderTrace; unknown cost, token, or wire-byte measurements stay null instead of being treated as zero.

  • Browser-rendered DOM size is not reported as bytesWire; browser paths use null when actual network transfer bytes cannot be proven. artifacts: [] means this run produced no screenshot or DOM artifact.

  • Multi-page crawls use bounded workers, enforce Frontier host concurrency and robots crawl delays, reuse channels and routing history within an API engine, and reuse browser processes while keeping fetch contexts isolated.

  • Earlier foundation review baseline: main@6965168 after PR #14 and PR #15 were merged; current freeze evidence is linked in Gate 2–4 acceptance.

0.3.0 — 2026-09-18

Programmable scrape and crawl. Same runner as the CLI.

  • REST: POST /v1/scrape, POST /v1/crawl (202 + taskId), GET /v1/crawl/:id. Responses are FetchResult / CrawlReport.
  • TypeScript SDK (@w2l/sdk, MIT): W2L.scrape / W2L.crawl / W2L.getCrawl. Server remains AGPL.
  • MCP stdio server (@w2l/mcp, MIT): tools scrape, crawl, get_crawl over the REST contract. No OAuth, no resources.
  • Firecrawl v1 shim (/fc/v1/scrape, /fc/v1/crawl): snapshot 2026-09-18, maps onto the native contract. Challenge pages are not success; no fire-engine; resume defaults to refetch. Not a compatibility layer.
  • 1000-page kill/resume probe: SIGKILL at 27 pages, resume to 1001 unique URLs, 0 lost, cached=0.
  • Workspace packages versioned 0.3.0.

0.2.0 — 2026-09-18

Crawl composes scrape. Checkpoint from day one.

  • w2l crawl <url> with --resume, --use-cached, --max-pages, --headed (browser arm only; CI stays headless).
  • @w2l/runtime: TaskStore (memory + SQLite next to the task dir), frontier, orchestrator.
  • Checkpoint is task → attempt → step at URL granularity. Caller-generated UUIDs; repeat writes are idempotent.
  • Resume restores the queue. Default is refetch; --use-cached is the only skip-fetch path and marks cached pages.
  • Link harvest from the full document after extract, before HTML is dropped. Nav links are kept; markdown stays chrome-free.
  • HTTP arm honours robots.txt with the same semantics as the browser arm.
  • Loop stop: two distinct canonical URLs with the same rawBodySha256 → loop_detected.
  • Page / time / cost / token budgets can set budget_exceeded.
  • Workspace packages versioned 0.2.0.

0.1.0 — 2026-09-18

First product-shaped cut of the identity ladder.

  • Anti-bot is a coverage ladder (ADR 0004), not an in-tree circumvention engine.
  • L0 identity bundle: UA, Client Hints, locale, timezone, viewport must agree; fail closed.
  • Product HTTP arms send that bundle (standard default; research is a declared bot).
  • Ladder refuses channels with a missing or contradictory identity before fetch.
  • w2l scrape <url> is the user entry (w2l-fetch remains an alias).
  • Extract-tf emits Markdown after extraction, not raw HTML.
  • Provider lane measures vendor identity and does not inject ours; HeadlessChrome / research-as-Chrome / UA-hint mismatch are not success.
  • Changing IP or session does not change identity (identityForRoute).
  • Workspace packages versioned 0.1.0.