Advanced reference
Use the browser preview or Codex MCP setup for a first result. REST, SDK, and self-hosted operation are available for developers who need explicit configuration and persistent task control; they currently require a checkout of this repository.
REST and SDK
The local API has POST /v1/scrape for one URL, POST /v1/batches for an explicit URL array, and GET /v1/batches/:id/items?limit=...&cursor=... for paginated outcomes. Monitor and Delivery have separate REST resources. The @w2l/sdk package is currently a private workspace package, not an independently published npm install.
A server started with tokens (--token, which can be repeated, or W2L_API_TOKEN and the comma-separated W2L_API_TOKENS) accepts a request only with Authorization: Bearer <token> naming one of them. Give each client its own token; restarting the server without a token revokes it. Tokens are compared as fixed-length SHA-256 digests in constant time. The SDK sends its token option, or W2L_API_TOKEN from the environment when no token is passed.
Scrape results, batch items and crawl pages carry metadata: the page’s <title>, <meta name="description">, language (<html lang>, else Content-Language), <meta name="keywords">, <meta name="robots">, the first <link rel="icon"> as an absolute URL, and <link rel="canonical">. Each is null when the page does not declare it. document.title is a different value: the content’s own title, usually its first heading.
A URL that answers with a file (PDF, CSV, JSON, plain text, XLSX, XLS or ZIP, by its Content-Type or, under application/octet-stream, by its bytes) is saved byte for byte as received under the task root, files/<sha256>.<ext>, and is never sent to the browser. The result’s file block says what it is (kind, detectedBy, contentType), its declared and received size (declaredBytes, bytes), the cap that applied (maxBytes), sha256 and path, and markdownFrom: pdf_text, text or null. A PDF’s markdown is its text layer with a line <!-- page N --> before each page, and file.pdf gives pageCount, pagesRead, each page’s number, printed label and start/end offsets in the Markdown, the PDF’s own info, and warnings (tables_unverified whenever there is text, no_text_layer per page without text: no OCR is run, page_cap, time_budget, page_error). A PDF with text is success, or partial when a cap or budget stopped it; one with no text layer at all is failed with empty_unverified; one that cannot be opened is failed with parse_error. CSV, JSON and text files give their text as markdown; XLSX, XLS and ZIP give none and are success. The operator’s W2L_MAX_FILE_BYTES (default 50 MiB, at most 500 MiB) caps a file; maxFileBytes on scrape, batch or crawl can lower it, and a file over the cap is failed with body_too_large and nothing saved. Images, audio, video, fonts and word-processor or presentation documents are failed with unsupported_content_type.
JSON extraction
A json format takes a JSON Schema: { "type": "json", "schema": { … } }, with prompt and modelFallback optional. OctoCrawl fills the schema from the page itself, and every value carries its evidence. A format without a schema is refused (invalid_request): OctoCrawl does not extract JSON from a prompt alone, which would need a model for every page. prompt only instructs the model fallback.
The schema may use this subset, which covers what Pydantic’s model_json_schema() and zod-to-json-schema usually write:
| Keywords | What OctoCrawl does with them |
|---|---|
type (one or a list), properties, required, items (one schema), additionalProperties, enum, const, $defs, definitions, local $ref (# or a pointer into the schema) |
Maps values by them and checks the result. |
anyOf / oneOf of a schema and { "type": "null" }, or of primitive types only |
Maps the value by the non-null schema or the types; a missing nullable field is null with a field_unavailable issue. |
minimum, maximum, exclusiveMinimum, exclusiveMaximum, multipleOf, minLength, maxLength, pattern, minItems, maxItems, uniqueItems |
Checks the result; never fills a value in. A page value that breaks a check makes the result incomplete with a field_unavailable issue. A pattern has at most 2,000 characters; one that can backtrack catastrophically, such as ^(a+)+$, is refused with invalid_request. A pattern with a lookaround, a backreference, a counted repetition above 16 (such as [0-9a-f]{32}), a \p{…} or \u{…} escape or a character outside the Basic Multilingual Plane, and any text holding such a character, is matched only against text of up to 2,048 UTF-16 code units within 100 ms; text it cannot decide counts as breaking it. |
title, description, $comment, examples, deprecated, readOnly, writeOnly, format, default; $schema (draft-07, 2019-09 or 2020-12) and $id at the root only |
Accepts and ignores them: format is not checked and default is never filled in. |
A schema is at most 64 KiB, 8 levels deep and 100 properties. Any other keyword, a union of objects, arrays or references, a keyword other than an annotation beside $ref, and $schema or $id below the root are refused with unsupported_parameter; details.parameters names the keyword where it was sent, such as formats[0].schema.properties.author.allOf.
A result is complete, incomplete (a required field has no source: missing_required; or the page was not a success: page_unsuccessful, page_partial) or invalid (the model’s answer did not fit the schema twice). json.evidence, like evidenceRecord.fieldEvidence, gives each value’s source: a label row, h1[0] or title for the content title (dom), finalUrl or requestedUrl (fetch), document.pageType (inferred), document.product.<name> for a product list or map the extractor reported empty (inferred), or model. For a number read from the page’s text, json.evidence also quotes that text (text).
Numbers are read as the page writes them: a decimal . or ,; thousands grouped by ., ,, a space or an apostrophe in groups of three (or India’s lakh groups); a currency symbol or code before or after; ,- for a whole amount. 12,99 € is 12.99, and 1.299,00 €, 1 299,00 € and CHF 1'299.– are 1299. A single . or , before exactly three digits (1.299 €, $1,299) is read only when the value settles it: a review count, an amount in a currency without minor units (JPY, KRW, ₩, 円, …), or a JSON-LD or product:price:amount price, whose format writes . as the decimal point. The page’s language, currency or domain is not used to guess. Otherwise the field is left out (null when nullable) with a field_unavailable issue quoting the text. The model fallback (modelFallback: true and W2L_EXTRACT_BASE_URL / W2L_EXTRACT_MODEL set) fills only what the page did not give, or a page value that breaks the schema; every other value keeps what the page said and its evidence. It asks the provider for strict structured outputs with a strict-safe copy of the schema, or for the schema as given when strict mode cannot express it; json.modelUsage.strict and strictReason say which and why.
Start the repository API only after reviewing its network and task-store settings. For full request shapes and examples, use the repository’s docs/onboarding.md, docs/batch-scrape.md, and examples/monitor-workflow.ts from the same checkout and commit as the running service. Mixing a guide from another branch with a local server can change the apparent contract.
Evidence Record
Every scrape result (full or compact, so MCP scrape too), batch item and crawl page carries evidenceRecord: the evidence for that page in one shape, whichever lane fetched it. Its JSON Schema (draft 2020-12) is packages/contracts/schemas/evidence-record.v1.json in the repository, which describes every field; each record names it with schemaVersion: "w2l.evidence/1". Every field is always present, and a value OctoCrawl did not observe is null, never 0 or a guess. The record sits beside the existing fields (evidence, snapshot, compliance, lane), which do not change. The Firecrawl-compatible /fc routes keep Firecrawl’s shape and do not carry it.
w2l.evidence/1 changes only additively: a later revision of the schema file accepts every record an earlier one accepted. A field added later is optional in the schema, so older records stay valid, though OctoCrawl writes it on every record from then on (so far bytes and contentType on artifacts, added with file download); enums only gain values; no field changes its meaning. Anything else would get a new schemaVersion and schema file.
| Field | Meaning |
|---|---|
schemaVersion |
w2l.evidence/1. |
requestedUrl |
The URL as requested. |
finalUrl |
The last URL OctoCrawl requested for the page, after redirects. On the browser lane a script or a meta refresh that loaded another document is a redirect; a URL the page set with the history API (pushState, replaceState), which nothing requested, is not. null when it sent none: a robots.txt disallow, a DNS failure or a policy refusal came first. |
redirectChain |
{ urls, complete }. urls runs from requestedUrl to finalUrl ([requestedUrl] without a redirect, [] when nothing was requested). complete is true when every hop is listed: on the HTTP lane, which requests each hop itself, and on the browser lane, which lists each redirect Chromium followed and each document a script or a meta refresh loaded (false there when a follow-up navigation did not start at requestedUrl, a document came without a request, or the chain of a page that kept moving on was cut to its first URL and last 20). It is false on the provider lane, which sees only where its vendor started and ended. |
fetchedAt |
When the reported response was received, UTC ISO 8601 with milliseconds: its headers on the HTTP lane, the capture of the rendered page on browser lanes, the vendor’s answer on the provider lane. null without a response. |
httpStatus |
The status of the response that answered finalUrl: on the browser lane, of the document the page shows when it is captured, not of the first navigation. null without one. |
status, reason |
OctoCrawl’s verdict, and the failure, block or budget reason (null for other statuses). |
lane |
http, browser_local, browser_local_authed, browser_proxy or provider. |
robotsDecision |
{ decision, robotsUrl, robotsSha256, unreachable, crawlDelayMs, userOverride }, or null when the result ended before robots.txt was checked. decision is allowed, disallowed or no_robots (the site has none); unreachable says why robots.txt could not be fetched, which counts as a disallow. userOverride is false: OctoCrawl has no override. |
rawSha256 |
SHA-256 of the body OctoCrawl read, as evidence.rawBodySha256: for a file (PDF, CSV, …), the bytes received in any lane, equal to its artifact’s sha256; for a web page, the response body as UTF-8 text on the HTTP lane, the rendered HTML on browser lanes, the vendor’s HTML on the provider lane. |
outputSha256 |
{ markdown, json }: SHA-256 of the UTF-8 bytes of the markdown delivered in this response, and of json.data as canonical JSON (RFC 8785: keys sorted at every level, no whitespace). null for what was not delivered. |
extractor |
{ name, version, commit }: extract-tf and its version (extract-tf/4) for a web page, pdf-text (pdf-text/1) for a PDF, file-text (file-text/1) for another file; and the source commit when the operator sets W2L_SOURCE_COMMIT, else null. |
fieldEvidence |
For a JSON request, each field’s JSON Pointer mapped to { source, locator }: the source is jsonld, microdata, meta, hydration, dom, text, inferred, model, fetch (the fetch’s own URL, not the page) or pdf, and the locator says where, such as table[0] tr[3] "UPC", h1[0], title, finalUrl, or page 14 "KPI 2" for a Label: value line of a PDF (page N is the <!-- page N --> marker). JSON-LD, microdata and meta values outside the Amazon adapter have no locator yet (null), nor has model. A field that is absent or null has no entry. null without a JSON request. |
artifacts |
Files saved for the result, each { kind, path, sha256, bytes, contentType }: a file saved as received (kind: "file", with its size and Content-Type), and the page snapshot (kind: "snapshot", size and type null) when W2L_CAPTURE_RAW_DIR is set; otherwise []. |
proxy |
host:port of the environment proxy the request went through (local mode). null when it went direct or the lane does not report its route, which the provider lane never does. |
identity |
{ userAgent, mode, contact }: the User-Agent observed on the request for the page (null when none was sent or observed), the crawl mode, and the contact a research-mode User-Agent declares (W2L_CONTACT). |
To check a cited value, hash the markdown you received as UTF-8 and compare it with outputSha256.markdown; for JSON, serialize json.data with sorted keys and no whitespace first.
Errors
When the API refuses or fails a request, it returns an error body instead of a result:
{ "error": "unsupported format: html (supported: markdown, links, json)", "code": "unsupported_format", "details": { "formats": ["html"] } }Branch on code; error is written for people. details appears only with unsupported_parameter and unsupported_format and lists the rejected names as sent, in parameters and formats. The Firecrawl shim (/fc) returns the same fields after "success": false.
| Code | HTTP | Meaning | Typical cause | What to do |
|---|---|---|---|---|
invalid_json |
400 | The body is not valid JSON. | A truncated or hand-edited body. | Send one JSON object. |
invalid_request |
400 | A value is missing, malformed or out of range. | No url, a non-HTTP(S) URL, an invalid includePaths pattern, a json format without a schema, a JSON Schema over its limits or with a malformed value. Also, for now, a batch refused because another is still being accepted or the active-batch limit is reached. |
Correct what the message names; retry a refused batch later. |
unsupported_parameter |
400 | The request names a parameter OctoCrawl does not support, or a value it cannot honour. Nothing is ignored silently. | A Firecrawl option such as actions or proxy, limit on native crawl (it takes maxPages), ignoreSitemap: false on /fc, a JSON Schema keyword outside the supported subset such as allOf. |
Remove or change what details.parameters lists. |
unsupported_format |
400 | A requested format is not produced. | html, rawHtml or screenshot. |
Ask for markdown, links or json (on /fc: markdown, links). |
unauthorized |
401 | The bearer token is missing or is not one of the server’s tokens. | A server started with --token, W2L_API_TOKEN or W2L_API_TOKENS, which --hosted requires. |
Send Authorization: Bearer <token>; the SDK reads W2L_API_TOKEN when no token is passed. |
not_found |
404 | The task, monitor, run, destination, delivery or session does not exist, or no route matches. | A mistyped ID, another task directory, a Firecrawl v2 path on /fc. |
Check the ID and the path. |
conflict |
409 | The resource’s current state does not allow the request. | A monitor revision out of sequence or on a running monitor; a run queued on a paused, busy or unknown monitor; a retry of a delivery that is not dead-lettered; resuming a crawl that is completed, cancelled or still running. | Read the resource’s state first; the same request fails again. |
internal_error |
500 | OctoCrawl failed unexpectedly. | A bug or a storage failure. | Retry once, then report it. A local server returns the underlying message; a hosted server (--hosted, remote MCP) returns internal error and writes the cause to its log. |
These codes describe the request, not the page. A page that was fetched but blocked or failed is a normal result with its status and reason (see result states); on /fc it is HTTP 200 with success: false, no code, and the reason in data.metadata.error.
The SDK throws W2LError for every error response. It carries status, code (when the body had one), method, path and the parsed body; its message is still <METHOD> <path> failed: <status> <body>, or crawl not found: <id> and batch not found: <id> for those lookups. A request that never reached the API throws the underlying fetch error. scrape waits for the answer until the request’s timeout (default 300 000 ms) plus 30 s; on Node this replaces fetch’s own 300 s wait for the response headers, so the API’s failed/timeout answer at the deadline arrives instead of a TypeError. waitBatch / waitCrawl (and batchAndWait / crawlAndWait) retry a status request that fails with a network error, 408, 429 or 5xx, up to maxRetries (default 5) in a row, before throwing its error; they throw other errors at once, and WaitTimeoutError (taskId, timeoutMs, last: the last status read or null) when timeoutMs runs out.
import { W2LError } from '@w2l/sdk'
try {
await w2l.getCrawl(taskId)
} catch (error) {
if (!(error instanceof W2LError) || error.code !== 'not_found') throw error
// The ID is wrong or belongs to another task directory.
}MCP tool errors start with the same code, for example unsupported_format: POST /v1/scrape failed: 400 {"error":...}, and their JSON-RPC error data is { "code": "unsupported_format", "status": 400 }. Successful tool results are unchanged.
Self-hosted operation
The local managed MCP service starts the API, scheduler, and delivery worker together; task state is SQLite-backed. Keep its task directory across restarts. The anonymous page preview is a separate request-based service and does not run persistent Monitor or Delivery tasks. A remote owner-only MCP implementation exists, but its WorkOS login, public URL, and hosted restart acceptance have not been completed.
If you are evaluating a hosted deployment, verify the authentication resource identifier, HTTPS receiver, persistent disk, outbound restrictions, quota store, and actual client flow before sharing a link. Do not point another user’s Codex installation at the loopback URL; 127.0.0.1 refers to their own computer.
Evidence and boundaries
The result states page explains user-facing outcomes. Repository evidence separates tested local transport, delivery recovery, Amazon field review, and still-open hosted gates. A green unit test or local sample is not evidence of a deployed service or an independent first-time user completing the flow.