One link.
Web data, ready.

Readable Markdown and checked fields from public pages, ready for an AI agent, a RAG pipeline or a citation. When a page can’t be read, you get the reason.

3 free previews a day · Public pages only · Use it from your agent

Get code

Run it on your computer

The same page, format and options, with no daily limit. It needs a checkout of the OctoCrawl repository. A local run can also use a local browser, so its result may differ from this preview.

    See what comes back

    RECORDED RESULTS

    What comes back.

    Readable Markdown, product fields checked against the page, and a refusal with its reason. Three results OctoCrawl recorded, shown as they were.

    requestedUrl
    https://docs.firecrawl.dev/introduction
    status
    success
    title
    Introduction
    finalUrl
    same as requested
    totalMs
    2509server time
    markdown
    Get Started
    # Introduction
    …
    the excerpt that was recorded
    Recorded 24 Sep 2026 on a local OctoCrawl at source commit 936fdf0.
    asin
    B000NI69YAmatched the captured page
    title
    Fluke 116 HVAC Multimeter, Standardmatched the captured page
    price
    290.67matched the captured page
    currency
    SGDmatched the captured page
    seller
    Amazon USmatched the captured page
    region
    Singapore 238823shown on the captured page
    Amazon.sg product B000NI69YA, recorded 23 Sep 2026 on a local OctoCrawl at source commit 991097f. Each value was matched against the captured page and signed off in a 100-product review, where 99 of 100 products came back complete and the one that did not withheld its fields. Amazon.sg support is in Beta.
    requestedUrl
    https://www.linkedin.com/feed/
    status
    blocked
    reason
    This site does not allow automated preview of this page.
    totalMs
    572server time
    Recorded 24 Sep 2026 on a local OctoCrawl at source commit 936fdf0. OctoCrawl stopped at the robots.txt check, before fetching the page, and returned no content instead of an empty page. At that commit the same reason was also given when robots.txt could not be read, and this record does not say which; OctoCrawl now reports the two apart.

    HOW IT WORKS

    From web page
    to usable content.

    1. Paste a public URL

      No install or sign-up. Paste any public http(s) address; you get 3 previews a day.

    2. OctoCrawl checks, then reads

      It respects robots.txt and reads only what anyone can open, then reports the status, final URL and time.

    3. Use the content

      Copy or download readable Markdown or the result JSON. Amazon.sg product pages add checked fields.

    A replay of two recorded runs on a local OctoCrawl: the example page (24 Sep 2026) and an Amazon.sg product (23 Sep 2026). Pages change, so your results may differ. Example results: https://docs.firecrawl.dev/introduction returned success in 2.51 seconds of server time, with Markdown that starts "Get Started" and "# Introduction". The Amazon.sg product B000NI69YA, a Fluke 116 HVAC Multimeter, was matched for Singapore 238823 at SGD 290.67, sold by Amazon US, in 4.00 seconds measured by the client.

    WHY OCTOCRAWL

    Can’t read a page?
    OctoCrawl tells you why.

    A crawler that only checks for a response can hand your agent a login wall, a challenge page or an empty shell as if it were the page. OctoCrawl reports what it actually read.

    • Fields you can check

      Fields you ask for are read from the page’s own JSON-LD, microdata, meta tags and tables, without a model, and each names where it came from. A field the page doesn’t state stays empty, with the reason.

    • Failures you can see

      A page OctoCrawl can’t read comes back blocked, incomplete, timed out or failed, with a reason and a diagnostic code. The preview reads robots.txt first, and never signs in or solves a CAPTCHA.

    • Open source, on your machine

      AGPL-3.0. Run it yourself with no daily limit, through your own network and proxy (HTTPS_PROXY), with a local Chromium for pages that need a browser.

    Measured on our own test sets

    0.0% false success
    on our 56-case test suite: no page OctoCrawl reported as read failed a ground-truth check.CI run 35423895294 · main@6dc2e6e · 19 Sep 2026
    103 of 106
    real-site cases passed, including sites where the right answer is an honest blocked.Local API · main@6024703 · 30 Sep 2026
    62 of 72
    URLs from our first research user were read. Each of the other 10 came back with its reason.Local API · main@6024703 · 30 Sep 2026

    Our own suites, not an independent audit. Method, and how other tools did on the same suite

    RUN IT YOURSELF

    Three a day here.
    Unlimited on yours.

    Clone the repository, start OctoCrawl on your computer, and call it from your agent through MCP, from your code through REST or the SDK, or from a Firecrawl v1 client. Results stay on your machine.

    1. git clone https://github.com/77777R7/w2l.git
    2. cd w2l && npm ci
    3. npx playwright install chromium
    4. npm run api
    5. # in another terminal
    6. curl -sS -X POST http://127.0.0.1:8787/v1/scrape \
    7. -H 'content-type: application/json' \
    8. -d '{"url":"https://example.com"}'

    Node.js 22.13 or later. The API listens on this computer only.

    1. # in the OctoCrawl checkout: the local MCP service
    2. npm run local:mcp
    3. # Codex
    4. codex mcp add w2l-local --url http://127.0.0.1:8791/mcp
    5. # Claude Code
    6. claude mcp add --transport http --scope local \
    7. w2l-local http://127.0.0.1:8791/mcp

    Local only: the service listens on 127.0.0.1. Verified with Codex; setups for Claude Code, Cursor and OpenCode are in the guide.

    1. // in the OctoCrawl checkout, with npm run api running;
    2. // run with: node --import tsx your-script.ts
    3. import { OctoCrawl } from '@w2l/sdk'
    4. const w2l = new OctoCrawl({ baseUrl: 'http://127.0.0.1:8787' })
    5. const page = await w2l.scrape('https://example.com', { debug: false })
    6. console.log(page.status, page.markdown)

    Local only. @w2l/sdk lives in the repository; it is not on npm yet.

    1. # a Firecrawl v1 client can point at your local OctoCrawl
    2. curl -sS -X POST http://127.0.0.1:8787/fc/v1/scrape \
    3. -H 'content-type: application/json' \
    4. -d '{"url":"https://example.com","formats":["markdown"]}'

    Partial and local only: Firecrawl v1 scrape and crawl, Markdown and links. No search, map, extract or v2. A migration aid, not a full compatibility layer.

    WHAT WORKS TODAY

    What’s here, what’s next.

    What you can use now, on this page and on your own computer, and what the roadmap has next or on hold.

    On this page

    • One public page per preview, three a day
    • Markdown, links, page info and up to 20 fields you name
    • PDFs and other files up to 5 MiB
    • Amazon.sg product records (Beta)
    • No account; results are not saved

    On your computer

    • REST API: scrape, batches of up to 1,000 URLs, crawl with resume
    • An evidence record with every result
    • Local MCP, verified with Codex
    • TypeScript SDK, from the repository
    • Monitor with HTTPS webhook delivery
    • Firecrawl v1 scrape and crawl (partial)
    • Your own proxy through HTTPS_PROXY

    Next

    • One-line install with npx
    • npm packages and a Python client
    • map and a maxAge cache
    • Tables to CSV
    • A browser extension that reads pages in your own browser

    On hold

    • Hosted API and hosted MCP: until people need runs while their computer is off
    • Search and agent features: until a paying user asks

    Not planned: stealth, fingerprint spoofing or proxy pools; getting past logins or CAPTCHAs; scraping sales leads. Roadmap

    FAQ

    Questions, answered plainly.

    Does OctoCrawl respect robots.txt? Is this legal?

    We can’t give legal advice; here is what the preview does. It reads a site’s robots.txt before it fetches a page. If the page is disallowed, or robots.txt can’t be reached (a server error, no answer or a timeout), it stops and reports the page as blocked; a robots.txt that answers with a 4xx status counts as no rules, as RFC 9309 provides. It never signs in, solves a CAPTCHA or gets past a verification page, and it refuses private network addresses. What you do with a page is up to you and the site’s terms.

    Which sites work?

    Public pages anyone can open without signing in. The preview reads them over HTTP without running JavaScript, so a page that only appears in a browser may come back incomplete. It reads pages up to 2 MiB and files such as PDFs up to 5 MiB, and stops after 40 seconds. Amazon.sg product pages (/dp/ASIN) are in Beta. X and Reddit posts often don’t come through: robots rules, sign-in walls or verification pages can stop the preview, and a hosted X or Reddit result hasn’t been verified yet. See Limits and result states.

    Why only three previews a day?

    The preview is a limited public trial: three previews per visitor and 100 for the whole site each UTC day. A request turned down before a preview starts, such as a malformed URL, or localhost or a private IP address typed into it, doesn’t count. Once a preview starts it counts, whatever the result, including a host name that turns out to point to a private network or a page stopped by robots.txt. OctoCrawl on your own computer has no daily limit.

    Do you store the URLs I submit, or the results?

    Results aren’t saved: your recent runs live in this page and are gone when you leave it. Each preview logs its state and the host of the page, never its path. While you type, the page asks the service for a short hint about the address, so that address appears in our hosting provider’s request log, kept for 30 days. The details are on the Privacy page.

    How is OctoCrawl different from Firecrawl or Crawl4AI?

    OctoCrawl reports what it actually read: a blocked, incomplete or timed-out page is a result with a reason, and checked fields carry their source. The benchmark notes compare the three tools on the same test suite, with the limits of that comparison. For moving off Firecrawl, OctoCrawl has a partial, local Firecrawl v1 shim.

    Can I call the preview from a script?

    Please don’t: the preview is for trying OctoCrawl in a browser. To automate, run OctoCrawl yourself and use REST, the SDK or MCP.

    Is it free?

    The preview is free and needs no account. OctoCrawl is open source under the AGPL-3.0, and running it yourself costs nothing but your own machine. There is no paid plan today.