One link.
Web data, ready.

Readable Markdown and checked fields from public pages, ready for an AI agent, a RAG pipeline or a citation. When a page can’t be read, you get the reason.

5 free previews a day · Public pages only · Use it from your agent

Get code

Use it in your code or agent

The same page, format and options, through hosted Octocrawl (api.octocrawl.dev and mcp.octocrawl.dev: no key within a daily allowance, a key for more and the browser lane) or on your own computer with no limit (npx octocrawl serve, npx -y @octocrawl/mcp). A browser-lane run may differ from this HTTP preview.

    See how it works

    HOW IT WORKS

    From web page
    to usable content.

    1. Paste a public URL

      No install or sign-up. Paste any public http(s) address; you get five previews a day.

    2. Octocrawl checks, then reads

      It respects robots.txt and reads only what anyone can open, then reports the status, final URL and time.

    3. Use the content

      Copy or download readable Markdown or the result JSON. Amazon.sg product pages add checked fields.

    A replay of two recorded runs: the example page on this site's preview (8 Oct 2026) and an Amazon.sg product on a local Octocrawl (23 Sep 2026). Pages change, so your results may differ. Example results: https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Overview returned success in 0.87 seconds of server time, with Markdown that starts "# Overview of HTTP" and then describes HTTP as a protocol for fetching resources such as HTML documents. The Amazon.sg product B000NI69YA, a Fluke 116 HVAC Multimeter, was matched for Singapore 238823 at SGD 290.67, sold by Amazon US, in 4.00 seconds measured by the client.

    WHAT IT DOES

    One link is
    only the start.

    The preview reads one page. The same engine maps whole sites, runs batches, watches pages for changes and reads with your own logins, and every result carries its Evidence Record.

    • Scrape a page

      • Markdown, links, tables and fields
      • From one URL
      • Over HTTP, or in a browser when the page needs one

      Hosted: yesYour computer: yes

    • Map a site

      • The URLs a site lists in its sitemap and links
      • Before you read a single page

      Hosted: yesYour computer: yes

    • Crawl and batch

      • Follow a site’s links
      • Or read a list of up to 1,000 URLs
      • A crawl that stops resumes where it left off

      Hosted: noYour computer: yes

    • Watch a page

      • A Monitor re-reads a page
      • Sends what changed to your HTTPS webhook, signed with your secret
      • Timed re-runs need a repository checkout

      Hosted: noYour computer: yes

    • Your own Chrome

      • Read pages signed in as you
      • A page that stops at a check goes to your Chrome
      • You get past it yourself

      Hosted: noYour computer: yes

    • An Evidence Record

      • Every result says where it came from
      • The final URL, HTTP status and robots.txt decision
      • A hash of what was read

      Hosted: yesYour computer: yes

    GET STARTED

    One line in,
    clean pages out.

    Use Octocrawl from your terminal, your code or your AI agent. Pick the one that fits how you work.

    terminal
    # Read one page as Markdown: nothing to install, no key
    npx octocrawl scrape https://example.com --markdown
    This domain is for use in documentation examples without needing permission. …
    
    # List the pages a site links to
    npx octocrawl map https://example.com
    
    # Read a list of pages and keep every result with its evidence
    npx octocrawl batch https://example.com https://example.org --out results
    ls results
    0001-example.com.md  0002-example.org.md  report.json  results.csv  results.jsonl

    Runs on your computer with Node.js 22.13 or later: no server, no key, no daily limit. Advanced reference

    # Add hosted Octocrawl to Claude Code: one URL, no key to start
    claude mcp add --transport http octocrawl https://mcp.octocrawl.dev/mcp
    
    # Then ask your agent
    Use Octocrawl to read https://example.com and give me its Markdown.

    Cursor, OpenCode and Codex take the same URL. Connect MCP

    // npm install @octocrawl/sdk
    import { W2L } from '@octocrawl/sdk'
    
    const octo = new W2L({ baseUrl: 'https://api.octocrawl.dev' })
    const page = await octo.scrape('https://example.com')
    console.log(page.markdown)

    Hosted, keyless within the daily allowance; or http://127.0.0.1:8787 after npx octocrawl serve. Advanced reference

    # pip install octocrawl-client
    from octocrawl_client import W2L
    
    octo = W2L(base_url="https://api.octocrawl.dev")
    page = octo.scrape("https://example.com")
    print(page["markdown"])

    Hosted, keyless within the daily allowance; or http://127.0.0.1:8787 after npx octocrawl serve. Advanced reference

    # Hosted: no key within the daily allowance
    curl -sS -X POST https://api.octocrawl.dev/v1/scrape \
      -H 'content-type: application/json' \
      -d '{"url":"https://example.com","formats":["markdown"]}'

    Add -H 'authorization: Bearer <key>' for more pages a day. Limits and result states

    FREE TIERS

    Free to start,
    four ways.

    None of them needs a card. Each mark that lights up on the planet stands for a page a day.

    Browser preview5 a day

    previews a day, per visitor

    Where
    This page
    Needs
    Nothing: no account, no install
    Reads
    One public page at a time
    Try a page

    Hosted, no key20 a day

    pages a day, per address

    Where
    api.octocrawl.dev, mcp.octocrawl.dev
    Needs
    Nothing
    Reads
    Scrape and map over HTTP, 10 a minute
    Connect MCP

    Hosted, with a key1,000 a day

    pages a day, to start

    Where
    The same two addresses
    Needs
    A key, issued by hand for now
    Reads
    HTTP, then a browser when the key allows it, 60 a minute
    Ask for a key

    Your computer∞ no daily limit

    no daily limit

    Where
    npx octocrawl serve
    Needs
    Node.js 22.13 or later; for a browser, npx playwright install chromium
    Reads
    Everything above; timed Monitors from a repository checkout
    Run it yourself

    The whole hosted service serves 1,500 pages a day; over an allowance the answer is HTTP 429 until 00:00 UTC. Credit packs are planned but not sold yet. Limits and result states

    FAQ

    Questions,
    answered plainly.

    Does Octocrawl respect robots.txt? Is this legal?

    We can’t give legal advice; here is what the preview does. It reads a site’s robots.txt before it fetches a page. If the page is disallowed, or robots.txt can’t be reached (a server error, no answer or a timeout), it stops and reports the page as blocked; a robots.txt that answers with a 4xx status counts as no rules, as RFC 9309 provides. It never signs in, solves a CAPTCHA or gets past a verification page, and it refuses private network addresses. What you do with a page is up to you and the site’s terms.

    Which sites work?

    Public pages anyone can open without signing in. The preview reads them over HTTP without running JavaScript, so a page that only appears in a browser may come back incomplete. It reads pages up to 2 MiB and files such as PDFs up to 5 MiB, and stops after 40 seconds. Amazon.sg product pages (/dp/ASIN) are in Beta. X and Reddit posts often don’t come through: robots rules, sign-in walls or verification pages can stop the preview, and a hosted X or Reddit result hasn’t been verified yet. See Limits and result states.

    Why only five previews a day?

    The preview is a limited public trial: five previews per visitor and 150 for the whole site each UTC day. A request turned down before a preview starts, such as a malformed URL, or localhost or a private IP address typed into it, doesn’t count. Once a preview starts it counts, whatever the result, including a host name that turns out to point to a private network or a page stopped by robots.txt. Octocrawl on your own computer has no daily limit.

    Do you store the URLs I submit, or the results?

    Results aren’t saved: your recent runs live in this page and are gone when you leave it. Each preview logs its state and the host of the page, never its path. While you type, the page asks the service for a short hint about the address, so that address appears in our hosting provider’s request log, kept for 30 days. The details are on the Privacy page.

    How is Octocrawl different from Firecrawl or Crawl4AI?

    Octocrawl reports what it actually read: a blocked, incomplete or timed-out page is a result with a reason, and checked fields carry their source. The benchmark notes compare the three tools on the same test suite, with the limits of that comparison. For moving off Firecrawl, Octocrawl has a partial, local Firecrawl v1 shim.

    Can I call it from a script or an agent?

    Not this preview page: it is for trying Octocrawl in a browser. Call hosted Octocrawl instead: https://api.octocrawl.dev/v1/scrape over REST, or https://mcp.octocrawl.dev/mcp from Claude Code, Cursor, OpenCode or Codex, keyless within a daily allowance and with a key for more. Or run Octocrawl yourself with no limit.

    Is it free?

    The preview and the keyless hosted allowance are free and need no account. A key, issued on request, gives more pages a day and the browser lane; credit packs are planned but not sold yet. Octocrawl is open source under the AGPL-3.0, and running it yourself costs nothing but your own machine.