One link.
Web data, ready.
Readable Markdown and checked fields from public pages, ready for an AI agent, a RAG pipeline or a citation. When a page can’t be read, you get the reason.
Readable Markdown and checked fields from public pages, ready for an AI agent, a RAG pipeline or a citation. When a page can’t be read, you get the reason.
YOUR RESULTS
This visit only · cleared when you leave the page
Use it in your agent: Connect MCP · Run it yourself
SELECTED RUN
HOW IT WORKS
No install or sign-up. Paste any public http(s) address; you get five previews a day.
It respects robots.txt and reads only what anyone can open, then reports the status, final URL and time.
Copy or download readable Markdown or the result JSON. Amazon.sg product pages add checked fields.
Extract structured data from a page Limits and result states
WHAT IT DOES
The preview reads one page. The same engine maps whole sites, runs batches, watches pages for changes and reads with your own logins, and every result carries its Evidence Record.
Hosted: yesYour computer: yes
Hosted: yesYour computer: yes
Hosted: noYour computer: yes
Hosted: noYour computer: yes
Hosted: noYour computer: yes
Hosted: yesYour computer: yes
GET STARTED
Use Octocrawl from your terminal, your code or your AI agent. Pick the one that fits how you work.
# Read one page as Markdown: nothing to install, no key npx octocrawl scrape https://example.com --markdown This domain is for use in documentation examples without needing permission. … # List the pages a site links to npx octocrawl map https://example.com # Read a list of pages and keep every result with its evidence npx octocrawl batch https://example.com https://example.org --out results ls results 0001-example.com.md 0002-example.org.md report.json results.csv results.jsonl
Runs on your computer with Node.js 22.13 or later: no server, no key, no daily limit. Advanced reference
# Add hosted Octocrawl to Claude Code: one URL, no key to start claude mcp add --transport http octocrawl https://mcp.octocrawl.dev/mcp # Then ask your agent Use Octocrawl to read https://example.com and give me its Markdown.
Cursor, OpenCode and Codex take the same URL. Connect MCP
// npm install @octocrawl/sdk import { W2L } from '@octocrawl/sdk' const octo = new W2L({ baseUrl: 'https://api.octocrawl.dev' }) const page = await octo.scrape('https://example.com') console.log(page.markdown)
Hosted, keyless within the daily allowance; or http://127.0.0.1:8787 after npx octocrawl serve. Advanced reference
# pip install octocrawl-client from octocrawl_client import W2L octo = W2L(base_url="https://api.octocrawl.dev") page = octo.scrape("https://example.com") print(page["markdown"])
Hosted, keyless within the daily allowance; or http://127.0.0.1:8787 after npx octocrawl serve. Advanced reference
# Hosted: no key within the daily allowance curl -sS -X POST https://api.octocrawl.dev/v1/scrape \ -H 'content-type: application/json' \ -d '{"url":"https://example.com","formats":["markdown"]}'
Add -H 'authorization: Bearer <key>' for more pages a day. Limits and result states
FREE TIERS
None of them needs a card. Each mark that lights up on the planet stands for a page a day.
previews a day, per visitor
pages a day, per address
api.octocrawl.dev, mcp.octocrawl.devpages a day, to start
no daily limit
npx octocrawl servenpx playwright install chromiumThe whole hosted service serves 1,500 pages a day; over an allowance the answer is HTTP 429 until 00:00 UTC. Credit packs are planned but not sold yet. Limits and result states
FAQ
We can’t give legal advice; here is what the preview does. It reads a site’s robots.txt before it fetches a page. If the page is disallowed, or robots.txt can’t be reached (a server error, no answer or a timeout), it stops and reports the page as blocked; a robots.txt that answers with a 4xx status counts as no rules, as RFC 9309 provides. It never signs in, solves a CAPTCHA or gets past a verification page, and it refuses private network addresses. What you do with a page is up to you and the site’s terms.
Public pages anyone can open without signing in. The preview reads them over HTTP without running JavaScript, so a page that only appears in a browser may come back incomplete. It reads pages up to 2 MiB and files such as PDFs up to 5 MiB, and stops after 40 seconds. Amazon.sg product pages (/dp/ASIN) are in Beta. X and Reddit posts often don’t come through: robots rules, sign-in walls or verification pages can stop the preview, and a hosted X or Reddit result hasn’t been verified yet. See Limits and result states.
The preview is a limited public trial: five previews per visitor and 150 for the whole site each UTC day. A request turned down before a preview starts, such as a malformed URL, or localhost or a private IP address typed into it, doesn’t count. Once a preview starts it counts, whatever the result, including a host name that turns out to point to a private network or a page stopped by robots.txt. Octocrawl on your own computer has no daily limit.
Results aren’t saved: your recent runs live in this page and are gone when you leave it. Each preview logs its state and the host of the page, never its path. While you type, the page asks the service for a short hint about the address, so that address appears in our hosting provider’s request log, kept for 30 days. The details are on the Privacy page.
Octocrawl reports what it actually read: a blocked, incomplete or timed-out page is a result with a reason, and checked fields carry their source. The benchmark notes compare the three tools on the same test suite, with the limits of that comparison. For moving off Firecrawl, Octocrawl has a partial, local Firecrawl v1 shim.
Not this preview page: it is for trying Octocrawl in a browser. Call hosted Octocrawl instead: https://api.octocrawl.dev/v1/scrape over REST, or https://mcp.octocrawl.dev/mcp from Claude Code, Cursor, OpenCode or Codex, keyless within a daily allowance and with a key for more. Or run Octocrawl yourself with no limit.
The preview and the keyless hosted allowance are free and need no account. A key, issued on request, gives more pages a day and the browser lane; credit packs are planned but not sold yet. Octocrawl is open source under the AGPL-3.0, and running it yourself costs nothing but your own machine.