OctoCrawl.
Browse docs: Extract a public page

OctoCrawl / Guides

Extract a public page

Use the OctoCrawl page when you want readable content, links, or a few fields from one public HTTP(S) URL without setting up a task. The browser submits the URL and, only when you change them, the two options below; capture settings and service secrets remain on the server.

Input

Open Try OctoCrawl, paste https://docs.firecrawl.dev/introduction, and press Extract page. This public documentation URL is the first-use example. When this site is served over public HTTPS, the page and extraction API use the same address; no local installation is needed for this single-page preview.

Expected output

Each extraction appears as a card under Recent runs with its status and total client-visible time; select a card to see its title, final URL, and the chosen output. Runs are kept only for the current visit and are never saved.

Choose a format

Format, at the left under the URL, chooses what the result panel shows. Every format comes from the same extraction, so switching (here or with View in the result panel) never spends another preview.

Format Shows Download
Markdown The page’s readable content. .md
Links The page’s links, deduplicated, up to 500, with the total found. .links.txt
Page info Title, description, language, canonical URL, robots, keywords and icon; for a file, its type, size, pages and warnings. .info.json
Fields The fields set in Options: each value, or why it is empty, and where on the page it was read. .fields.json
JSON The whole result as the server sent it, including the requested URL. .json

Links and page metadata come with ordinary pages that were read, and fields with those pages and with files. Amazon.sg products, X and Reddit posts return their checked record instead, and a blocked or failed result has only its reason and JSON.

Options

Options holds two settings. The button shows the number of fields, or a dot when only Read is changed.

  • Read: Main content (the default) leaves out headers, menus and footers. Whole page keeps them; a page with no clear main content then returns everything instead of failing.
  • Fields: up to 20 fields, each a name and a type (text, number, yes / no, or a list of text). Type a name on the add line and press Enter to add it, and set its type beside it. + product fields fills in name, brand, price, currency, availability, SKU, rating and review count. Names use letters, digits, spaces, dots, dashes and underscores.

The Fields view lists each field in the order you set them, marked ✓ when the page states it, · when it does not, and : when the page states more than one value. Under each value it shows where on the page it was read, as the downloaded .fields.json records it. Links and Page info read the same way: links are numbered and marked · on the page’s own site or ↗ elsewhere, and each page-info value names the tag it was read from, or the tag OctoCrawl looked for when the page does not declare it.

Fields are read from the page itself: JSON-LD, microdata, meta tags, table rows, definition lists, and a PDF’s Label: value lines. No AI model is used. A field the page does not state comes back empty with its reason, never guessed. Amazon.sg product pages take no options; they always return their checked product record.

Run it on your computer

Get code shows the same extraction for the local API (a cURL command for POST /v1/scrape) and for the local MCP service (the scrape tool call, or scrape_product for an Amazon.sg product), with the URL, format and options you chose. Both need a checkout of the OctoCrawl repository and have no daily limit. A local run can also use a local browser, so its result may differ from this preview. See the API reference and Connect MCP.

For the exact recorded success sample and its observation time, see Introduction. The result may differ when the source page changes.

If extraction does not complete

  • Blocked: the site denied automated access, required login, or returned a verification page. Try another permitted public page; OctoCrawl does not solve a challenge.
  • Timed out: the source or local outbound path did not finish within the preview deadline. Check the final URL and reason. A timeout alone does not prove that the site’s parser is wrong.
  • Incomplete: content or identity could not be fully verified. Read what is available and its missing reason; do not treat it as a complete record.
  • Daily limit reached: stop until the applicable quota resets. The local review server’s counters reset on restart; the hosted preview uses durable counters.

See limits and result states before relying on an extracted field in another system.

The page preview runs at this site's URL. MCP setup is local; hosted MCP is paused.Try a page ↗