Back to catalog

Context Dev · Company

Context API / Crawl Website & Scrape Markdown

Crawls from a start URL, extracts each page's content as Markdown, and returns results for all crawled pages. maxPages caps the pages (up to 100) and maxDepth the link depth (up to 5). 1 Context.dev credit per page.

Private Gateway connection

Data below comes from the configured private Gateway. Provider activation remains governed by its evidence and policy gates.

unverifiedRequest shape unavailableUnverified
Published API evidence
Public metadata only. Response bodies, credentials, and internal review notes are never displayed.

No verification date is claimed. No response capture is claimed.

Request parameters

  • country body · optional

    Fetch the target page through a residential proxy in this country (ISO 3166-1 alpha-2). Example: "de"

  • excludeSelectors body · optional

    CSS selectors to remove before each crawled page is converted to Markdown. Applied after includeSelectors. Exclusion takes precedence: an element matching both is removed. Examples: "nav", "footer", ".ad-banner", "[aria-hidden=true]".

  • followSubdomains body · optional

    When true, follow links on subdomains of the starting URL's domain (e.g. docs.example.com when starting from example.com). www and apex are always treated as equivalent.

  • includeFrames body · optional

    When true, the contents of iframes are rendered to Markdown for each crawled page.

  • includeImages body · optional

    Include image references in the Markdown output

  • includeLinks body · optional

    Preserve hyperlinks in the Markdown output

  • includeSelectors body · optional

    CSS selectors. When provided, only matching HTML subtrees (and their descendants) are kept before each crawled page is converted to Markdown. When omitted, the entire document is kept. Examples: "article.main", "#content", "[role=main]".

  • maxAgeMs body · optional

    Return a cached result if a prior scrape for the same parameters exists and is younger than this many milliseconds. Defaults to 1 day (86400000 ms) when omitted. Max is 30 days (2592000000 ms). Set to 0 to always scrape fresh.

  • maxDepth body · optional

    Maximum link depth from the starting URL (0 = only the starting page)

  • maxPages body · optional

    Maximum number of pages to crawl. Hard cap: 500. Example: 10

  • pdf body · optional

    PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range.

  • settleAnimations body · optional

    When true, waits briefly for CSS and transition animations to settle before extracting each crawled page. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages.

  • shortenBase64Images body · optional

    Truncate base64-encoded image data in the Markdown output

  • stopAfterMs body · optional

    Soft time budget for the crawl in milliseconds. After each scrape, the crawler checks the elapsed time and, if exceeded, returns the pages collected so far instead of continuing. Min: 10000 (10s). Max: 110000 (110s). Default: 80000 (80s).

  • tags body · optional

    Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.

  • timeoutMS body · optional

    Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).

  • url body · required

    The starting URL for the crawl (must include http:// or https:// protocol). Example: "https://www.example.com"

  • urlRegex body · optional

    Regex pattern. Only URLs matching this pattern will be followed and scraped. An automatic prefix scope in the form ^<starting URL> follows a redirect of the starting page. Example: "^https?://[^/]+/blog/"

  • useMainContentOnly body · optional

    Extract only the main content, stripping headers, footers, sidebars, and navigation

  • waitForMs body · optional

    Browser wait time in milliseconds after initial page load for each crawled page. Defaults to 3500 (3.5 seconds). Min: 0. Max: 30000 (30 seconds).

  • zdr body · optional

    Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact support@context.dev), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true.

Sign in to run this operation, inspect live eligibility, and see governed execution and audit evidence. Sign in.

Context API / Crawl Website & Scrape Markdown · looot