Back to catalog

Context Dev · Company

Context API / Extract Structured Website Data

Crawl a website, use the provided JSON Schema and instructions to prioritize relevant internal links, and extract structured data from the selected pages.

Private Gateway connection

Data below comes from the configured private Gateway. Provider activation remains governed by its evidence and policy gates.

unverifiedRequest shape unavailableUnverified
Published API evidence
Public metadata only. Response bodies, credentials, and internal review notes are never displayed.

No verification date is claimed. No response capture is claimed.

Request parameters

  • actions body · optional

    Optional browser actions executed in order on the requested page after it loads, before links are discovered or additional pages are crawled. Requires a paid plan. When actions are provided and stopAfterMs is omitted, the crawl budget defaults to 110000 ms.

  • factCheck body · optional

    When true, every returned value must be grounded in facts stated on the page; fields that cannot be supported by the page are returned as null/empty. When false (default), the model may make reasonable inferences and derivations from the page content (e.g. ideal customer, competitor analysis, recommendations) while keeping verifiable specifics (names, quotes, URLs, dates, metrics) faithful to the source.

  • followSubdomains body · optional

    When true, follow links on subdomains of the starting URL's domain.

  • includeFrames body · optional

    When true, iframe contents are included in Markdown before extraction.

  • instructions body · optional

    Optional extraction guidance, such as which facts to prioritize or how to interpret fields in the schema.

  • maxAgeMs body · optional

    Return cached scrape results if a prior scrape for the same parameters is younger than this many milliseconds. Defaults to 7 days (604800000 ms).

  • maxDepth body · optional

    Optional maximum link depth from the starting URL (0 = only the starting page). If omitted, there is no crawl depth limit.

  • maxPages body · optional

    Maximum number of pages to analyze for extraction. Hard cap: 50. Defaults to 5.

  • pdf body · optional
  • schema body · required

    JSON Schema for the returned data object. Image fields such as `image_urls` or `product_photos` automatically make page image references available to extraction, so product data and photos can be returned in one call. TypeScript Zod users can pass a JSON Schema generated from a Zod object; Python users can pass the equivalent JSON Schema object.

  • settleAnimations body · optional

    When true, waits briefly for CSS and transition animations to settle before extracting each crawled page. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages.

  • stopAfterMs body · optional

    Soft time budget for the crawl in milliseconds. Min: 10000 (10s). Max: 110000 (110s). Defaults to 80000 (80s), or 110000 (110s) when browser actions are provided.

  • tags body · optional

    Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.

  • timeoutMS body · optional

    Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).

  • url body · required

    The starting website URL to crawl and extract from. Must include http:// or https://.

  • waitForMs body · optional

    Optional browser wait time in milliseconds after initial page load for each crawled page.

Sign in to run this operation, inspect live eligibility, and see governed execution and audit evidence. Sign in.

Context API / Extract Structured Website Data · looot