Back to catalog

Context Dev · Company

Context API / Crawl Sitemap

Crawl an entire website's sitemap and return all discovered page URLs. Pass search to have the crawled sitemap filtered down to the pages about a phrase (for example pricing and plans or api authentication docs), most relevant first — a searched crawl scans the whole sitemap and costs 2 credits instead of 1.

Private Gateway connection

Data below comes from the configured private Gateway. Provider activation remains governed by its evidence and policy gates.

unverifiedRequest shape unavailableUnverified
Published API evidence
Public metadata only. Response bodies, credentials, and internal review notes are never displayed.

No verification date is claimed. No response capture is claimed.

Request parameters

  • domain query · required

    Domain to build a sitemap for. Example: "example.com"

  • headers query · optional

    Optional outbound HTTP headers forwarded only to the target URL, sent as deep-object query params such as headers[X-Custom]=value. When provided, caching is bypassed: the result is neither read from nor written to cache.

  • maxLinks query · optional

    Maximum number of links to return from the sitemap crawl. Defaults to 10,000. Minimum is 1, maximum is 100,000.

  • search query · optional

    Optional search phrase. When provided, the crawled sitemap is filtered to the pages whose URLs are about that phrase, most relevant first, and the request costs 2 credits instead of 1. Example: "help center and troubleshooting articles"

  • sitemapUrl query · optional

    Optional explicit sitemap URL. When provided, exactly this sitemap is crawled instead of discovering the domain's sitemaps.

  • tags query · optional

    Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.

  • timeoutMS query · optional

    Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).

  • urlRegex query · optional

    Optional RE2-compatible regex pattern. Only URLs matching this pattern are returned and counted against maxLinks. Example: "^https?://[^/]+/blog/"

  • zdr query · optional

    Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact support@context.dev), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true.

Sign in to run this operation, inspect live eligibility, and see governed execution and audit evidence. Sign in.