Back to catalog

Tinyfish · Web Data

TinyFish Fetch API / Fetch and extract content from URLs

Fetches web pages, renders JavaScript-heavy pages when needed, and returns clean extracted content in your preferred format. Submit up to 10 URLs, get back structured content. Per-URL failures appear in errors[] and do not fail the entire request. Per-URL error codes (in errors[].error): - targethttperror — target server returned a non-2xx HTTP status other than 404/410; the raw status code is in errors[].status - pagenotfound — target URL returned HTTP 404 or 410; the raw status code is in errors[].status - targetunreachable — connection refused, TLS failure, DNS failure, or other network error - timeout — request timed out - proxyerror — proxy tunnel failure - botblocked — bot-challenge page detected (Cloudflare, etc.) - emptycontent — page loaded but no extractable text was found - loginrequired — the target redirected to a login or account wall; retrying without credentials will not help - contenttoolarge — the document exceeds the 20MB size limit for generic documents (PDF and CSV have their own 50MB budgets); not retried in the browser - invalidurl — malformed URL or SSRF-blocked address - invalidredirecturl — redirect target rejected before fetch - conditionalunsupported — conditional requests (ifnonematch / ifmodifiedsince) are supported on the fast path only; this URL requires browser rendering - selectornotmatched — no elements matching any includeselectors entry remained after excludeselectors was applied; the error carries unmatchedselectors plus candidateselectors retry hints (a partial miss is not an error — it's reported on the result's unmatchedselectors) - selectorunsupported — includeselectors / exclude_selectors sent for a URL that resolves to a direct PDF/CSV download (no HTML to scope)

Private Gateway connection

Data below comes from the configured private Gateway. Provider activation remains governed by its evidence and policy gates.

unverifiedRequest shape unavailableUnverified
Published API evidence
Public metadata only. Response bodies, credentials, and internal review notes are never displayed.

No verification date is claimed. No response capture is claimed.

Request parameters

  • exclude_selectors body · optional

    Array of CSS selectors (1-20 entries, each 1-1000 characters) for elements to remove before extraction — applied before `include_selectors` scopes what remains, so it also prunes inside selected regions. Entries may themselves use CSS comma-grouping. Entries that match nothing are a no-op, never an error, but URLs that resolve to direct PDF/CSV downloads fail with `selector_unsupported`. Invalid CSS selector syntax is rejected with a 422. Applied post-fetch: caching and routing are unchanged.

  • format body · optional

    Output format for extracted content. "markdown" (default) is ideal for LLM consumption. "html" returns cleaned semantic HTML. "json" returns a structured document tree. Example: "markdown"

  • if_modified_since body · optional

    Last-Modified validator from a prior fetch of this URL, forwarded verbatim as the If-Modified-Since header on the origin request. Only valid with a single URL — combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them. Example: "Wed, 21 Oct 2015 07:28:00 GMT"

  • if_none_match body · optional

    ETag validator from a prior fetch of this URL, forwarded verbatim as the If-None-Match header on the origin request. Only valid with a single URL — combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them. Example: "W/\"abc123\""

  • image_links body · optional

    Extract all image URLs (<img src>) from each page. Useful for finding visual content or media assets. Image links are returned as absolute URLs in the image_links array of each result. Example: false

  • include_etag_and_last_modified body · optional

    Opt-in to receiving `etag` / `last_modified` validators (and `not_modified` detection) on each result. Defaults to false — tf-fetch omits these fields unless requested. Independent of `if_none_match` / `if_modified_since`: works with a single URL or a batch. Example: true

  • include_selectors body · optional

    Array of CSS selectors (1-20 entries, each 1-1000 characters) that scope extracted content (`text`, `links`, `image_links`) to elements matching ANY entry, concatenated in document order. Tag selectors cover semantic sections (`main`, `article`, `nav`); entries may themselves use CSS comma-grouping. Selected content is returned verbatim in the requested format (scripts/styles stripped) — automatic boilerplate removal is bypassed. Page-level metadata (`title`, `description`, `language`, `autho...

  • links body · optional

    Extract all outbound links (<a href>) from each page. Useful for discovering related pages or navigating to specific content. Links are returned as absolute URLs in the links array of each result. Example: false

  • page_metadata body · optional

    Return page-head metadata for each page in the page_metadata object of each result: canonical URL, favicon, robots directive, generator, viewport, keywords, all Open Graph (og), Twitter card (twitter), and article tags, and remaining named meta tags under `other`. Useful for SEO/technical audits and link-preview generation. Example: false

  • per_url_timeout_ms body · optional

    Wall-clock timeout budget in milliseconds applied independently to each URL. If one URL exceeds this budget, it returns a per-URL timeout error while other URLs in the same request continue. Example: 45000

  • purpose body · optional

    Why these URLs are being fetched — the underlying goal or task the content will be used for. Used to better tailor fetching and extraction to your intent. Example: "Compare pricing tiers across vendors for a procurement report"

  • ttl body · optional

    Caller freshness tolerance in seconds for the cached entry. Omit (default) for unlimited tolerance — any cached entry is acceptable. Set to 0 to prefer a live fetch; a cached entry is still served if the origin's Cache-Control: max-age covers its age, or the host is in the small allowlist of operator-pinned never-expire domains. Set to N > 0 to accept a cached entry whose age is below N; the upstream Cache-Control: max-age and the never-expire allowlist may extend (never shorten) t... Example: 0

  • urls body · required

    Array of URLs to fetch (1-10). All URLs are fetched in parallel. Each URL is processed independently — if one fails, others still return successfully. Errors are reported per-URL in the errors array.

Sign in to run this operation, inspect live eligibility, and see governed execution and audit evidence. Sign in.

TinyFish Fetch API / Fetch and extract content from URLs · looot