Back to catalog

Tavily · Search

Tavily Search and Extract API / Initiate a web crawl from a base URL

Tavily Crawl is a graph-based website traversal tool that can explore hundreds of paths in parallel with built-in extraction and intelligent discovery.

Private Gateway connection

Data below comes from the configured private Gateway. Provider activation remains governed by its evidence and policy gates.

unverifiedRequest shape unavailableUnverified
Published API evidence
Public metadata only. Response bodies, credentials, and internal review notes are never displayed.

No verification date is claimed. No response capture is claimed.

Request parameters

  • allow_external body · optional

    Whether to include external domain links in the final results list.

  • chunks_per_source body · optional

    Chunks are short content snippets (maximum 500 characters each) pulled directly from the source. Use `chunks_per_source` to define the maximum number of relevant chunks returned per source and to control the `raw_content` length. Chunks will appear in the `raw_content` field as: `<chunk 1> [...] <chunk 2> [...] <chunk 3>`. Available only when `instructions` are provided. Must be between 1 and 5.

  • exclude_domains body · optional

    Regex patterns to exclude specific domains or subdomains from crawling (e.g., `^private\.example\.com$`).

  • exclude_paths body · optional

    Regex patterns to exclude URLs with specific path patterns (e.g., `/private/.*`, `/admin/.*`).

  • extract_depth body · optional

    Advanced extraction retrieves more data, including tables and embedded content, with higher success but may increase latency. `basic` extraction costs 1 credit per 5 successful extractions, while `advanced` extraction costs 2 credits per 5 successful extractions.

  • format body · optional

    The format of the extracted web page content. `markdown` returns content in markdown format. `text` returns plain text and may increase latency.

  • include_favicon body · optional

    Whether to include the favicon URL for each result.

  • include_images body · optional

    Whether to include images in the crawl results.

  • include_usage body · optional

    Whether to include credit usage information in the response. `NOTE:`The value may be 0 if the total use of /extract and /map have not yet reached minimum requirements. See our [Credits & Pricing documentation](https://docs.tavily.com/documentation/api-credits) for details.

  • instructions body · optional

    Natural language instructions for the crawler. When specified, the mapping cost increases to 2 API credits per 10 successful pages instead of 1 API credit per 10 pages. Example: "Find all pages about the Python SDK"

  • limit body · optional

    Total number of links the crawler will process before stopping.

  • max_breadth body · optional

    Max number of links to follow per level of the tree (i.e., per page).

  • max_depth body · optional

    Max depth of the crawl. Defines how far from the base URL the crawler can explore.

  • select_domains body · optional

    Regex patterns to select crawling to specific domains or subdomains (e.g., `^docs\.example\.com$`).

  • select_paths body · optional

    Regex patterns to select only URLs with specific path patterns (e.g., `/docs/.*`, `/api/v1.*`).

  • timeout body · optional

    Maximum time in seconds to wait for the crawl operation before timing out. Must be between 10 and 150 seconds.

  • url body · required

    The root URL to begin the crawl. Example: "docs.tavily.com"

Sign in to run this operation, inspect live eligibility, and see governed execution and audit evidence. Sign in.