Back to catalog

Zyte Api · Web

Web Data Extraction API / Process a single URL, return the result

Process a single URL, return the result. This endpoint blocks until the result is ready. It is intended for short-running operations. At least one of the following request fields must be set to true: - browserHtml - httpResponseBody - httpResponseHeaders - screenshot - An automatic extraction request field: - article - articleList - articleNavigation - forumThread - jobPosting - jobPostingNavigation - pageContent - product - productList - productNavigation - serp All automatic extraction data types support performing extraction using either a browser request or an HTTP request. Choose which using extractFrom; for serp use serpOptions.extractFrom instead. When no option is specified, currently automatic extraction defaults to using a browser request, except for serp, where an HTTP request is used by default instead. In the future, however, the default value may depend on the target website. When automatic extraction uses a browser request, it can be combined with any fields compatible with browserHtml, e.g. screenshot. When automatic extraction uses an HTTP request, it can be combined with any fields compatible with httpResponseBody. serp cannot be combined with any other fields besides serpOptions and url. You cannot combine multiple automatic extraction request fields (e.g. product and productList) on the same request. You cannot combine httpResponseBody with a request field that is exclusive of browser requests (e.g. httpResponseBody and browserHtml). httpResponseHeaders can be requested alone or with any other valid combination of request fields except for serp. The request body size limit is 5MiB.

Private Gateway connection

Data below comes from the configured private Gateway. Provider activation remains governed by its evidence and policy gates.

unverifiedRequest shape unavailableUnverified
Published API evidence
Public metadata only. Response bodies, credentials, and internal review notes are never displayed.

No verification date is claimed. No response capture is claimed.

Request parameters

  • actions body · optional

    Sequence of browser actions to execute. Select an action below to see its API reference. When using actions, you get the actions response field with debug information about action execution. [See an example](/zyte-api/usage/browser.md).

  • article body · optional

    Set to `true` to get article data in the article response field. The target page should only contain a single article, such as a blog post or a news article. For pages with multiple articles consider using articleList instead. To combine this field with [HTTP requests](/zyte-api/usage/http.md), set extractFrom to `"httpResponseBody"`. If you use actions, data extraction happens *after* action execution has finished or timed out. See also: [List of all automatic extraction request fields](/zyt...

  • articleList body · optional

    Set to `true` to get article list data in the articleList response field. The target page should contain multiple articles, usually as links or short snippets. Examples of such pages are main or category pages of news sites, main pages of blogs showing multiple posts, and other pages with multiple articles. Article list data is especially useful to get basic information about articles on a website, like a headline and a link to the article details, using a smaller number of requests, when art...

  • articleListOptions body · optional

    Options for automatic extraction.

  • articleNavigation body · optional

    Set to `true` to get article navigation data in the articleNavigation response field. The target page should contain multiple articles and/or subcategories that can be followed. Article navigation data is especially useful for implementing article crawling, i.e. following links to article pages, as well as to subcategories and pagination that can in turn link to more article pages. Article navigation data can also be used to get basic information of articles and subcategories on a website, ob...

  • articleNavigationOptions body · optional

    Options for automatic extraction.

  • articleOptions body · optional

    Options for automatic extraction.

  • browserHtml body · optional

    Set to `true` to get the [browser HTML](/zyte-api/usage/browser.md) in the browserHtml response field. This field is not compatible with [HTTP requests](/zyte-api/usage/http.md). If you use actions, the browser HTML is generated *after* action execution has finished or timed out. By default, [iframes](https://developer.mozilla.org/en-US/docs/Web/HTML/Element/iframe) are empty. See includeIframes. To access content from the [shadow DOM](https://developer.mozilla.org/en-US/docs/Web/Web_Componen...

  • cookieManagement body · optional

    Cookie management method It determines how to handle user cookies, defined through requestCookies, and automatic cookies, cookies automatically generated by Zyte API. `auto` (default) uses user cookies if defined, or automatic cookies otherwise. `discard` uses user cookies if defined, or no cookies otherwise.

  • customAttributes body · optional

    Schema of the custom attributes to extract. This is a subset of the OpenAPI specification, using JSON syntax. Zyte custom attributes extraction uses a Large Language Model (LLM) operated by Zyte to obtain any structured data specified by this schema from any unstructured web page. This allows to perform extraction similar to standard schemas, such as article or product, but much more flexibly. When this field is specified, the customAttributes.values field in the response would contain the ex...

  • customAttributesOptions body · optional

    Additional options for custom attributes extraction.

  • customHttpRequestHeaders body · optional

    HTTP request headers. Can only be used in combination with httpResponseBody. To set headers with other outputs, see requestHeaders. Setting HTTP request headers has some caveats: - Zyte API sends some headers automatically for [ban avoidance](/zyte-api/usage/errors.md), and may silently override or drop some of your custom headers for that purpose. However, your custom headers may override those automatic headers, and in doing so they can break the ban avoidance capabilities of Zyte API, as s...

  • device body · optional

    Type of device to emulate during your request. A desktop device is emulated by default. Can only be used in combination with httpResponseBody.

  • echoData body · optional

    This field is returned in the echoData response field, verbatim. This field can be useful, for example, to keep track of the original request order when [sending multiple requests in parallel](/zyte-api/usage/optimize.md). The request can be rejected if the data is too big. [See an example](/zyte-api/usage/features.md). See also: jobId.

  • extractFrom body · optional

    [Extraction source](/zyte-api/usage/extract/index.md). `httpResponseBody` extracts from httpResponseBody. It is usually faster and cheaper. `browserHtml` extracts from both browserHtml and screenshot. It typically improves quality over `httpResponseBody`, but is not as robust in case of rendering issues. `userHtml` extracts from user-provided HTML content. When this option is set, the userHtml field must contain the HTML to extract from. No download request is performed and no browser renderi...

  • followRedirect body · optional

    Whether to follow [HTTP redirection](https://developer.mozilla.org/en-US/docs/Web/HTTP/Redirections) or not. Only supported in [HTTP requests](/zyte-api/usage/http.md), [browser requests always follow redirection](/zyte-api/usage/browser.md).

  • forumThread body · optional

    Set to `true` to get forum threads data in the forumThread response field. The target page should contain an individual forum thread page on a forum website. To combine this field with [HTTP requests](/zyte-api/usage/http.md), set extractFrom to `"httpResponseBody"`. If you use actions, data extraction happens *after* action execution has finished or timed out. See also: [List of all automatic extraction request fields](/zyte-api/usage/extract.md), article, browserHtml, screenshot, requestHea...

  • forumThreadOptions body · optional

    Options for automatic extraction.

  • geolocation body · optional

    [ISO 3166-1 alpha-2](https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2) code of a country from which the request should be sent, i.e. the request [geolocation](/zyte-api/usage/features.md). If not specified, Zyte API will use a geolocation that, for the target website, does not cause bans or unexpected locale changes in the response data, such as the wrong language, currency, date format, time zone, etc. If you believe Zyte API is using the wrong default geolocation for a website... Example: "US"

  • httpRequestBody body · optional

    [Base64](https://en.wikipedia.org/wiki/Base64)-encoded data to send as request body. Can only be used in combination with httpResponseBody. It usually needs to be used in combination with httpRequestMethod. If you only need to send UTF-8-encoded text, use httpRequestText instead to skip Base64-encoding. Note that you cannot combine both fields on the same request. [See an example](/zyte-api/usage/http.md). See also: customHttpRequestHeaders.

  • httpRequestMethod body · optional

    Request [HTTP method](https://developer.mozilla.org/en-US/docs/Web/HTTP/Methods). Can only be used in combination with httpResponseBody. [See an example](/zyte-api/usage/http.md). See also: httpRequestText, httpRequestBody, customHttpRequestHeaders, httpResponseHeaders.

  • httpRequestText body · optional

    UTF-8 text to send as request body. Can only be used in combination with httpResponseBody. It usually needs to be used in combination with httpRequestMethod. If you need to send a binary or non-UTF-8 request body, use httpRequestBody instead. Note that you cannot combine both fields on the same request. [See an example](/zyte-api/usage/http.md). See also: customHttpRequestHeaders.

  • httpResponseBody body · required

    Set to `true` to get the HTTP response body in the httpResponseBody response field. This field is not compatible with [browser automation](/zyte-api/usage/browser.md). [See an example](/zyte-api/usage/http.md). See also: httpRequestMethod, httpRequestText, httpRequestBody, customHttpRequestHeaders.

  • httpResponseHeaders body · optional

    Set to `true` to get the HTTP response headers in the httpResponseHeaders response field. [See an example](/zyte-api/usage/features.md). See also: customHttpRequestHeaders, requestHeaders.

  • includeIframes body · optional

    Whether to add the content of [iframes](https://developer.mozilla.org/en-US/docs/Web/HTML/Element/iframe) into browserHtml. Note that iframes are visible in screenshots even if this is set to `false`. See also: browserHtml.

  • ipType body · optional

    [Type of IP address](/zyte-api/usage/features.md) from which the request should be sent. If not specified, Zyte API will use an IP type that, for the target website, does not cause bans or unexpected response data. If you believe Zyte API is using the wrong default IP type for a website, please [reach out to our expert anti-ban team](https://support.zyte.com/support/tickets/new). [See an example](/zyte-api/usage/features.md).

  • javascript body · optional

    Forces JavaScript execution on a [browser request](/zyte-api/usage/browser.md) to be enabled (`true`) or disabled (`false`). By default Zyte API enables or disables JavaScript execution for a request depending on which option makes it easier to avoid bans. Use this request field to override that choice. Passing this request field when requesting automatic extraction ( product, article, etc.) may impact the quality of the returned data, as it might override the optimal value for automatic extr...

  • jobId body · optional

    ID of the [Scrapy Cloud](/scrapy-cloud/get-started.md) job from which this request has been sent, to be returned in the jobId response field. This field is meant to help with request tracking. [scrapy-zyte-api](https://scrapy-zyte-api.readthedocs.io/en/latest/index.html) fills this request field automatically. [See an example](/zyte-api/usage/features.md). See also: echoData. Example: "example-job-1"

  • jobPosting body · optional

    Set to `true` to get job posting data in the jobPosting response field. The target page should contain individual job posting page on a company website or on a job website. To combine this field with [HTTP requests](/zyte-api/usage/http.md), set extractFrom to `"httpResponseBody"`. If you use actions, data extraction happens *after* action execution has finished or timed out. See also: [List of all automatic extraction request fields](/zyte-api/usage/extract.md), browserHtml, screenshot, requ...

  • jobPostingNavigation body · optional

    Set to `true` to get job posting navigation data in the jobPostingNavigation response field. The target page should contain multiple job postings and/or subcategories that can be followed. Job posting navigation data is especially useful for implementing job posting crawling, i.e. following links to job posting pages, as well as pagination that can in turn link to more job posting pages. Job posting navigation data can also be used to get basic information of job postings on a website, obtain...

  • jobPostingNavigationOptions body · optional

    Options for automatic extraction.

  • jobPostingOptions body · optional

    Options for automatic extraction.

  • networkCapture body · optional

    Filters to capture browser network responses. HTTP responses received during browser rendering (including action execution) will be returned in the networkCapture response field if they match any of the filters defined here. You can capture up to 10 responses, provided the sum of their bodies does not exceed 5 MiB. If they do exceed that limit, only the first captured responses within the limit are returned. [See an example](/zyte-api/usage/browser.md).

  • pageContent body · optional

    Set to `true` to get page content data in the pageContent response field. The target page can contain any type of data. Page content data is especially useful for understanding the layout and hierarchy of information on a page, enabling advanced processing such as content extraction, user experience analysis, and automated page summarization. Page content data can also be used to capture the main content intended for users, along with auxiliary navigation components such as headers, footers,...

  • pageContentOptions body · optional

    Options for automatic extraction.

  • product body · optional

    Set to `true` to get product data in the product response field. The target page should only contain a single product. For pages with multiple products consider using productList instead. To combine this field with [HTTP requests](/zyte-api/usage/http.md), set extractFrom to `"httpResponseBody"`. If you use actions, data extraction happens *after* action execution has finished or timed out. [See an example](/zyte-api/usage/extract.md). See also: [List of all automatic extraction request field...

  • productList body · optional

    Set to `true` to get product list data in the productList response field. The target page should contain a list or a grid of products. Product list data is especially useful to get basic information about products on a website using a smaller number of requests, when product attributes are extracted directly from a product list page, without making individual product requests. To implement product crawling from product list pages, use productNavigation, which also enables navigation through p...

  • productListOptions body · optional

    Options for automatic extraction.

  • productNavigation body · optional

    Set to `true` to get product navigation data in the productNavigation response field. The target page should contain multiple products and/or subcategories that can be followed. Product navigation data is especially useful for implementing product crawling, i.e. following links to product pages, as well as to subcategories and pagination that can in turn link to more product pages. Product navigation data can also be used to get basic information of products and subcategories on a website, ob...

  • productNavigationOptions body · optional

    Options for automatic extraction.

  • productOptions body · optional

    Additional options for product extraction.

  • requestCookies body · optional

    A list of cookies to be sent with a request. You can use the contents of the responseCookies response field as a value for this request field. [See an example](/zyte-api/usage/features.md).

  • requestHeaders body · optional

    HTTP request headers. Can only be used in a [browser request](/zyte-api/usage/browser.md). For [HTTP requests](/zyte-api/usage/http.md), see customHttpRequestHeaders. At the moment it only supports the [Referer header](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Referer). [See an example](/zyte-api/usage/browser.md).

  • responseCookies body · optional

    Set to `true` to get the list of cookies set during a request in the responseCookies response field. [See an example](/zyte-api/usage/features.md). See also: requestCookies.

  • screenshot body · optional

    Set to `true` to get a page screenshot in the screenshot response field. This field is not compatible with [HTTP requests](/zyte-api/usage/http.md). To adjust the screenshot contents you can use screenshotOptions and viewport. If you use actions, the screenshot is generated *after* action execution has finished or timed out. [See an example](/zyte-api/usage/browser.md). See also: browserHtml, requestHeaders.

  • screenshotOptions body · optional

    Options for the screenshot taken when the screenshot request field is `true`.

  • serp body · optional

    Set to `true` to get the data of a search engine results page (SERP) in the serp response field. The target URL should be a search engine URL. Currently, you cannot combine this field with any other request fields besides serpOptions and url. See also: [List of all automatic extraction request fields](/zyte-api/usage/extract.md).

  • serpOptions body · optional

    Options for SERP extraction.

  • session body · optional

    Parameters to create or reuse a [client-managed session](/zyte-api/usage/features.md). If `id` does not match one of your running sessions, a new session is created with that session ID. Otherwise, the matching running session is reused. Client-managed sessions may expire due to any of the following: - 15 minutes (900 seconds) have passed since the session was created. - 2 minutes (120 seconds) have passed since the session use. - For 3 times in a row, requests using this session got banned....

  • sessionContext body · optional

    User-defined name-value pairs to [request a server-managed session](/zyte-api/usage/features.md) initialized with sessionContextParameters). For every subsequent request with the same session context, Zyte API will either reuse an available session created for the same session context or create a new session using sessionContextParameters). Server-managed sessions expire after 4 hours or 3 ban responses. If you are targeting websites that silently expire their sessions before the 4-hour mark,...

  • sessionContextParameters body · optional

    Parameters to create a server-managed session for a given sessionContext). [See an example](/zyte-api/usage/features.md). See also: actions.

  • tags body · optional

    Assign arbitrary key-value pairs to the request that you can use for filtering in the [Stats API](/zyte-api/usage/stats.md). Keys must be strings. Values must be strings or `null`. For example: `{"tags": {"foo": "bar", "baz": null}}`.

  • url body · required

    An absolute URL to extract data from. The host name must be a domain name, it cannot be an IP address. Example: "https://example.com/item-page"

  • verifyCertificate body · optional

    Whether to validate TLS certificates of the target website. When omitted or false, certificates are not validated (matching historical behavior); set to `true` to fail with an error response instead of returning page content when validation fails. Only supported in [HTTP requests](/zyte-api/usage/http.md), in [browser requests](/zyte-api/usage/browser.md) certificates are always validated and this parameter is silently ignored.

  • viewport body · optional

    [Browser viewport](https://developer.mozilla.org/en-US/docs/Glossary/Viewport).

Sign in to run this operation, inspect live eligibility, and see governed execution and audit evidence. Sign in.

Web Data Extraction API / Process a single URL, return the result · looot