Skip to main content
Use POST /html to fetch the fully rendered DOM of any page without spinning up or maintaining a browser yourself. AdsCrawl runs a real Chromium instance behind residential proxies, waits for the page to finish loading, and returns the content in whichever format your workflow needs.

Choose a content mode

The contentMode field controls what you get back: markdown and json modes both use the Mozilla Readability algorithm. If the page has no recognisable article content, the API returns 422 READABILITY_CONTENT_NOT_FOUND.

Send your first request

1

Choose your contentMode

Decide whether you need the full page DOM (html), clean readable text (markdown), or a structured article object (json). For most LLM and summarisation tasks, markdown is the right starting point.
2

Set waitUntil

Use domcontentloaded for speed — it returns as soon as HTML is parsed without waiting for images and stylesheets to load. Switch to load when you need styles to be applied, or networkidle when the page makes follow-up XHR requests to fill in content.
3

Optionally scope with selector

Set selector to a CSS selector to wait for that element and return only its HTML (in html mode) or run Readability against it (in markdown/json mode). A selector that doesn’t match returns 422.
4

Send the request

Include your API key in the x-api-key header and POST a JSON body to https://api.adscrawl.net/html.

JSON response (contentMode=json)

When you use contentMode: "json", the response body is a structured Readability article object:

Tips and common options

Set countryCode to a two-letter region code (e.g. "US", "GB") or "GLOBAL" to route your request through a residential proxy in that region. This lets you bypass geo-restrictions and retrieve the localised version of a page.
Pass a cookies array to fetch pages that require an authenticated session. Each cookie needs at minimum name, value, and domain.
If you request contentMode: "markdown" or contentMode: "json" on a page that contains no recognisable article body — such as a login screen or a data-heavy dashboard — the API returns 422 READABILITY_CONTENT_NOT_FOUND. Switch to contentMode: "html" to retrieve the raw DOM for those pages.

Full request body reference

string
required
Target page URL. Only ports 80 and 443 are supported.
"html" | "markdown" | "json"
Output format. Defaults to "html". "markdown" and "json" apply Readability to extract the article body.
string
CSS selector. Waits for the first matching element; html mode returns only that element’s HTML. markdown/json modes run Readability against it. Missing selectors return 422.
"domcontentloaded" | "load" | "networkidle"
Navigation wait condition. Defaults to "load". Use "domcontentloaded" for faster HTML extraction.
string
Managed proxy region. "GLOBAL" picks a dynamic exit node from 15 popular regions. A two-letter code (e.g. "US") prefers a trusted proxy with dynamic fallback. Cannot be combined with a custom proxy.
cookies[]
Cookie list injected into the browser context before navigation.
string
Browser locale, such as "en-US" or "zh-CN".
string
IANA timezone ID, such as "Asia/Shanghai".
object
Viewport size, e.g. { "width": 1280, "height": 720 }.
number
Navigation timeout in milliseconds. Must be positive and no greater than 3,600,000.