> ## Documentation Index
> Fetch the complete documentation index at: https://docs.adscrawl.net/llms.txt
> Use this file to discover all available pages before exploring further.

# Ruby SDK for AdsCrawl

> Install and use the AdsCrawl Ruby SDK to fetch rendered content, screenshots, structured data, and manage remote CDP sessions. Ruby 3.1+, standard library only.

The AdsCrawl Ruby SDK works with Ruby 3.1 and later. It uses only the Ruby standard library, supports hash and keyword arguments that mirror the HTTP API's camelCase field names, and provides explicit HTTP deadlines with credential-redacted errors.

<Note>
  Keep your API key in server-side code. Never expose it in client-side applications.
</Note>

## Install

```bash theme={null}
gem install adscrawl
```

Or add it to your bundle:

```bash theme={null}
bundle add adscrawl
```

<Card title="adscrawl-ruby on GitHub" icon="github" href="https://github.com/AdsCrawl/adscrawl-ruby">
  View source, runnable examples, and releases.
</Card>

## Authenticate and make your first request

[Create an API key](https://app.adscrawl.net/register/?utm_source=rubygems\&utm_medium=sdk\&utm_campaign=adscrawl-ruby) in the dashboard and set `ADSCRAWL_API_KEY` in your server environment. You can also pass the key to `AdsCrawl::Client.new(api_key: "...")`.

```ruby theme={null}
require "adscrawl"

client = AdsCrawl::Client.new
markdown = client.markdown(
  url: "https://www.adscrawl.net",
  waitUntil: "domcontentloaded"
)
puts markdown
```

## Rendered content and screenshots

```ruby theme={null}
require "adscrawl"

client = AdsCrawl::Client.new
html = client.html(url: "https://www.adscrawl.net")
article = client.article(url: "https://www.adscrawl.net")
puts [article["title"], article["textContent"]]

png = client.screenshot(
  url: "https://www.adscrawl.net",
  viewport: {width: 1440, height: 900},
  fullPage: true,
  waitUntil: "load"
)
File.binwrite("page.png", png)
```

`html` returns HTML, `markdown` returns Markdown, `article` returns a Hash, and `screenshot` returns binary PNG data. Page options include `viewport`, `locale`, `cookies`, custom `proxy`, managed `countryCode`, `fingerprint`, server-side `timeoutMs`, and navigation `waitUntil`. A custom proxy and managed country cannot be combined.

## Proxy and fingerprint

```ruby theme={null}
require "adscrawl"

client = AdsCrawl::Client.new
server = ENV["ADSCRAWL_PROXY_SERVER"]
username = ENV["ADSCRAWL_PROXY_USERNAME"]
password = ENV["ADSCRAWL_PROXY_PASSWORD"]
raise "Set both proxy username and password, or neither." if username.nil? != password.nil?
raise "Set ADSCRAWL_PROXY_SERVER with proxy credentials." if server.nil? && (username || password)

routing = {countryCode: "GLOBAL"}
if server
  proxy = {server: server}
  proxy.merge!(username: username, password: password) if username && password
  routing = {proxy: proxy}
end

png = client.screenshot(
  **routing,
  url: "https://www.browserscan.net/",
  viewport: {width: 1440, height: 900},
  fullPage: true,
  waitUntil: "networkidle",
  timeoutMs: 60_000,
  userAgentMode: "random",
  userAgentOs: "windows",
  fingerprint: {
    webRtc: "forward", webGl: "random", webGpu: "random",
    webGlImage: "random", canvas: "random", audioContext: "random",
    clientRects: "random", speechVoices: "random", fonts: "random",
    hardware: "random", doNotTrack: "random"
  },
  timeout_ms: 75_000
)
File.binwrite("browserscan.png", png)
```

AdsCrawl uses real browsers with configurable routing and browser fingerprints. Randomized settings are generated as a coherent profile across the operating system, GPU, hardware, fonts, and related signals. This browser workflow has been verified to access and render BrowserScan, Pixelscan, and IPhey and return screenshots.

## Structured extraction

```ruby theme={null}
require "adscrawl"

client = AdsCrawl::Client.new
puts client.spa.templates["templates"]

result = client.spa.extract(
  url: "https://www.adscrawl.net",
  fields: {
    title: {source: "dom", selector: "h1", value: "text", required: true}
  }
)
puts result.dig("data", "title")

inspection = client.spa.inspect(url: "https://www.adscrawl.net")
puts inspection["candidates"]
```

## Remote CDP browsers

```ruby theme={null}
require "adscrawl"

client = AdsCrawl::Client.new
session = client.cdp.create(
  idleTimeoutMs: 600_000,
  maxSessionMs: 3_600_000,
  browserSettings: {viewport: {width: 1440, height: 900}}
)
begin
  version = client.cdp.get_version(session)
  puts version["Browser"]
  # Connect a CDP-compatible Ruby library to session["cdpBaseUrl"].
ensure
  client.cdp.close(session["sessionId"])
end
```

`client.cdp.list` returns active sessions. `get_version` validates the session token URL and deliberately omits the API key. CDP connection URLs contain secrets; do not log them.

## Persistent cloud browsers

```ruby theme={null}
require "adscrawl"

client = AdsCrawl::Client.new
launched = client.cloud_browsers.launch(
  proxy: {server: ENV.fetch("ADSCRAWL_PROXY_SERVER")},
  tabs: ["https://www.adscrawl.net"]
)
browser_id = launched["id"]
begin
  profile = client.cloud_browsers.get(browser_id)
  puts profile.dig("runtime", "status")
ensure
  client.cloud_browsers.stop(browser_id)
end
```

API-key starts and launches require an explicit top-level custom proxy. A `stopping` response means shutdown is pending; poll `get` until `stopped`. A failed launch may expose a cleanup profile ID as `APIError#resource_id`.

## Configuration and errors

```ruby theme={null}
require "adscrawl"

client = AdsCrawl::Client.new(base_url: "https://api.adscrawl.net", timeout_ms: 90_000)
begin
  puts client.markdown(
    {url: "https://www.adscrawl.net", timeoutMs: 60_000},
    timeout_ms: 75_000
  )
rescue AdsCrawl::APIError => error
  puts [error.status, error.api_code, error.trace_id]
rescue AdsCrawl::TimeoutError
  puts "Inspect remote sessions before retrying."
end
```

The API key defaults to `ADSCRAWL_API_KEY`. The base URL defaults to `ADSCRAWL_BASE_URL`, then `ADSCRAWL_API_URL`, then `https://api.adscrawl.net`. Errors include `APIError`, `TimeoutError`, `ConnectionError`, and `ResponseError`. API response bodies are credential-redacted. The HTTP deadline includes response body reading. Requests are not automatically retried.
