Metadata-Version: 2.4
Name: azcrawlerpy
Version: 1.5.0
Summary: Agentic Crawler Discovery Framework.
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.13
Classifier: Operating System :: OS Independent
Requires-Python: <3.14,>=3.13
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: playwright==1.62.0
Requires-Dist: pydantic>=2.11.10
Requires-Dist: camoufox[geoip]==0.5.6
Requires-Dist: pillow>=10
Provides-Extra: recaptcha
Requires-Dist: onnxruntime>=1.18; extra == "recaptcha"
Requires-Dist: numpy>=1.26; extra == "recaptcha"
Requires-Dist: azcrawlerpy-recaptcha-google==2.0.3; extra == "recaptcha"
Provides-Extra: text
Requires-Dist: onnxruntime>=1.18; extra == "text"
Requires-Dist: numpy>=1.26; extra == "text"
Requires-Dist: azcrawlerpy-recaptcha-text==1.0.1; extra == "text"
Provides-Extra: captcha
Requires-Dist: azcrawlerpy[recaptcha,text]; extra == "captcha"
Dynamic: license-file

# azcrawlerpy

A framework for navigating and filling multi-step web forms programmatically. Supports Camoufox anti-detect browser (built on Firefox with C++ level fingerprint spoofing) and standard Chromium via Playwright. Uses JSON instruction files to define form navigation workflows, making it ideal for automated form submission, web scraping, and AI agent-driven web interactions.

## Table of Contents

- [Installation](#installation)
- [Quick Start](#quick-start)
- [Logging & Telemetry](#logging--telemetry)
- [Core Concepts](#core-concepts)
- [Instructions Schema](#instructions-schema)
  - [Top-Level Structure](#top-level-structure)
  - [Browser Configuration](#browser-configuration)
  - [URL Rewrites](#url-rewrites)
  - [Cookie Consent Handling](#cookie-consent-handling)
  - [Step Definitions](#step-definitions)
  - [Field Types](#field-types)
  - [Optional Fields](#optional-fields)
  - [Action Types](#action-types)
  - [Final Page Configuration](#final-page-configuration)
  - [Post Final Page Steps](#post-final-page-steps)
  - [Per-Step Page Extraction](#per-step-page-extraction)
  - [Data Extraction Configuration](#data-extraction-configuration)
- [Data Points (input_data)](#data-points-input_data)
- [Element Discovery](#element-discovery)
- [AI Agent Guidance](#ai-agent-guidance)
- [Retry and Resilience](#retry-and-resilience)
- [Error Handling and Diagnostics](#error-handling-and-diagnostics)
- [Debugging and Crawl Output](#debugging-and-crawl-output)
  - [Debug Mode](#debug-mode)
  - [Where the Output Lands](#where-the-output-lands)
  - [The Crawl Result](#the-crawl-result)
  - [Screenshots](#screenshots)
  - [Error Diagnostics](#error-diagnostics)
  - [Captured Logs](#captured-logs)
  - [Debug Recording](#debug-recording)
  - [Network Exit](#network-exit)
  - [Azure App Service Log Capture](#azure-app-service-log-capture)
  - [Persisting the Output](#persisting-the-output)
  - [Sensitive Data in the Output](#sensitive-data-in-the-output)
  - [Reading a Failed Crawl](#reading-a-failed-crawl)
- [Browser Profile Building (Profiler)](#browser-profile-building-profiler)
- [Examples](#examples)

## Installation

Requires Python 3.13.

```bash
uv add azcrawlerpy
```

Or install from source:

```bash
uv pip install -e .
```

## Quick Start

`FormCrawler` **requires** a logger. In production, pass an `AzureLogger` from
[`azpaddypy`](../azpaddypy/README.md) so crawl telemetry (spans, correlation
IDs, structured records) flows into Application Insights. For local scripts
and notebooks, any stdlib-shaped logger works — anything that implements
`debug`/`info`/`warning`/`error`/`exception`/`critical`.

```python
import asyncio

from azcrawlerpy import CrawlerBrowserConfig, DebugMode, FormCrawler, HumanizeConfig
from azpaddypy.mgmt.logging import AzureLogger, bootstrap_azure_monitor


async def main():
    # 1. Wire telemetry once per process (idempotent; no-op without a
    #    connection string so it's safe in local runs).
    bootstrap_azure_monitor(service_name="crawler-demo")
    logger = AzureLogger(__name__)

    # 2. CrawlerBrowserConfig controls runtime browser settings
    #    (proxy, stealth, humanize).
    browser_config = CrawlerBrowserConfig(humanize=HumanizeConfig(enabled=True))

    # 3. FormCrawler now requires a logger -- pass AzureLogger for App Insights
    #    integration, or a stdlib `logging.getLogger(__name__)` for local work.
    crawler = FormCrawler(
        headless=True,
        browser_config=browser_config,
        logger=logger,
    )

    instructions = {
        "url": "https://example.com/form",
        "browser_config": {
            "browser_type": "camoufox",
            "viewport_width": 1920,
            "viewport_height": 1080
        },
        "steps": [
            {
                "name": "step_1",
                "wait_for": "input[name='email']",
                "timeout_ms": 15000,
                "fields": [
                    {
                        "type": "text",
                        "selector": "input[name='email']",
                        "data_key": "email"
                    }
                ],
                "next_action": {
                    "type": "click",
                    "selector": "button[type='submit']"
                }
            }
        ],
        "final_page": {
            "wait_for": ".success-message",
            "timeout_ms": 60000
        }
    }

    input_data = {
        "email": "user@example.com"
    }

    result = await crawler.crawl(
        url=instructions["url"],
        input_data=input_data,
        instructions=instructions,
        debug_mode=DebugMode.ALL,
        # Optional wall-clock cap on the entire crawl. Exceeding it raises
        # CrawlerTimeoutError with a partial result attached. Omit or set
        # to None to disable.
        global_timeout_ms=300_000,
    )

    # The crawler returns every artifact in memory -- the caller owns the
    # sink (disk, Azure Blob, stdout, ...). Apart from browser profiles and
    # Playwright's temporary debug file, nothing is written to disk on your behalf.
    print(f"Final URL: {result.final_url}")
    print(f"Steps completed: {result.steps_completed}")
    print(f"Screenshots captured: {len(result.screenshots)}")
    print(f"HTML bytes: {len(result.html)}")
    print(f"Extracted data: {result.extracted_data}")

asyncio.run(main())
```

## Logging & Telemetry

`FormCrawler` accepts any logger that matches the stdlib `logging.Logger`
shape (structural `LoggerLike` protocol in
[`crawling/utils.py`](azcrawlerpy/crawling/utils.py)): `debug`, `info`,
`warning`, `error`, `exception`, `critical`.

### Production: AzureLogger

Pairing the crawler with `AzureLogger` from `azpaddypy` gives you, with no
code in the crawler itself:

- **Application Insights** export of all crawl log records via the OTel
  handler that azpaddypy attaches to the root logger; the crawler's
  module-level loggers propagate to root automatically.
- **Trace context** (`operation_Id` / `operation_ParentId`) is set by the
  Azure Monitor exporter from the active OTel span — no per-record
  duplication.
- **Built-in OpenTelemetry spans** on the crawler's hot paths so the App
  Insights Application Map and Performance views show structure inside the
  crawl rather than one opaque blob:
  - `FormCrawler.crawl` — root span; carries `crawl.url`, `crawl.headless`,
    `crawl.debug_mode`, `crawl.strict`, `crawl.steps_planned`, plus
    post-success `crawl.steps_completed` and `crawl.profiler_visited_count`.
    On timeout / error it carries `outcome=timeout|error` with `Status.ERROR`
    and the captured exception.
  - `FormCrawler.attempt` — one span per restart-loop iteration (`attempt.index`,
    `browser.type`); restart paths add `attempt.outcome=restart` and
    `attempt.restart_step`.
  - `profiler.run` — wraps the profiler phase (`profiler.visit_count`,
    `profiler.in_memory`, `profiler.browser_type`).
  - `ElementDiscovery.discover` — wraps page element discovery
    (`discovery.url`, `discovery.browser_type`, `discovery.explore_iframes`,
    plus post-run `discovery.total_elements` and `discovery.iframes_count`).
- **Correlation IDs** if the caller sets one via
  `AzureLogger.set_correlation_id(...)` — it propagates through the
  `contextvars`-backed OTel context into all crawler spans automatically.

```python
from azpaddypy.mgmt.logging import AzureLogger, bootstrap_azure_monitor, trace_function

bootstrap_azure_monitor(service_name="crawler-api", service_version="0.5.1")
logger = AzureLogger(__name__)

# Optional: generate a correlation ID per request + emit a parent span
@trace_function(name="crawl_request")
async def handle_crawl(url: str, input_data: dict, instructions: dict):
    crawler = FormCrawler(headless=True, browser_config=None, logger=logger)
    return await crawler.crawl(
        url=url, input_data=input_data, instructions=instructions,
    )
```

The crawler does **not** import `opentelemetry` directly. Spans are produced
through a soft-import shim: when OTel is installed (the typical case in a
function-app that pulls in azpaddypy) spans flow to Application Insights;
when not, the shim is a zero-cost no-op so standalone scripts and minimal
test environments keep working unchanged.

### Local / scripts: stdlib logger

If you don't need App Insights, a standard library logger is enough:

```python
import logging

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(name)s: %(message)s")
crawler = FormCrawler(
    headless=True,
    browser_config=None,
    logger=logging.getLogger("my_crawler"),
)
```

Any logger that satisfies the `LoggerLike` Protocol is accepted — there is no
hard runtime dependency on `azpaddypy`.

## Core Concepts

The framework operates on two primary inputs:

1. **Instructions (instructions.json)**: Defines the form structure, selectors, navigation flow, and field types
2. **Data Points (input_data)**: Contains the actual values to fill into form fields

The crawler processes each step sequentially:
1. Wait for the step's `wait_for` selector to become visible
2. Fill all fields defined in the step using values from `input_data`
3. Execute the `next_action` to navigate to the next step
4. Repeat until all steps are complete
5. Wait for and capture the final page

## Instructions Schema

### Top-Level Structure

```json
{
  "url": "https://example.com/form",
  "browser_config": { ... },
  "cookie_consent": { ... },
  "steps": [ ... ],
  "final_page": { ... },
  "post_final_page_steps": [ ... ],
  "data_extraction": { ... }
}
```

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `url` | string | Yes | Starting URL for the form; `crawl(url=...)` must pass the same URL or it raises `ValueError` |
| `browser_config` | object | Yes | Browser engine and viewport settings |
| `cookie_consent` | object | No | Cookie banner handling configuration |
| `captcha` | object | No | CAPTCHA handling configuration (e.g., Cloudflare Turnstile) |
| `steps` | array | Yes | Ordered list of form steps |
| `final_page` | object | Yes | Configuration for the result page |
| `post_final_page_steps` | array | No | Steps executed on the result page after its wait/settle, before `data_extraction` (see [Post Final Page Steps](#post-final-page-steps)) |
| `data_extraction` | object | No | Configuration for extracting data from final page |
| `profiler` | object | No | Browser profile building configuration (visit sites to accumulate cookies before crawling) |

### Browser Configuration

```json
{
  "browser_config": {
    "browser_type": "camoufox",
    "viewport_width": 1920,
    "viewport_height": 1080
  }
}
```

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `browser_type` | string | Yes | Browser engine: `camoufox` (anti-detect Firefox) or `chromium` (standard Playwright) |
| `viewport_width` | integer | Yes | Browser viewport width in pixels |
| `viewport_height` | integer | Yes | Browser viewport height in pixels |
| `blocked_url_patterns` | array | No | URL glob patterns to block via `page.route()` (e.g., `**/analytics/**`) |
| `url_rewrites` | object | No | Substring replacements applied to request URLs before they are sent (see [URL Rewrites](#url-rewrites)) |
| `url_rewrite_mode` | string | No | How rewrites are applied: `redirect` (default) or `mirror` (see [URL Rewrites](#url-rewrites)) |
| `capture_url_substrings` | array | No | Requests whose URL contains one of these substrings are recorded with headers and bodies in `network_log` (see [Debug Recording](#debug-recording)) |
| `allow_synthetic_events` | boolean | No | Allow JS fallbacks a page can tell apart from a user: clicks dispatched from JS (`isTrusted=false`) and removing a cookie banner no accept button dismissed. Default `false` (see [Click Fallback Chain](#click-fallback-chain)) |

Both engines render the configured width exactly. The height differs by engine: chromium sets it as the page's inner height, while camoufox sizes the browser window, leaving the page shorter by the height of the browser chrome. Camoufox ignores a Playwright viewport, so the size goes in through the fingerprint: the generated screen is pinned to the viewport so camoufox cannot shrink the window to a smaller rolled screen, and every camoufox crawl pins a fingerprint preset whose available screen area (the screen less its taskbar) holds the window, and sets the window explicitly.

### URL Rewrites

Redirects matching requests to a different host. The main use case is a proxy that cannot reach a
third-party host the page depends on, but which has a reachable mirror.

```json
{
  "browser_config": {
    "browser_type": "chromium",
    "viewport_width": 1920,
    "viewport_height": 1080,
    "url_rewrites": {
      "://www.google.com/recaptcha/": "://www.recaptcha.net/recaptcha/"
    }
  }
}
```

Each key is matched as a substring of the request URL and replaced with its value. Requests are
routed through `page.route()`, so the page's own loader keeps working — including any `onload`
callback it registered on the original script tag.

`url_rewrite_mode` controls how the replacement is reached:

| Mode | Behaviour | Use when |
|------|-----------|----------|
| `redirect` (default) | Sends the request to the replacement URL via `route.continue_(url=...)`. The document's origin becomes the replacement host. | Plain assets, and anything on chromium |
| `mirror` | Fetches the replacement and serves its content under the **original** URL via `route.fulfill()`. The origin is unchanged, and so are URLs inside text responses (see notes). | The resource performs same-origin checks, or the target is firefox/camoufox |

Prefer `redirect` — it is a single request. Reach for `mirror` when a rewritten resource loads
same-origin sub-resources: firefox refuses to start a web worker from a different origin than its
document, so a redirected script that spawns one fails with a `Security Error` and never finishes
initialising. `mirror` keeps the origin intact and avoids the check entirely.

Notes:

- With `redirect` the replacement must keep the same protocol; cross-origin redirects are allowed.
- Rewrites apply for the whole page lifetime, so a rewritten script's own follow-up requests are
  rewritten too when they match.
- With `mirror`, text responses (HTML, JavaScript, CSS, JSON, XML) get the original back in place of
  the replacement before the page sees them. A mirrored script that names its own host, as
  reCAPTCHA's `api.js` does, would otherwise make the page load from a host its
  Content-Security-Policy may not allow. Those URLs match the rule again, so they are still fetched
  from the replacement. Binary responses, and bodies that do not contain the replacement, pass
  through unchanged.
- Write rules as specific URL fragments such as `://www.google.com/recaptcha/`: a `mirror` rule is
  reversed in response bodies exactly as broadly as it is applied to URLs, so a bare hostname is
  reversed wherever it appears.
- Prefer either mode over an init script that patches URLs in JavaScript. Camoufox runs injected
  scripts in an isolated world that cannot touch page globals, so a JS-based rewrite silently does
  nothing there; routing works on both engines and leaves no page-visible trace.

Example — reCAPTCHA behind a proxy whose zone blocks `google.com`. Google serves reCAPTCHA from
`recaptcha.net` for exactly this situation, so rewriting the host lets the widget load and render
normally. On camoufox this needs `"url_rewrite_mode": "mirror"`, because the reCAPTCHA frame spawns
a worker from a hardcoded `google.com` URL. Note this only restores *reachability*; whether the
challenge then passes still depends on the exit IP's reputation.

### Cookie Consent Handling

The framework supports two modes for handling cookie consent banners:

**Standard Mode** (regular DOM elements):
```json
{
  "cookie_consent": {
    "banner_selector": "dialog:has-text('cookies')",
    "accept_selector": "button:has-text('Accept')"
  }
}
```

**Shadow DOM Mode** (for Usercentrics, OneTrust, etc.):
```json
{
  "cookie_consent": {
    "banner_selector": "#usercentrics-cmp-ui",
    "shadow_host_selector": "#usercentrics-cmp-ui",
    "accept_button_texts": ["Accept All", "Alle akzeptieren"]
  }
}
```

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `banner_selector` | string | Yes | CSS selector for the banner container |
| `accept_selector` | string | No | CSS selector for accept button (standard mode) |
| `shadow_host_selector` | string | No | CSS selector for shadow DOM host |
| `accept_button_texts` | array | No | Text patterns to match accept buttons in shadow DOM (case-insensitive, at the start of a word) |
| `banner_settle_delay_ms` | integer | No | Wait time before checking for banner |
| `banner_visible_timeout_ms` | integer | No | Timeout for banner visibility |
| `accept_button_timeout_ms` | integer | No | Timeout for accept button visibility |
| `post_consent_delay_ms` | integer | No | Wait time after handling consent |
| `js_fallback_texts` | array | No | Custom text patterns for JS fallback button matching (overrides defaults) |

The JS fallback matches button text against: `klar`, `akzept`, `accept`, `agree`, `ok`, `verstanden`, `einverstanden`, `zustimm`. Set `js_fallback_texts` to override this list with site-specific patterns. Both `js_fallback_texts` and `accept_button_texts` match case-insensitively at the start of a word, so `ok` matches "OK" and "Okay" but not "Cookie-Einstellungen", and `akzept` matches "Alle akzeptieren". A text that only occurs inside a word ("verstanden" in "Einverstanden") needs its own entry.

Shadow DOM accept buttons get a real click. The JS fallback, and removing a banner that no accept button dismissed, only run when `browser_config.allow_synthetic_events` is on.

### Step Definitions

Each step represents a form page or section:

```json
{
  "name": "personal_info",
  "wait_for": "input[name='firstName']",
  "timeout_ms": 15000,
  "fields": [ ... ],
  "next_action": { ... }
}
```

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `name` | string | Yes | Unique identifier for the step |
| `wait_for` | string | Yes | CSS selector to wait for before processing |
| `timeout_ms` | integer | Yes | Timeout in milliseconds for wait condition |
| `fields` | array | Yes | List of field definitions |
| `next_action` | object | Yes | Action to navigate to next step |
| `data_extraction` | array | No | Data to extract BEFORE field handling in this step |
| `post_field_extraction` | array | No | Data to extract AFTER field handling (for modal results, dynamic values) |
| `page_extraction` | object | No | Full page extraction run at the end of this step, same schema as `data_extraction` on the instructions (see [Per-Step Page Extraction](#per-step-page-extraction)) |
| `optional` | boolean | No | Skip the step when it fails: its wait_for is not found, its budget runs out or an extraction fails (see [Optional Steps](#optional-steps)) |
| `restart_crawl` | boolean | No | Trigger a full browser restart from the start URL when the step fails: its wait_for times out, its budget runs out or an extraction fails (see [Restart Crawl](#restart-crawl)). Requires `max_restarts`. |
| `max_restarts` | integer | No | Maximum number of full crawl restart attempts when `restart_crawl` is true (must be >= 1). Each step tracks its own restart budget. |

### Restart Crawl

Steps marked with `"restart_crawl": true` raise an internal `CrawlRestartError` when the step fails: its `wait_for` selector times out, its `timeout_ms` budget runs out, or one of its extractions fails (the error message names which). The crawler then closes the browser, recreates the context (reusing any profiler storage state so cookie consent does not re-run), and replays the entire crawl from the start URL. Each step tracks its own restart budget against `max_restarts`; once exhausted, the original timeout error is raised.

```json
{
  "name": "results_page",
  "wait_for": "[data-cy='quote-result']",
  "timeout_ms": 15000,
  "restart_crawl": true,
  "max_restarts": 2,
  "fields": [],
  "next_action": { "type": "delay", "delay_ms": 0 }
}
```

`restart_crawl` takes priority over `optional` and `strict`. Use it for steps where the page may render broken or never finalize due to flaky SPA/backend behavior, where a fresh browser context is the most reliable recovery. Prefer `retry_config` on individual fields/actions for transient interaction failures; reserve `restart_crawl` for full-page load failures.

### Optional Steps

Steps marked with `"optional": true` are skipped gracefully when the `wait_for` selector is not found within `timeout_ms`, the step's budget runs out, or one of its extractions fails. The entire step (fields, data extraction, next_action) is skipped with an info log. This is useful for conditional workflow pages that only appear depending on prior input values or dynamic form behavior.

```json
{
  "name": "additional_driver_details",
  "wait_for": "#additional-driver-form",
  "timeout_ms": 5000,
  "optional": true,
  "fields": [ ... ],
  "next_action": { "type": "click", "selector": "#next" }
}
```

The `optional` check takes precedence over the `strict` parameter. When a step is optional and the wait_for selector times out, the step is always skipped regardless of strict mode. Non-optional steps follow the existing behavior: raise `CrawlerTimeoutError` when `strict=True`, or log a warning and skip the rest of the step when `strict=False`.

**Tip:** Use a shorter `timeout_ms` (e.g., 3000-5000ms) on optional steps to avoid waiting the full timeout when the step is absent.

### Field Types

#### TEXT

For text inputs, email fields, phone numbers, and similar:

```json
{
  "type": "text",
  "selector": "input[name='email']",
  "data_key": "email"
}
```

#### TEXTAREA

For multi-line text areas:

```json
{
  "type": "textarea",
  "selector": "textarea[name='message']",
  "data_key": "message"
}
```

#### DROPDOWN / SELECT

For native `<select>` elements:

```json
{
  "type": "dropdown",
  "selector": "select[name='country']",
  "data_key": "country",
  "type_config": {
    "select_by": "text"
  }
}
```

| type_config Parameter | Values | Description |
|-----------------------|--------|-------------|
| `select_by` | `text`, `value`, `index` | How to match the option |
| `option_visible_timeout_ms` | integer | Timeout in ms for option visibility |

#### RADIO

For radio button groups:

```json
{
  "type": "radio",
  "selector": "input[type='radio'][value='${value}']",
  "data_key": "payment_method"
}
```

**Pattern A - Value-driven selector**: Use `${value}` placeholder in selector, data provides the value:
```json
{
  "type": "radio",
  "selector": "input[type='radio'][value='${value}']",
  "data_key": "gender"
}
// data: { "gender": "male" }
```

**Pattern B - Boolean flags**: Use explicit selectors with boolean data values:
```json
{
  "type": "radio",
  "selector": "[role='radio']:has-text('Yes')",
  "data_key": "accept_terms",
  "force_click": true
}
// data: { "accept_terms": true }  // clicks if truthy, skips if null
```

`force_click` is supported on `radio`, `click_only`, and `click_select` field types. When set, the click fallback chain tries force click before normal click.

#### CHECKBOX

For checkbox inputs:

```json
{
  "type": "checkbox",
  "selector": "input[type='checkbox'][name='newsletter']",
  "data_key": "subscribe_newsletter"
}
```

Data value `true` checks the box, `false` unchecks it, and `null` skips the field entirely (no interaction). Only JSON booleans are accepted: any other value, such as the string `"false"`, raises `FieldInteractionError` (or skips an `optional` field) instead of being read as truthy.

#### DATE

For date inputs with format conversion:

```json
{
  "type": "date",
  "selector": "input[name='birthdate']",
  "data_key": "birthdate",
  "type_config": {
    "format": "DD.MM.YYYY"
  }
}
```

Supported formats (mapped to strftime internally):

| Format | Example | Description |
|--------|---------|-------------|
| `DD.MM.YYYY` | 15.06.1985 | Day.Month.Year |
| `MM/DD/YYYY` | 06/15/1985 | Month/Day/Year |
| `YYYY-MM-DD` | 1985-06-15 | ISO format |
| `DD/MM/YYYY` | 15/06/1985 | Day/Month/Year |
| `YYYY/MM/DD` | 1985/06/15 | Year/Month/Day |
| `DD-MM-YYYY` | 15-06-1985 | Day-Month-Year |
| `MM-DD-YYYY` | 06-15-1985 | Month-Day-Year |

Data must be provided in ISO format (`YYYY-MM-DD`) in `input_data`. The `type_config.format` specifies the output format for typing into the field; it takes the patterns above, not strftime directives. A value that is not ISO 8601 raises `FieldInteractionError` before the page is touched (an `optional` field is skipped instead), and an unsupported `format` raises `InvalidInstructionError`. Native `<input type="date">` fields are auto-detected and use the value as-is, no `type_config` needed.

#### SLIDER

For range inputs:

```json
{
  "type": "slider",
  "selector": "input[type='range'][name='coverage']",
  "data_key": "coverage_amount"
}
```

#### FILE

For file upload fields:

```json
{
  "type": "file",
  "selector": "input[type='file']",
  "data_key": "document_path"
}
```

Data value should be the absolute file path.

#### COMBOBOX

For autocomplete/typeahead inputs:

```json
{
  "type": "combobox",
  "selector": "input[aria-label='City']",
  "data_key": "city",
  "press_enter": true,
  "type_config": {
    "option_selector": ".autocomplete-option",
    "type_delay_ms": 50,
    "wait_after_type_ms": 500
  }
}
```

| type_config Parameter | Description |
|-----------------------|-------------|
| `option_selector` | CSS selector for dropdown options (required) |
| `type_delay_ms` | Delay between keystrokes (simulates human typing) |
| `wait_after_type_ms` | Wait time for options to appear |
| `clear_before_type` | Clear field before typing |
| `option_visible_timeout_ms` | Timeout in ms for option visibility |

> **Breaking change in 1.1.6:** `press_enter` moved from the combobox `type_config` to the field itself, where it now also works for `text`, `date`, `iframe_field` and `click_select`. Move the key up one level; leaving it inside `type_config` fails validation.

#### CLICK_SELECT

For custom dropdowns requiring click-then-select:

```json
{
  "type": "click_select",
  "selector": ".custom-dropdown-trigger",
  "data_key": "option_value",
  "post_click_delay_ms": 300,
  "type_config": {
    "option_selector": ".dropdown-item:has-text('${value}')"
  }
}
```

#### CLICK_ONLY

For elements that only need clicking (no data input):

```json
{
  "type": "click_only",
  "selector": "button.expand-section"
}
```

With conditional clicking based on data:

```json
{
  "type": "click_only",
  "selector": "button:has-text('${value}')",
  "data_key": "selected_option"
}
```

#### IFRAME_FIELD

For fields inside iframes (alternative to `iframe_selector`):

```json
{
  "type": "iframe_field",
  "selector": "input[name='card_number']",
  "iframe_selector": "iframe#payment-frame",
  "data_key": "card_number"
}
```

### Common Field Parameters

| Parameter | Type | Description |
|-----------|------|-------------|
| `data_key` | string | Key in input_data to get value from |
| `selector` | string | CSS/Playwright selector for the element |
| `type_config` | object | Type-specific configuration (see field type sections above) |
| `iframe_selector` | string | Selector for parent iframe if field is embedded |
| `field_visible_timeout_ms` | integer | Timeout for field to become visible |
| `post_click_delay_ms` | integer | Wait after clicking the field |
| `click_timeout_ms` | integer | Timeout for each click or check the field makes, per strategy where a click falls back; unset, clicks that fall back use the library's click timeout and the rest Playwright's. A click timeout error says whether the default applied |
| `skip_verification` | boolean | Skip value verification after filling |
| `force_click` | boolean | Use force click to bypass overlays (`click_only`, `radio`, and `click_select`) |
| `press_enter` | boolean | Press Enter after the interaction to commit the value (`text`, `date`, `iframe_field`, `combobox`, `click_select`) |
| `press_keys` | array of strings | [Playwright key names](https://playwright.dev/python/docs/api/class-keyboard#keyboard-press) pressed at page level after the field's interaction (any field type) |
| `optional` | boolean | Skip field gracefully if element is not found or interaction fails |
| `retry_config` | object | Retry configuration for transient failures (see [Retry and Resilience](#retry-and-resilience)) |

### Pressing Keys

`press_keys` sends keys to the page once the field's own interaction has finished, in the order given. Its main use is dismissing a widget that swallows the next click — a date picker left open by a `force_click`, or a custom select with no close control:

```json
{
  "type": "click_select",
  "selector": "#vehicle-fuel",
  "data_key": "fuel",
  "press_keys": ["Escape"],
  "type_config": {
    "option_selector": "ul.dropdown-content li"
  }
}
```

The keys go to whatever the page currently focuses, as a real key press would, so the field's own element does not need to still exist. This is what separates it from `press_enter`, which presses on the field itself to commit a value; the two are independent and can be combined.

An overlay that outlives the last field of a step blocks the step's `next_action`. Attach the keys to that last field rather than adding a field for them.

### Optional Fields

Fields marked with `"optional": true` are skipped gracefully when the element is not found on the page or when interaction fails. This is useful for elements that may or may not appear depending on dynamic page behavior, A/B tests, or conditional rendering that cannot be predicted by data alone.

```json
{
  "type": "click_only",
  "selector": "#promotional-banner button.dismiss",
  "data_key": "dismiss_promo",
  "optional": true,
  "field_visible_timeout_ms": 2000
}
```

**How `optional` differs from `null` values:**

| Mechanism | Element lookup | Use case |
|-----------|---------------|----------|
| Value = `null` in input_data | No (skipped immediately) | Field exists but should not be filled for this data row |
| `"optional": true` | Yes (waits for visibility) | Field may or may not exist on the page |
| Key absent from input_data | No | Raises `MissingDataError` unless the field is `optional`; use `null` to skip |

When `optional` is set and the element is not found or interaction fails (including a value that does not stick after filling, or an unusable input value), the handler logs an info message and continues to the next field. No error is raised regardless of the `strict` mode. The message says why: `Optional field not found` when nothing matches the selector, `Optional field present but not visible (N matching)` when it matches only hidden elements, and `Optional field interaction failed` with Playwright's reason (for example a disabled element) when the interaction itself failed. A required field that matches only hidden elements raises `FieldNotFoundError` with `Field not visible: ...` instead of `Field not found: ...`.

**Tip:** Set a low `field_visible_timeout_ms` (e.g., 1000-2000ms) on optional fields to avoid waiting the full default timeout when the element is absent.

**Note:** For skipping entire steps (all fields + next_action), use [Optional Steps](#optional-steps) instead.

### Action Types

#### CLICK

Click a button or link:

```json
{
  "type": "click",
  "selector": "button[type='submit']"
}
```

With iframe support:

```json
{
  "type": "click",
  "selector": "button:has-text('Next')",
  "iframe_selector": "iframe#form-frame"
}
```

#### WAIT

Wait for an element to appear:

```json
{
  "type": "wait",
  "selector": ".loading-complete"
}
```

#### WAIT_HIDDEN

Wait for an element to disappear:

```json
{
  "type": "wait_hidden",
  "selector": ".loading-spinner"
}
```

#### SCROLL

Scroll to an element:

```json
{
  "type": "scroll",
  "selector": "#section-bottom"
}
```

#### DELAY

Wait for a fixed time:

```json
{
  "type": "delay",
  "delay_ms": 2000
}
```

#### CONDITIONAL

Execute actions based on conditions:

```json
{
  "type": "conditional",
  "condition": {
    "type": "selector_visible",
    "selector": ".error-message"
  },
  "actions": [
    {
      "type": "click",
      "selector": "button.dismiss-error"
    }
  ]
}
```

Condition types:
- `selector_visible`: True if any element matching the selector is visible on the page
- `selector_hidden`: True if no element matching the selector is visible on the page
- `data_equals`: True if `input_data[key]` equals `value`
- `data_exists`: True if `input_data[key]` is truthy

#### SOLVE_CAPTCHA

Clear a captcha that appears mid-flow, typically after a submit. Place it where the challenge
appears, usually as a step's `next_action` once the submit button has been clicked as a
`click_only` field. `captcha.kind` names the challenge and picks the solver; each kind has its
own settings and its own default model.

| `kind` | Challenge | Extra |
|--------|-----------|-------|
| `recaptcha_v2` | reCAPTCHA v2 image grids ("select all images with…") | `azcrawlerpy[recaptcha]` |
| `text` | a distorted-text image with an answer box | `azcrawlerpy[text]` |

`azcrawlerpy[captcha]` installs every solver. Both share these fields:

| Field | Default | Description |
|-------|---------|-------------|
| `weights` | the kind's weights package | Path to model weights overriding the default for this kind |
| `max_rounds` | 18 | Rounds before `CaptchaNotSolvedError` |
| `appear_timeout_ms` | 15000 | Wait for a challenge; none appearing counts as success |
| `round_delay_ms` | 3500 | Pause between rounds while the page re-renders |

Every kind's default model ships in its own package, installed by the kind's extra:
`azcrawlerpy-recaptcha-google` for `recaptcha_v2` and `azcrawlerpy-recaptcha-text` for `text`.
They are published apart from azcrawlerpy, so a code release does not upload the weights again.
Each extra pins an exact version of its weights package, so every install of one azcrawlerpy
release runs the same models, and a retrained model reaches installs through an azcrawlerpy
release that bumps the pin. There is no download at run time. A
provenance card sits next to each weights file (same name, `.json`): which model variant, which
training run, which commit, and its score. The solver logs that line when it loads
the model, so a deployment's model is identifiable from its log. The
models are trained in the `captcha-training` project, which scores candidates against the pinned
release and writes a winner into a weights package checkout it is given. Override per action with
`weights`, or per machine with `AZCRAWLERPY_CAPTCHA_RECAPTCHA_V2` and `AZCRAWLERPY_CAPTCHA_TEXT`.

A challenge that does not clear raises `CaptchaNotSolvedError` naming the kind, the rounds spent
and the reason. It is never reported as cleared, so a page guarded by something the solver cannot
handle fails loudly instead of failing later for an unrelated-looking reason. If no challenge
appears at all the action succeeds, which is the normal case for invisible reCAPTCHA that passes
silently.

##### Kind `recaptcha_v2`

```json
{ "type": "solve_captcha", "captcha": { "kind": "recaptcha_v2" } }
```

| Challenge | Handled |
|-----------|---------|
| reCAPTCHA v2 image grid, 3x3 dynamic (clicked tiles are replaced; re-classified until none match) | yes |
| reCAPTCHA v2 image grid, 4x4 one-photo ("select all squares") | yes - the whole photo is read by the 4x4 graph |
| reCAPTCHA v2 **invisible**, when it escalates to an image grid | yes |
| reCAPTCHA v2 **audio** challenge | no - reported as unsupported |
| reCAPTCHA v3 (score only, no challenge) | nothing to solve; the action succeeds |
| hCaptcha, Cloudflare Turnstile, Arkose/FunCaptcha | no - reported as unsupported |

The default model is **TesseraNet**, one network for both grid kinds: a 3x3 round's tiles and a
4x4 round's photo, read whole, go through a graph per grid kind (`tesseranet_3x3.onnx`,
`tesseranet_4x4.onnx`) that takes pixels to every category's probability per square, and
`cell_head.json` beside them names each channel's category and threshold. It runs on CPU through
onnxruntime with no torch, in tens of milliseconds per round, over the reCAPTCHA categories listed below.
It has no zero-shot fallback: a category it was not trained on is skipped. `captcha-training` trains,
exports and promotes it.

A `weights` path (or `AZCRAWLERPY_CAPTCHA_RECAPTCHA_V2`) must name a TesseraNet 3x3 graph with its
`tesseranet_4x4.onnx` and `cell_head.json` beside it. Category labels are matched to the head's
channel labels ignoring case, spaces and underscores.

| Field | Default | Description |
|-------|---------|-------------|
| `prompt_labels` | — | Extra category-noun stems to labels, for widget languages beyond the built-in English, German and Italian |
| `dynamic_reload_delay_ms` | 4000 | Longest wait for replacement tiles to load |
| `dynamic_settle_delay_ms` | 5000 | Extra pause after replacement tiles load, so they are classified and clicked in turn |

##### How the category is read, in any language

reCAPTCHA renders in the page's language (the browser locale, or the `hl` parameter the site
passes). Each round resolves its category in this order, and its log line says which step decided:
`reCAPTCHA round: prompt='…' target='Bridge' via=challenge-id id=/m/015kr tiles=9`.

1. `challenge-id`: the widget's own challenge response names the category by an id that is the same
   in every language. The crawler listens for it from the moment a page opens, and a round only uses
   the response that delivered the grid on screen, never an earlier one.
2. `challenge-name`: the English name Google sends with 3x3 challenges, for an id not in the table yet.
3. `prompt-noun`: the word the widget sets in bold ("buses", "Brücken", "biciclette"), matched
   against built-in English, German and Italian stems plus `prompt_labels`. Used when the challenge
   response was not seen.
4. `unresolved`: the English name, or else the noun itself, is passed on as is; a label the weights do
   not know is skipped.

The labels are `Bicycle`, `Boat`, `Bridge`, `Bus`, `Car`, `Chimney`, `Crosswalk`, `Hydrant`,
`Motorcycle`, `Mountain`, `Palm`, `Parking Meter`, `Stairs`, `Taxi`, `Tractor`, `Traffic Light` and
`Truck`. The weights' `cell_head.json` lists the categories they know; a round asking for any other is
skipped.

**Another widget language**

`prompt_labels` only matters for rounds without a challenge id. Take the bold words from the round
log and add a stem per category, short enough to cover inflections and long enough not to appear
inside another category's word; longer stems are tried first, so `"autobus"` wins over `"auto"`:

```json
{
  "type": "solve_captcha",
  "captcha": {"kind": "recaptcha_v2", "prompt_labels": {"feux": "Traffic Light", "vélo": "Bicycle"}}
}
```

##### What determines the success rate

The approach follows the ETH Zurich *Breaking reCAPTCHAv2* paper (COMPSAC 2024), which reports a
100% solve rate on the Google demo widget with a YOLOv8 tile classifier; the default model is
TesseraNet, a different network trained on the same public tile sets plus photo datasets. Their
measurements show that the model is the smallest part of that result; what reCAPTCHA thinks of the *session* decides how many rounds it
serves before accepting an answer:

| Condition (paper, median challenges per pass) | Without | With |
|-----------------------------------------------|---------|------|
| Natural (Bezier-curve) mouse movement | 13 | 5 |
| Cookies and browsing history in the profile | 5 | 2 |
| Residential-looking IP (VPN) | flagged after ~20 runs | 100/100 passes |

Humans in the same study needed a median of 2 challenges, so a session that keeps being served
rounds after a handful of correct answers is being judged on its reputation, not its answers.
Practical consequences for a crawl:

- Keep `humanize` enabled in the browser config; Camoufox then moves the cursor along a curve for
  every tile click, which is the mouse-movement condition above.
- A warmed profile (see [Browser Profile Building](#browser-profile-building-profiler)) is the
  cookie/history condition. Dropping it costs rounds even with a perfect classifier.
- A datacenter exit IP is the dominant negative signal. The paper's bot passed every run on a
  residential-looking IP and was flagged within ~20 runs without one.
- The paper's threshold on the target-class probability is 0.2; the default weights carry their own
  per-category cuts for each grid kind in `cell_head.json`. A refused empty Verify means the grid holds
  a match no tile cleared its cut for, so the single likeliest tile is clicked.
- For 4x4 "select all squares" grids the paper segments the *whole* image and maps the mask onto
  cells. This action likewise reads the whole photo, with TesseraNet's 4x4 graph scoring every square.

##### Notes and limitations (recaptcha_v2)

- Only the **visible** challenge frame is solved. Pages rendering several widgets have several hidden
  frames holding blank tiles; those are ignored.
- Tiles are cropped from one screenshot of the grid per pass, so solving never scrolls the page.
- Clicks are paced like a hand: a pause to look at the grid, a varying gap between tiles, a pause
  before Verify, and each click lands at a random point in the middle of the tile rather than its
  exact centre. Camoufox's `humanize` supplies the cursor path; this supplies where and when it ends.
  The ETH paper measured natural pointer behaviour as the second-largest factor after the IP.
- Unknown categories are skipped (targets are sparse), which keeps the round loop alive rather than
  aborting on the first unfamiliar prompt.
- An empty answer is submitted at most once in a row. If no tile clears the confidence threshold
  and the previous round already submitted nothing, or the widget answers "Please select all
  matching images", the likeliest tile is clicked instead. Repeating an empty answer cannot clear a
  challenge that is still on screen, and it burns the round budget.
- Dynamic grids are recognised by the widget marking clicked tiles while replacements load; the
  replacements are classified again until none match, then Verify is pressed.
- **Solving is not the same as passing.** A session reCAPTCHA distrusts receives visually degraded
  tiles and effectively unlimited rounds; no classifier reliably beats that. The action gives up
  after `max_rounds` rather than looping forever, so the failure stays visible. Proxy reputation and
  session history matter at least as much as the model.
- 4x4 grids are one photo cut into squares. The squares are cropped out of the widget's gaps,
  reassembled into the photo and read whole by the 4x4 graph, which scores every square at once.
- Classifiers are cached per configuration, so repeated challenges in one process reuse loaded weights.
- The default weights are TesseraNet, trained from scratch in `captcha-training` with no pretrained
  backbone. Its training data includes public reCAPTCHA tile sets (MIT, CC BY 4.0, CC0, and one with no
  stated licence) and COCO, LVIS, ADE20K and Open Images photos, and some of its labels were produced
  by OWLv2 (Apache-2.0) and YOLOv8-seg (**AGPL-3.0**); review those terms before redistributing, or
  point `weights` at your own model.


##### Kind `text`

```json
{
  "type": "solve_captcha",
  "captcha": {
    "kind": "text",
    "image_selector": "img.captcha",
    "input_selector": "#captcha-code",
    "submit_selector": "button[type=submit]",
    "refresh_selector": "a.captcha-reload",
    "min_confidence": 0.6
  }
}
```

The image is screenshotted, read by the default SVTRv2 recogniser (ONNX, runs on CPU through
onnxruntime, no torch), and the reading is typed into the input one character at a time with a
human pace. The alphabet is digits, ASCII letters in both cases and ASCII punctuation, up to the length the recogniser was trained for (see `azcrawlerpy.captcha.text.charset`).
Before reading, a small chooser (`selector.onnx`) looks at a thumbnail and the crop's aspect ratio
and names one of four input sizes; the recogniser then reads the crop once at that size. The two are
one pair, so a `weights` path (or `AZCRAWLERPY_CAPTCHA_TEXT`) must name a recogniser with its
`selector.onnx` beside it.

| Field | Default | Description |
|-------|---------|-------------|
| `image_selector` | required | The captcha image |
| `input_selector` | required | The answer input |
| `submit_selector` | — | Clicked after typing. When set, the image disappearing counts as cleared and a new image starts another round; when unset the answer is typed and the rest of the flow submits it |
| `refresh_selector` | — | Serves a new image; clicked instead of typing when the reading is below `min_confidence` |
| `iframe_selector` | — | Frame holding the captcha, if any |
| `min_confidence` | 0.0 | Mean per-character probability under which a fresh image is requested; requires `refresh_selector` |

Rounds only exist with `submit_selector`: a page that shows a new image after a wrong answer gets
another reading until `max_rounds`. Without it the action ends after typing, so a wrong reading
surfaces wherever the flow submits.

### Common Action Parameters

| Parameter | Type | Description |
|-----------|------|-------------|
| `selector` | string | Target element selector |
| `iframe_selector` | string | Selector for parent iframe (`click`, `wait`, `wait_hidden`, `scroll`) |
| `delay_ms` | integer | Delay in ms for `delay` (randomized up to 2x); exact timeout in ms for `wait` and `wait_hidden` |
| `condition` | object | Condition definition (for `conditional` actions) |
| `actions` | array | Nested actions to execute if condition is met (for `conditional` actions) |
| `pre_action_delay_ms` | integer | Wait before executing action (`click`, `wait`, `wait_hidden`, `scroll`; `click` has a default) |
| `post_action_delay_ms` | integer | Wait after executing action (`click`, `wait`, `wait_hidden`, `scroll`; `click` has a default) |
| `retry_config` | object | Retry configuration for transient failures (see [Retry and Resilience](#retry-and-resilience)) |

### Final Page Configuration

```json
{
  "final_page": {
    "wait_for": ".result-container, .confirmation",
    "timeout_ms": 60000,
    "post_wait_delay_ms": 2000,
    "screenshot_selector": ".result-panel"
  }
}
```

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `wait_for` | string | Conditional | CSS selector to wait for (required if `wait_for_data_key` not set) |
| `wait_for_data_key` | string | Conditional | Data key for exact text match (uses `text="{value}"` selector). Required if `wait_for` not set. |
| `timeout_ms` | integer | Yes | Timeout in milliseconds for waiting |
| `post_wait_delay_ms` | integer | No | Delay in ms after selector found, for SPA content to render (default: 0) |
| `screenshot_selector` | string | No | Element to screenshot (null for full page) |

Only one of `wait_for` or `wait_for_data_key` can be set (not both).

### Post Final Page Steps

`post_final_page_steps` is an optional list of regular step definitions that run on the result page after its `wait_for`/`post_wait_delay_ms` settle and before `data_extraction`, the end screenshot, and the HTML capture. Use it to interact with the results (load more entries, switch coverage tabs) that a normal step can never reach because the result page only exists after all steps and CAPTCHA handling.

Each entry supports the full step semantics (`wait_for`, `optional`, `restart_crawl`, extraction, retries). Use `{"type": "delay", "delay_ms": 0}` as `next_action` when a step only extracts. Step extractions additionally support a repeat-click loop for load-more buttons:

```json
{
  "post_final_page_steps": [
    {
      "name": "load_all_results",
      "wait_for": "div[data-test^='tarifkachel/']",
      "timeout_ms": 60000,
      "fields": [],
      "post_field_extraction": [
        {
          "name": "tariff_count",
          "selector": "div[data-test^='tarifkachel/']",
          "click_before": "button:has-text('weitere Ergebnisse laden')",
          "repeat_click_until_gone": true,
          "max_repeat_clicks": 20,
          "wait_after_click_ms": 1500
        }
      ],
      "next_action": { "type": "delay", "delay_ms": 0 }
    }
  ]
}
```

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `repeat_click_until_gone` | boolean | No | Keep clicking `click_before` until it disappears, `max_repeat_clicks` is reached, or the extraction's `selector` match count stops growing after a click (no-progress guard). Set `wait_after_click_ms` high enough for new content to render, or the guard fires early. |
| `max_repeat_clicks` | integer | No | Safety cap on clicks when `repeat_click_until_gone` is true (must be >= 1) |

The repeat-click loop is available on any step extraction, but is most useful here. The whole loop runs inside the step's `timeout_ms` budget.

### Per-Step Page Extraction

`instructions.data_extraction` runs once, after every step, so it only ever captures the UI state the last step left behind. When a results page changes in response to clicks — a coverage tab, a filter, a sort order — each state needs its own list snapshot. `page_extraction` puts a full `DataExtractionConfig` on a step and runs it at the end of that step, after `post_field_extraction` and before `next_action`, so it sees the DOM the step's own clicks produced.

```json
{
  "name": "screen_vk_sb500",
  "wait_for": "[data-test='deckung/VK']",
  "timeout_ms": 120000,
  "optional": true,
  "data_extraction": [
    { "name": "_switch_vk", "selector": "[data-test='deckung/VK']",
      "click_before": "[data-test='deckung/VK']", "wait_after_click_ms": 6000 },
    { "name": "_load_more", "selector": "div[data-test^='tarifkachel/']",
      "click_before": "button:has-text('weitere Ergebnisse laden')",
      "repeat_click_until_gone": true, "max_repeat_clicks": 20 }
  ],
  "page_extraction": {
    "fields": {
      "insurer_names": { "selector": "div[data-test^='tarifkachel/'] img[alt]", "attribute": "alt", "multiple": true }
    }
  },
  "fields": [],
  "next_action": { "type": "delay", "delay_ms": 2000 }
}
```

Results are namespaced under the step name, so several steps can extract the same field names without overwriting each other:

```python
result.extracted_data["screen_vk_sb500"]["insurer_names"]  # [... 98 items ...]
```

Because every extraction shares one flat result dict and `instructions.data_extraction` is merged over it last, `Instructions` rejects configs where a `page_extraction` step name would collide with another snapshot step, a step extraction name, or a final-page field name.

Failures follow the same `restart_crawl > optional > strict` chain as the other extraction slots, so a snapshot step marked `optional: true` cannot fail the crawl. Extraction cost scales with fields × matched elements and counts against the step's `timeout_ms`, so budget list-heavy snapshots accordingly — an exhausted budget on an `optional` step silently drops the snapshot.

### Data Extraction Configuration

Extract structured data from the final page using CSS selectors:

```json
{
  "data_extraction": {
    "fields": {
      "tier_prices": {
        "selector": ".price-value",
        "attribute": null,
        "regex": "([0-9]+[.,][0-9]{2})",
        "multiple": true,
        "iframe_selector": "iframe#form-frame"
      },
      "selected_price": {
        "selector": "#total-amount",
        "attribute": "data-value",
        "regex": null,
        "multiple": false
      }
    }
  }
}
```

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `selector` | string | Yes | CSS selector to locate element(s) |
| `attribute` | string | No | Element attribute to extract (null for text content) |
| `regex` | string | No | Regex pattern to apply (uses first capture group if present) |
| `multiple` | boolean | No | True for list of all matches, False (default) for first match only |
| `iframe_selector` | string | No | CSS selector for iframe if element is inside one |
| `pad_missing` | boolean | No | With `multiple: true`, append `null` for missing/non-matching values instead of dropping them, keeping list indices aligned with the matched elements (safe for `nested_fields` over sparse fields) |

The `data_extraction` config also supports `nested_fields` to combine flat extracted arrays into structured outputs (`paired_dict` or `object_list`). See [docs/README.md](azcrawlerpy/docs/README.md) for details.

Extracted data is available in the crawl result:

```python
result = await crawler.crawl(...)
print(result.extracted_data)
# {'tier_prices': ['32,28', '35,26', '50,34'], 'selected_price': '35,26'}
```

## Data Points (input_data)

The `input_data` dictionary provides values for form fields. Keys must match `data_key` values in the instructions.

### Structure

```json
{
  "email": "user@example.com",
  "first_name": "John",
  "last_name": "Doe",
  "birthdate": "1985-06-15",
  "country": "Germany",
  "accept_terms": true,
  "newsletter": false,
  "premium_option": null
}
```

### Value Types

| Type | Description | Example |
|------|-------------|---------|
| String | Text values, dropdown selections | `"John"` |
| Boolean | Checkbox/radio toggle | `true`, `false` |
| Null | Skip this field entirely (no DOM interaction) | `null` |
| Integer/Float | Numeric inputs, sliders | `12000`, `99.99` |

Note: Setting a value to `null` skips the field without any DOM interaction. This is different from `"optional": true` on the field definition, which attempts the interaction but tolerates failure. See [Optional Fields](#optional-fields).

Every `data_key` a non-optional field references must be present in the data row: an absent key raises `MissingDataError`, so a misspelt key cannot silently leave a field unfilled. Use `null` to skip a field.

### Radio Button Patterns

**Pattern A - Mutually exclusive options with value selector**:
```json
{
  "gender": "male"
}
```
Selector uses `${value}` placeholder: `input[value='${value}']`

**Pattern B - Boolean flags for each option**:
```json
{
  "option_a": true,
  "option_b": null,
  "option_c": null
}
```
Only the option with `true` gets clicked.

### Date Handling

Dates in input_data should use ISO format (`YYYY-MM-DD`):
```json
{
  "birthdate": "1985-06-15",
  "start_date": "2024-01-01"
}
```

The framework converts to the format specified in the field definition.

## Element Discovery

The `ElementDiscovery` class scans web pages to identify interactive elements, helping build instructions.json files. It launches the browser the way the crawler does, so an optional `browser_config` (`CrawlerBrowserConfig`) applies its proxy, `ignore_https_errors` and, on chromium, `context_options`.

```python
from pathlib import Path
from azcrawlerpy import BrowserType, ElementDiscovery

async def discover_elements():
    discovery = ElementDiscovery(headless=False, browser_type=BrowserType.CAMOUFOX)

    report = await discovery.discover(
        url="https://example.com/form",
        output_dir=Path("./discovery_output"),
        cookie_consent={
            "banner_selector": "#cookie-banner",
            "accept_selector": "button.accept"
        },
        explore_iframes=True,
        screenshot=True,
        viewport_width=1920,
        viewport_height=1080,
        wait_timeout_ms=None,
    )

    print(f"Found {report.total_elements} elements")

    for text_input in report.text_inputs:
        print(f"Text input: {text_input.selector}")
        print(f"  Suggested type: {text_input.suggested_field_type}")

    for dropdown in report.selects:
        print(f"Dropdown: {dropdown.selector}")
        print(f"  Options: {dropdown.options_preview}")

    for radio_group in report.radio_groups:
        print(f"Radio group: {radio_group.group_name}")
        for option in radio_group.options:
            print(f"  - {option.label_text}: {option.selector}")
```

### Discovery Report Contents

- `text_inputs`: Text, email, phone, password fields
- `textareas`: Multi-line text areas
- `selects`: Native dropdown elements with options
- `radio_groups`: Grouped radio buttons
- `checkboxes`: Checkbox inputs
- `buttons`: Clickable buttons
- `links`: Anchor elements
- `date_inputs`: Date picker fields
- `file_inputs`: File upload fields
- `sliders`: Range inputs
- `custom_components`: Non-standard interactive elements
- `iframes`: Discovered iframes with their elements

## AI Agent Guidance

This section provides instructions for AI agents tasked with creating `instructions.json` and `input_data` files.

### Workflow for Creating Instructions

1. **Discovery Phase**: Use `ElementDiscovery` to scan each page/step of the form
2. **Mapping Phase**: Map discovered elements to field definitions
3. **Flow Definition**: Define step transitions and actions
4. **Data Schema**: Create the input_data structure

### Step-by-Step Process

#### 1. Analyze the Form Structure

- Identify how many pages/steps the form has
- Note the URL pattern changes (if any)
- Identify what element appears when each step loads

#### 2. For Each Step, Define:

```json
{
  "name": "<descriptive_step_name>",
  "wait_for": "<selector_that_confirms_step_loaded>",
  "timeout_ms": 15000,
  "fields": [...],
  "next_action": {...}
}
```

**Naming conventions**:
- Use snake_case for step names: `personal_info`, `payment_details`
- Use descriptive data_keys: `first_name`, `email_address`, `accepts_terms`

#### 3. Selector Priority

When choosing selectors, prefer in order:
1. `[data-testid='...']` or `[data-cy='...']` - Most stable
2. `[aria-label='...']` or `[aria-labelledby='...']` - Accessible and stable
3. `input[name='...']` - Form field names
4. `:has-text('...')` - Text content (use for buttons/labels)
5. CSS class selectors - Least stable, avoid if possible

#### 4. Handle Dynamic Content

For AJAX-loaded content:
- Use `wait` action before interacting
- Add `field_visible_timeout_ms` to field definitions
- Use `post_click_delay_ms` for fields that trigger updates

#### 5. Radio Button Strategy

**Option A - When radio values are meaningful**:
```json
{
  "type": "radio",
  "selector": "input[type='radio'][value='${value}']",
  "data_key": "payment_type"
}
// data: { "payment_type": "credit_card" }
```

**Option B - When you need individual control**:
```json
{
  "type": "radio",
  "selector": "[role='radio']:has-text('Credit Card')",
  "data_key": "payment_credit_card",
  "force_click": true
},
{
  "type": "radio",
  "selector": "[role='radio']:has-text('PayPal')",
  "data_key": "payment_paypal",
  "force_click": true
}
// data: { "payment_credit_card": true, "payment_paypal": null }
```

#### 6. Iframe Handling

When elements are inside iframes:
```json
{
  "type": "text",
  "selector": "input[name='card_number']",
  "iframe_selector": "iframe#payment-iframe",
  "data_key": "card_number"
}
```

### Creating input_data

#### 1. Analyze Required Fields

From the instructions, extract all unique `data_key` values:
```python
data_keys = set()
for step in instructions["steps"]:
    for field in step["fields"]:
        if field.get("data_key"):
            data_keys.add(field["data_key"])
```

#### 2. Determine Value Types

| Field Type | Data Type | Example |
|------------|-----------|---------|
| text, textarea | string | `"John Doe"` |
| dropdown | string | `"Germany"` |
| radio (value-driven) | string | `"option_a"` |
| radio (boolean) | boolean/null | `true` or `null` |
| checkbox | boolean | `true` / `false` |
| date | string (ISO) | `"1985-06-15"` |
| slider | number | `50000` |
| file | string (path) | `"/path/to/file.pdf"` |

#### 3. Handle Mutually Exclusive Options

For radio groups with boolean flags, only ONE should be `true`:
```json
{
  "employment_fulltime": true,
  "employment_parttime": null,
  "employment_selfemployed": null,
  "employment_unemployed": null
}
```

#### 4. Date Format

Always provide dates in ISO format in input_data:
```json
{
  "birthdate": "1985-06-15",
  "policy_start": "2024-01-01"
}
```

The instructions specify the output format for the specific form.

### Common Patterns

#### Multi-Step Wizard
```json
{
  "steps": [
    {
      "name": "step_1_personal",
      "wait_for": "input[name='firstName']",
      "timeout_ms": 15000,
      "fields": [...],
      "next_action": { "type": "click", "selector": "button:has-text('Next')" }
    },
    {
      "name": "step_2_address",
      "wait_for": "input[name='street']",
      "timeout_ms": 15000,
      "fields": [...],
      "next_action": { "type": "click", "selector": "button:has-text('Next')" }
    }
  ]
}
```

#### Form with Loading States
```json
{
  "next_action": {
    "type": "click",
    "selector": "button[type='submit']",
    "post_action_delay_ms": 1000
  }
}
```

#### Conditional Fields
```json
{
  "type": "conditional",
  "condition": {
    "type": "data_equals",
    "key": "has_additional_driver",
    "value": true
  },
  "actions": [
    {
      "type": "click",
      "selector": "button:has-text('Add Driver')"
    }
  ]
}
```

## Retry and Resilience

The framework provides built-in retry and fallback mechanisms for handling transient failures during web interactions.

### Retry Configuration

Add `retry_config` to any field, action, step extraction, or data extraction to enable automatic retry with exponential backoff:

```json
{
  "type": "text",
  "selector": "input[name='email']",
  "data_key": "email",
  "retry_config": {
    "max_attempts": 3,
    "base_delay_ms": 500,
    "backoff_multiplier": 1.5
  }
}
```

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `max_attempts` | integer | Yes | Total attempts (1 = no retry, 2 = one retry, etc.) |
| `base_delay_ms` | integer | Yes | Base delay between retries in milliseconds |
| `backoff_multiplier` | float | No | Multiplier per retry (default: 1.5). Delay = `base_delay_ms * backoff_multiplier^(attempt-1)` |

Retryable errors include Playwright `TimeoutError` and `Error` (element not found, intercepted clicks, etc.), and on fields a value that does not stick after filling.

### Click Fallback Chain

All click interactions (`click` actions, `click_only`, `radio`, `click_select` fields) use an escalating fallback strategy when a click fails:

1. **Normal click** (or force click if the field sets `force_click`)
2. **Force click** (or normal click if the field sets `force_click`) -- bypasses actionability checks
3. **JS dispatchEvent** -- fires a `MouseEvent` via `element.dispatchEvent()`
4. **JS element.click()** -- calls `element.click()` via `document.querySelector()`

This handles common real-world issues like overlay interception, sticky headers, and elements that pass Playwright's visibility checks but fail to receive pointer events.

Steps 3 and 4 only run when `browser_config.allow_synthetic_events` is on: a click dispatched from JS reaches the page with `isTrusted=false`, which bot detection reads as automation.

### Value Verification

`text` and `textarea` fields verify the actual input value after filling. If the value written to the field does not match the expected value, a `FieldInteractionError` is raised once `retry_config` attempts are exhausted (an `optional` field is skipped instead). This catches silent data corruption from autofill interference, input masks, or JavaScript reformatting.

Set `skip_verification: true` on the field to disable this check for fields where the site intentionally transforms the input (e.g., phone number formatting).

### ComboBox Degraded Retry

When `retry_config` is set on a `combobox` field, retries use degraded parameters to handle slow autocomplete responses:
- Typing delay scales by `1.5^(attempt-1)` (slower typing gives autocomplete more time)
- The field is always cleared before retyping on retry attempts

## Error Handling and Diagnostics

The framework provides detailed error information when failures occur.

### Exception Types

Every `CrawlerError` that `crawl()` raises carries a `partial_result`: a
`CrawlResult` holding whatever had been collected before the failure
(screenshots, `extracted_data`, `steps_completed`, `profiler_visited_urls`,
`platform_logs`, `crawler_log`, `stdout_log`, `stderr_log`,
`playwright_debug_log`, `error_diagnostics`, and the
[debug recording](#debug-recording): `run_info`, `page_events`,
`network_log`, `route_hits`). That holds for failures before the page loop
too, such as a profiler error. Catch the exception and read
`e.partial_result` to recover that data — see
[Debugging and Crawl Output](#debugging-and-crawl-output) below.

| Exception | When Raised |
|-----------|-------------|
| `FieldNotFoundError` | Selector doesn't match any element (not raised for `optional` fields) |
| `FieldInteractionError` | Element found but interaction failed, or value mismatch after fill (not raised for `optional` fields) |
| `CrawlerTimeoutError` | A step's `timeout_ms` budget was exhausted (initial `wait_for` OR the rest of the step body — extractions, fields, action, retries), a restart budget was exhausted, OR the crawl's `global_timeout_ms` wall-clock cap was hit |
| `NavigationError` | Navigation action failed |
| `MissingDataError` | A non-optional field's data_key is absent from input_data (`null` skips a field instead) |
| `InvalidInstructionError` | An instruction the models accept but the crawl cannot use, such as an unsupported date format. Malformed instructions raise pydantic's `ValidationError` before the crawl starts |
| `UnsupportedFieldTypeError` | Not raised: validation rejects an unknown field type. Kept for compatibility |
| `UnsupportedActionTypeError` | Not raised: validation rejects an unknown action type. Kept for compatibility |
| `IframeNotFoundError` | Not raised: a missing iframe surfaces as the timeout of the element inside it. Kept for compatibility |
| `DataExtractionError` | Data extraction from final page failed |
| `CaptchaNotSolvedError` | A `solve_captcha` action could not clear its challenge within `max_rounds`, or met a challenge its solver does not support |
| `CrawlerError` | Anything that fails outside the page loop, e.g. the profiler; `original_error` holds the cause |

### Timeout Budgets

Every timeout is used exactly as configured (or as its built-in default). Interaction delays are randomized between 1x and 2x to pace interactions like a person; the extraction waits `wait_before_ms` and `wait_after_click_ms` and `final_page.post_wait_delay_ms` are exact.

The crawler enforces two independent wall-clock budgets. Both translate
their exhaustion into `CrawlerTimeoutError` (or `CrawlRestartError` when
restart is configured) and attach a partial result.

**1. Per-step budget — `step.timeout_ms`**

Hard upper bound on the total wall-clock duration of one step. Covers
`wait_for`, `data_extraction`, field handlers (including retries),
`next_action` (including retries and delays), `post_field_extraction` and
`page_extraction`. The same field is also used as the timeout for the initial
`wait_for_selector` call, so a step that uses its full budget waiting for
the selector leaves no time for the rest of the body. On exhaustion, the
priority chain `restart_crawl > optional > strict` decides the outcome:

- `restart_crawl=True` → `CrawlRestartError` → full browser restart (up to
  `max_restarts`).
- `optional=True` → step is skipped silently.
- `strict=True` → `CrawlerTimeoutError`.
- `strict=False` → warning logged, step returns.

When the budget runs out, the crawler reports where it went, phase by phase:
the log line carries `time_spent=[...]` and the `CrawlerTimeoutError`
message ends with, for example,
`Time spent: wait_for 812ms, field 'birth_date' 1203ms, field 'brand' 6120ms (did not finish)`.
Retries and debug screenshots count against the budget and show up in that
breakdown. Every completed step logs the same breakdown in
`Step completed: name='...' time_spent=[...]`, so a failing run can be
compared with a successful one.

**2. Crawl-level budget — `crawler.crawl(..., global_timeout_ms=...)`**

Optional wall-clock cap on the entire crawl (profiler, navigation, all
steps and the final-page capture). `None` (the default) disables it. When
the deadline fires, the crawl task is cancelled and
`CrawlerTimeoutError(step_name="<global_timeout>", selector="<global>",
timeout_ms=...)` is raised with a partial `CrawlResult` attached.
The partial carries `screenshots`, `extracted_data`, `steps_completed`,
`profiler_visited_urls`, `platform_logs`, `crawler_log`, `stdout_log`,
`stderr_log`, `playwright_debug_log` and the debug recording collected up
to the moment of cancellation. `html` is empty and `final_url` falls back to the start URL
because the live page is being torn down by the cancellation and cannot
be read.

```python
from azcrawlerpy.crawling.exceptions import CrawlerTimeoutError

try:
    result = await crawler.crawl(
        url=url,
        input_data=input_data,
        instructions=instructions,
        debug_mode=DebugMode.ALL,
        global_timeout_ms=300_000,  # 5 minutes
    )
except CrawlerTimeoutError as e:
    if e.step_name == "<global_timeout>":
        # Global budget hit -- inspect partial_result for what was collected.
        partial = e.partial_result
        if partial is not None:
            print(f"Timed out after {partial.steps_completed} step(s)")
            print(f"Screenshots: {len(partial.screenshots)}")
    else:
        # Per-step budget or restart-budget exhaustion.
        raise
```

## Debugging and Crawl Output

A crawl hands back everything it saw in memory, on a `CrawlResult`. It writes
nothing to disk apart from browser profiles and a temporary file for
Playwright's debug stream, which it reads back and deletes. Where the output
ends up (a folder, blob storage, a database) is the caller's decision. One
argument, `debug_mode`, sets how much is recorded, and the same output comes
back whether the crawl succeeds or fails.

### Debug Mode

```python
from azcrawlerpy import DebugMode

result = await crawler.crawl(
    url=url,
    input_data=input_data,
    instructions=instructions,
    debug_mode=DebugMode.ALL,  # or "all"; the default is DebugMode.END
)
```

`debug_mode` takes a `DebugMode` or its value as a string: `"none"`,
`"start"`, `"end"` or `"all"`.

| Output | `NONE` | `START` | `END` (default) | `ALL` |
|--------|:------:|:-------:|:---------------:|:-----:|
| `html`, `final_url`, `extracted_data`, `steps_completed`, `profiler_visited_urls` | ✓ | ✓ | ✓ | ✓ |
| [Screenshot](#screenshots) `start`, once the start URL has loaded and cookie consent has run | | ✓ | | ✓ |
| A screenshot after every field and after every step's action | | | | ✓ |
| Screenshot `end`, once the final page and the post steps are done | | | ✓ | ✓ |
| Screenshot `final`, after the final extraction | | ✓ | ✓ | ✓ |
| [Error screenshot and `error_diagnostics`](#error-diagnostics) when the crawl fails | | ✓ | ✓ | ✓ |
| [Captured logs](#captured-logs): `crawler_log`, `stdout_log`, `stderr_log`, `playwright_debug_log` | | ✓ | ✓ | ✓ |
| [Debug recording](#debug-recording): `run_info`, `page_events`, `network_log`, `route_hits` | | ✓ | ✓ | ✓ |
| [`platform_logs`](#azure-app-service-log-capture), on Azure App Service only | | | | ✓ |

Choosing a mode:

- **`NONE`** records nothing beyond the result itself. The logger still
  receives every line. Use it when nothing will be kept.
- **`END`**, the default, is the cheapest mode that still explains a failure:
  the error screenshot, diagnostics, logs and recording, plus two pictures of
  a successful end.
- **`START`** takes a picture of the landing page instead of the `end` one, to
  see what the site served: a block page, a consent wall or another variant of
  the form.
- **`ALL`** takes a picture after every field and every action, which is what
  writing or repairing instructions needs. It is also the most expensive
  mode. Each screenshot is taken inside its step and counts against the step's
  `timeout_ms`, showing up as `screenshot` in the step's `time_spent`. Each
  one holds a full-page JPEG and the full HTML, so a long form makes a large
  result. On Azure App Service it also collects the platform log files.

Recording is best effort throughout. A screenshot or capture that fails, for
example while a click is navigating the page, is skipped and logged as a
warning. Debug capture never fails a crawl, and never replaces the error of
one that failed.

### Where the Output Lands

| How the crawl ends | What you get |
|--------------------|--------------|
| Success | The `CrawlResult` that `crawl()` returns |
| A failure while a page is open: a field, an action, a step or final-page timeout, an extraction | A `CrawlerError` whose `partial_result` holds the `html` and URL of the page it failed on, the error screenshot as the last screenshot, and `error_diagnostics`. Exceptions that are not a `CrawlerError` are wrapped in one, with the cause on `original_error` |
| A step's restart budget runs out | A `CrawlerTimeoutError` whose `partial_result` is built from the last attempt's page, as above |
| `global_timeout_ms` runs out | A `CrawlerTimeoutError` with `step_name="<global_timeout>"`. Its `partial_result` holds everything recorded so far, but the page is being torn down: `html` is empty, `final_url` is the start URL, and there is no error screenshot or diagnostics |
| A failure before the first page opens, such as a profiler error | A `CrawlerError` whose `partial_result` holds the logs and the recording, but no page output |
| Invalid instructions, or a `url` that differs from the instructions' | pydantic's `ValidationError` or a `ValueError`, raised before anything is recorded |
| The task is cancelled from outside, or the process is killed | Nothing. The output lives in memory until `crawl()` returns |

The last row is why a harness with its own time limit, such as a function
timeout or a queue lease, should set `global_timeout_ms` below it. The crawl
then ends itself and hands back a partial result.

`partial_result` has the same shape as a successful result, so one function
can store both:

```python
from azcrawlerpy import CrawlerError, DebugMode

try:
    result = await crawler.crawl(
        url=url,
        input_data=input_data,
        instructions=instructions,
        debug_mode=DebugMode.END,
        global_timeout_ms=240_000,  # below the host's own limit
    )
except CrawlerError as error:
    if error.partial_result is not None:
        save_crawl_output(error.partial_result, run_folder)  # see Persisting the Output
    raise
save_crawl_output(result, run_folder)
```

### The Crawl Result

| Field | Type | Contents |
|-------|------|----------|
| `url` | `str` | The start URL |
| `input_data` | `dict` | The data the crawl filled in, as passed |
| `instructions` | `dict` | The instructions, as passed |
| `final_url` | `str` | The page URL at the end, or where the crawl failed; the start URL after a global timeout |
| `html` | `str` | The page HTML at the end, or where the crawl failed; empty after a global timeout |
| `steps_completed` | `int` | Steps that ran to their end. Skipped steps and post steps are not counted |
| `extracted_data` | `dict` | Values from step extractions, page extraction snapshots and the final extraction |
| `profiler_visited_urls` | `list[str]` | The URLs the profiler chose to visit |
| `screenshots` | `list[Screenshot]` | See [Screenshots](#screenshots) |
| `error_diagnostics` | `dict \| None` | See [Error Diagnostics](#error-diagnostics); `None` on success |
| `crawler_log`, `stdout_log`, `stderr_log`, `playwright_debug_log` | `bytes` | See [Captured Logs](#captured-logs) |
| `run_info`, `page_events`, `network_log`, `route_hits` | models | See [Debug Recording](#debug-recording) |
| `platform_logs` | `dict[str, bytes]` | See [Azure App Service Log Capture](#azure-app-service-log-capture) |

`print(result)` prints a one-screen summary: the URLs, the steps completed,
the input and instruction keys, the number and total size of the screenshots,
the HTML length, the extracted keys, and the number of page events and
requests.

### Screenshots

Each `Screenshot` is one capture of the page:

| Field | Contents |
|-------|----------|
| `label` | `NNN_name`: a sequence number, then what was captured, with every character other than letters, digits, `_` and `-` replaced by `_`. Labels are unique within a crawl, so they can serve as file names |
| `attempt` | The browser attempt the screenshot was taken in, starting at 1 |
| `image` | JPEG bytes of the whole page. A page taller or wider than `MAX_CAPTURE_DIMENSION_PX` is captured as its top-left slice |
| `html` | The page HTML at the same moment; empty when it could not be read |
| `form_state` | Every `input`, `select`, `textarea` and `button` of the main frame at the same moment |

The names after the sequence number say when the screenshot was taken:

| Name | Taken | Modes |
|------|-------|-------|
| `start` | Once the start URL has loaded and cookie consent has run | `START`, `ALL` |
| `{step}_{field}` | After each field that did not fail. `{field}` is the field's `data_key`, or its selector when it has none | `ALL` |
| `{step}_after_action` | After the step's `next_action`, when it did not fail | `ALL` |
| `end` | Once the final page's selector has appeared and the post steps have run | `END`, `ALL` |
| `final` | After the final extraction. When `final_page.screenshot_selector` is set, the image shows only that element, while `html` and `form_state` still cover the page | `START`, `END`, `ALL` |
| `error_{step}` | When the crawl fails, with the error diagnostics. `{step}` is the step running, or the last one run when the final page fails; `error_unknown` before the first step | all but `NONE` |

Post steps take field and action screenshots like any other step. A restart
keeps the screenshots of the attempts before it, and their numbering carries
on, so `002_start` with `attempt=2` follows `001_start` with `attempt=1`.
A restart with budget left takes no error screenshot: with `END`, the attempt
it ended leaves no picture at all, while `START` and `ALL` keep that
attempt's own screenshots.

`form_state` lists each control with `tag`, `type`, `id`, `name`, `value`,
`text` (the selected option of a select, or a button's text), `checked`
(checkboxes and radios only), `disabled` and `visible` (a non-empty box that
is not `visibility: hidden`). Values are cut to `MAX_FORM_VALUE_CHARS`, and
the value of a password input is never recorded. Comparing the `form_state`
of two screenshots shows which value changed, or which control appeared or
went away, without reading the HTML.

### Error Diagnostics

When a crawl fails while a page is open, and `debug_mode` is not `NONE`, the
crawler inspects the failing page before building the partial result. It
captures, each on its own, so one that fails costs none of the others:

- a full-page error screenshot, first, while the page is closest to the failure;
- the page URL, title and time, the step, and the selector that failed;
- every `[data-cy]` element with its tag, type and text;
- the visible buttons and input fields;
- the `data-cy` values closest to the failed selector, as suggested fixes.

The findings land in three places:

1. **`partial_result.error_diagnostics`**, a JSON-ready dict:

   ```json
   {
     "timestamp": "2026-01-01T12:00:03.456+00:00",
     "success": false,
     "error": {
       "type": "FieldNotFoundError",
       "message": "Field not found: selector='[data-cy='hsn-input']' type='text' in step='vehicle'",
       "step_name": "vehicle",
       "failed_selector": "[data-cy='hsn-input']"
     },
     "diagnostics": {
       "url": "https://example.com/quote/vehicle",
       "title": "Your vehicle",
       "timestamp": "2026-01-01T12:00:03.456+00:00",
       "failed_selector": "[data-cy='hsn-input']",
       "step_name": "vehicle",
       "available_data_cy_selectors": [
         {"selector": "hsn", "tag": "input", "type": "text", "text": null}
       ],
       "visible_buttons": ["[data-cy='next'] \"Next\""],
       "visible_inputs": ["[data-cy='hsn'] (text) HSN"],
       "suggested_selectors": ["hsn"],
       "has_screenshot_bytes": true
     }
   }
   ```

2. **The exception.** `error.diagnostics` is the same capture as a
   `PageDiagnostics` object, with the image in `screenshot_bytes`. `str(error)`
   appends a readable `=== AI DEBUG INFO ===` block listing the selectors,
   buttons, inputs and suggestions. `error.to_dict()` returns a JSON-ready
   summary of the error, its diagnostics, the partial result and the original
   error.

3. **The screenshots.** The error screenshot is also the last entry of
   `partial_result.screenshots`, labelled `NNN_error_{step}`, with the page's
   `html` and `form_state`.

```python
from pathlib import Path

from azcrawlerpy import CrawlerError

try:
    result = await crawler.crawl(url=url, input_data=input_data, instructions=instructions)
except CrawlerError as error:
    print(error)  # the message, then the AI DEBUG INFO block
    if error.diagnostics is not None and error.diagnostics.screenshot_bytes is not None:
        Path("error.jpg").write_bytes(error.diagnostics.screenshot_bytes)
    partial = error.partial_result
    if partial is not None and partial.error_diagnostics is not None:
        print(partial.error_diagnostics["diagnostics"]["suggested_selectors"])
        print(partial.steps_completed, partial.final_url)
    raise
```

No diagnostics are taken after a global timeout, on a restart that still has
budget, or with `NONE`. Console output and network traffic are not part of
them; they are in the [debug recording](#debug-recording).

### Captured Logs

With any mode other than `NONE`, the crawl copies its own output into memory
from the moment `crawl()` starts until it returns:

| Field | Contents |
|-------|----------|
| `crawler_log` | Every log record emitted during the crawl, by azcrawlerpy and by any library it calls, one line each: `2026-01-01 12:00:03.456 - azcrawlerpy.crawling.crawler - INFO - Step completed: ...` |
| `stdout_log`, `stderr_log` | Everything written to `sys.stdout` and `sys.stderr` during the crawl. It is copied, not diverted: the process streams still receive it |
| `playwright_debug_log` | Playwright's own debug stream (`pw:api`, `pw:browser*`) from the driver process: every API call it received, and the browser's launch and exit messages |

All four are UTF-8 bytes, ready to write to a file.

- **Levels.** A record reaches `crawler_log` only if its logger lets it
  through. The logger passed to `FormCrawler` decides for the crawl's own
  lines, most of which are `INFO`, so a logger set to `WARNING` leaves out
  most of the useful ones. The library's module loggers (`azcrawlerpy.*`),
  which report the click ladder, retries and page captures, follow the level
  of the `azcrawlerpy` logger or the root, `WARNING` by default; set
  `logging.getLogger("azcrawlerpy").setLevel(logging.INFO)` to capture their
  `INFO` lines too. Other libraries' records appear at the levels their own
  loggers pass.
- **Concurrent crawls.** Each crawl captures only its own output, even when
  several run at once in one process: capture follows the crawl's asyncio
  context rather than the process.
- **Playwright's stream.** The crawler starts its drivers with `DEBUG` and
  `DEBUG_FILE` pointing at a temporary file, reads the file back when the
  crawl ends and deletes it. Values already set in the environment win, so
  setting `DEBUG_FILE` yourself leaves `playwright_debug_log` empty. The
  profiler's browsers are captured too.

Lines worth searching `crawler_log` for:

| Line | Meaning |
|------|---------|
| `Starting crawl: url=... steps=... debug=...` | The crawl began |
| `Camoufox identity: exit_ip=... locale=... preset=...` | The device and location a camoufox crawl presents |
| `Egress check: attempt=... status=... response=...` | What the IP echo service saw (see [Network Exit](#network-exit)) |
| `Processing step: name=...`, `Wait condition met: ...` | A step started, and its `wait_for` matched |
| `Step completed: name=... time_spent=[...]` | Where the step's time went, phase by phase (see [Timeout Budgets](#timeout-budgets)) |
| `Page URL changed: url=... attempt=...` | Every main-frame URL change, including client-side routing (not with `NONE`) |
| `Click failed via ...`, `JS click fallbacks skipped ...` | A rung of the [click ladder](#click-fallback-chain) failed |
| `Retry n/max: ...`, `All n attempts exhausted: ...` | A retry, and the retry that gave up |
| `Field failed (non-strict)`, `Action failed (non-strict)`, `Step wait timeout (non-strict, skipping step)`, `Step budget exhausted (non-strict)` | What a lenient crawl stepped over |
| `Optional step skipped (...)` | An optional step gave up, and why |
| `Restart requested (...)`, `Restarting crawl: step=... attempt=n/max`, `Max restarts exhausted` | The restart path |
| `Debug screenshot skipped`, `Page capture failed`, `Error diagnostics could not be captured` | A debug capture that did not happen, and why |
| `Blocked URL pattern matched no request`, `URL rewrite matched no request` | A URL rule that never acted: probably misconfigured (not with `NONE`) |
| `Network capture: n body read(s) still unfinished ...` | Bodies still being read when the browser closed |
| `Crawl recording: attempts=... page_events=... (dropped n) network_entries=... (dropped n) ...` | The recording's summary at the end of the crawl |
| `Global timeout exhausted` | The crawl-wide budget ran out |
| `Platform log capture preflight: active=...` | Whether platform logs will be collected, and which condition stopped them |

### Debug Recording

With any mode other than `NONE`, the result, or the `partial_result` of a
failed crawl, carries a record of what the crawl's page did, across every
browser attempt. Every event and request carries the `attempt` it belongs
to, like the screenshots, so a restart keeps what the earlier attempts saw, and a retry that succeeded still
shows why the first attempt failed.

| Field | Contents |
|-------|----------|
| `run_info` | How the crawl ran (see below) |
| `page_events` | Console messages, uncaught page errors and main-frame URL changes, oldest first |
| `network_log` | The page's requests in the order they were seen |
| `route_hits` | How many requests each URL rule acted on |

**`run_info`** holds what differs between two runs of the same instructions:

| Field | Contents |
|-------|----------|
| `started_at` | UTC start time in ISO 8601. Every `at_ms` in the recording counts from it |
| `duration_ms` | Wall-clock duration of the crawl |
| `versions` | Python, azcrawlerpy, Playwright, camoufox and the launched browser build; `None` when unknown |
| `platform` | The operating system the crawl ran on |
| `headless`, `debug_mode`, `strict`, `global_timeout_ms` | The `FormCrawler` and `crawl()` settings |
| `crawler_browser_config` | The `CrawlerBrowserConfig`, with the proxy password masked |
| `attempts` | Browsers launched; above 1 after restarts |
| `geoip_ip` | The exit IP camoufox's geoip placed the fingerprint at (see [Network Exit](#network-exit)) |
| `egress_checks` | One per attempt when `egress_check_url` is set: `attempt`, `url`, `status`, `response` (cut to a fixed length) or `error` |

**`page_events`** entries have `at_ms`, `attempt`, `kind`, `level` and `text`:

| `kind` | `level` | `text` |
|--------|---------|--------|
| `console` | The console message type: `log`, `warning`, `error`, ... | The message |
| `pageerror` | `None` | The uncaught error, with its stack |
| `navigation` | `None` | The URL the main frame changed to, including `history.pushState` routing |

**`network_log`** holds one `NetworkEntry` per request:

| Field | Contents |
|-------|----------|
| `at_ms`, `attempt` | When the request was seen, and in which attempt |
| `method`, `url`, `resource_type` | The request, and Playwright's resource type: `document`, `xhr`, `fetch`, `image`, ... |
| `status` | The response status; `None` when no response arrived |
| `failure` | Why the request failed: the browser's error text, such as `net::ERR_CONNECTION_REFUSED`, or `blocked by pattern '...'` |
| `duration_ms` | From the request to its end; `None` while it was still running when the attempt ended |
| `request_headers`, `response_headers` | Only for URLs in `capture_url_substrings` |
| `request_body`, `response_body` | For URLs in `capture_url_substrings`, and for failed requests (see below) |
| `capture_note` | Why a header or body is missing or cut |

Which requests get an entry, and how much of each is kept:

| Request | Entry | Headers | Bodies |
|---------|-------|---------|--------|
| URL contains one of `browser_config.capture_url_substrings`, any type | ✓ | ✓ | request and response |
| `document`, `xhr` or `fetch` answered with a 4xx or 5xx status | ✓ | | request and response |
| `document`, `xhr` or `fetch` the network dropped | ✓ | | request |
| `document`, `xhr` or `fetch` aborted by `blocked_url_patterns` | ✓, `failure` names the pattern | | |
| Any other `document`, `xhr` or `fetch` | ✓ | | |
| Any other type (image, script, font, ...) that fails | ✓ | | |
| Any other type blocked by `blocked_url_patterns` | Counted in `route_hits` only | | |
| Any other type that succeeds | | | |

A failed request keeps its bodies without anyone having to guess its URL in
advance. The server's own answer, such as a `400` with a trace id, is often
the whole explanation of a failure. Headers are kept only for listed URLs,
because they carry session cookies and tokens.

Bodies are stored as text when their content type is text-like (`text/*`,
JSON, XML, JavaScript, form-encoded, multipart or GraphQL); any other body is
replaced by a placeholder such as `<5120 bytes of image/png>`. Response bodies
are read once the response has finished. Reads still running when an attempt
ends get `DRAIN_TIMEOUT_S` before its browser closes.

**`route_hits`** has `blocked`, the requests aborted per `blocked_url_patterns`
entry, and `rewritten`, the requests changed per `url_rewrites` key. A rule
that matched nothing in the whole crawl is also logged as a warning, because
it is probably misconfigured.

The recording covers the crawl's page only: the profiler's pages and the
egress check's tab are not recorded. It is bounded, so a chatty page cannot
exhaust memory. The limits are named in
[`crawling/constants.py`](azcrawlerpy/crawling/constants.py):

| Limit | What happens past it |
|-------|----------------------|
| `MAX_PAGE_EVENTS`, `MAX_NETWORK_ENTRIES` | The oldest entries are dropped; the summary line counts them |
| `MAX_TEXT_CHARS` | An event's text is cut, noting how many characters were cut |
| `MAX_BODY_CHARS` | A body is cut, and `capture_note` says so |
| `BODY_BUDGET_CHARS` | All kept bodies share this budget; the oldest are dropped first, and their `capture_note` says so |

Reading the recording:

```python
for event in result.page_events:
    if event.kind == "navigation":
        print(f"{event.at_ms:>7} ms  attempt {event.attempt}  -> {event.text}")

for entry in result.network_log:
    if entry.failure or (entry.status or 0) >= 400:
        print(entry.attempt, entry.method, entry.url, entry.status or entry.failure)
        if entry.response_body:
            print("   ", entry.response_body[:500])
```

### Network Exit

Set `egress_check_url` on `CrawlerBrowserConfig` to an IP echo service, for
example `https://ipinfo.io/json`, which also returns the location and network
owner. Before each attempt starts, the crawler loads that URL in a separate
tab of the same browser, so it leaves through the same proxy exit as the
crawl. The response is logged in every mode, and recorded in
`run_info.egress_checks` unless the mode is `NONE`.

With `geoip` enabled on camoufox, the crawler looks up the exit IP once per
crawl and passes it to every launch, so the profiler warmup and every
restart claim the same location; `run_info.geoip_ip` records it. Comparing
it with the egress check shows whether the timezone and locale the
fingerprint claims match the exit the site actually sees. A proxy that
rotates the exit per connection makes them differ.

The locale follows `CrawlerBrowserConfig.locale` (e.g. `"de-DE"`). Left unset with `geoip` on, it is the exit country's most spoken language, where camoufox would otherwise draw one at random by population share and present rare locales such as `bar-DE`. Each camoufox crawl also fixes one device for all its launches: a fingerprint preset plus the canvas, audio and font noise seeds, font list and voices that camoufox would redraw on every launch.

```python
from azcrawlerpy import CrawlerBrowserConfig, ProxyConfig

browser_config = CrawlerBrowserConfig(
    proxy=ProxyConfig(server="http://proxy.example.com:8080", username="user", password="pass"),
    geoip=True,
    egress_check_url="https://ipinfo.io/json",
)
```

### Azure App Service Log Capture

When the crawler runs on an Azure App Service Plan target (Web App, Function App, or Web App / Function App in a container), `debug_mode=DebugMode.ALL` automatically captures the platform log tree so it can be attached to the run's output artifacts. It holds the host-side view of the process, including the worker's stdout and stderr, so it is often the only way to see low-level browser-launch failures in post-mortem. Playwright's own debug stream is captured separately into `playwright_debug_log`.

**When it activates (two gating conditions):**

- `debug_mode=DebugMode.ALL`, AND
- `$HOME` is set **and** `$HOME/LogFiles/` is a directory.

If either gating check fails, the crawl is otherwise untouched and a
clear reason is logged (see "Log output" below).

**Sink:**

The full `$HOME/LogFiles/` tree is read into a `dict[str, bytes]` and
returned on `CrawlResult.platform_logs`. Keys are forward-slash relative
paths (e.g. `"Application/Functions/Worker/worker.log"`); values are the
raw file bytes. On graceful failure the same dict is attached to
`CrawlerError.partial_result.platform_logs`, so harnesses that catch
`CrawlerError` can still recover them. The crawler never writes to disk
— the caller owns the sink.

**What gets collected:**

The entire `$HOME/LogFiles/` tree, recursively. No filtering. Files include:

- `Application/Functions/Worker/*.log` (Python worker stdout/stderr on Functions)
- `Application/Functions/Host/` (Functions host logs)
- `<instance>_default_docker.log` (Web App container stdout/stderr)
- `kudu/`, `Nginx/`, `docker.log`, etc.

On Linux App Service that path is `/home/LogFiles/`; on Windows App Service it is typically `C:\home\LogFiles\` or `D:\home\LogFiles\`. The detection uses `$HOME` so both work without configuration.

**Usage example (harness uploads to its own storage):**

```python
result = await crawler.crawl(
    url=url,
    input_data=input_data,
    instructions=instructions,
    debug_mode=DebugMode.ALL,  # activates platform log capture
)

# result.platform_logs is a dict[str, bytes]
for rel_path, content in result.platform_logs.items():
    blob_client.upload_blob(
        name=f"{run_id}/platform_logs/{rel_path}",
        data=content,
        overwrite=True,
    )
```

**Crash-safe:**

The collection is wrapped in `try / finally` around the entire crawl body, so it runs on:

- Successful return
- `CrawlerTimeoutError` from an exhausted restart budget
- Any `CrawlerError` wrapping a browser crash
- Profiler-phase failures (the profiler is inside the `try`)
- `KeyboardInterrupt` / task cancellation

Per-file errors (Windows log rotation, permission issues, races) are logged as warnings and counted; they never mask the crawl's original exception.

**Log output:**

Every invocation emits one start line plus one outcome line via the configured logger (flows to Application Insights through `azpaddypy`):

| Situation | Level | Message |
|---|---|---|
| Helper invoked | INFO | `Platform log capture starting` |
| `$HOME` not set (local dev, non-Azure) | INFO | `Platform log capture skipped: $HOME not set (not an Azure App Service target)` |
| `$HOME/LogFiles` missing | INFO | `Platform log capture skipped: '<path>' not found (not an Azure App Service target)` |
| `$HOME/LogFiles` exists but is a file | WARNING | `Platform log capture skipped: '<path>' exists but is not a directory` |
| Cannot stat `$HOME/LogFiles` (permissions) | WARNING | `Platform log capture skipped: cannot stat '<path>' error=<exc>` |
| Source found but empty | WARNING | `Platform log capture found source but collected nothing: source='<path>' (directory is empty or contains no regular files)` |
| Files collected | INFO | `Platform logs collected: source='<s>' files=<N> bytes=<B> errors=<E>` |

**Deployment requirements per target:**

| Target | Works out of the box? | Notes |
|---|---|---|
| Azure Functions on App Service Plan | Yes | Functions host always writes to `/home/LogFiles/Application/Functions/Worker/`. Playwright Node stderr inherits that fd. |
| Web App for Containers on App Service Plan | Yes, with two app settings | `WEBSITES_ENABLE_APP_SERVICE_STORAGE=true` (default) so `/home` is the Azure Files mount, **and** App Service Logs → Filesystem enabled so stdout/stderr is written there. Enable via `az webapp log config --docker-container-logging filesystem --level information`. |
| Pure Azure Container Apps | No | Container Apps does not mount `/home/LogFiles/`; logs go to Log Analytics only. The helper correctly INFO-logs "not found" and skips. |
| Local dev (`docker run`, native Python) | No (by design) | `$HOME/LogFiles/` doesn't exist; helper skips and logs the reason. |

**Delivery:**

The collected bytes land on `CrawlResult.platform_logs` on success. On graceful failure (any `CrawlerError`, including `CrawlerTimeoutError` from an exhausted restart budget or the `global_timeout_ms` wall-clock cap), they land on `exception.partial_result.platform_logs`.

**Shared-worker noise:**

`/home/LogFiles/Application/Functions/Worker/*.log` is shared across all concurrent function invocations on the same worker instance. The copied snapshot therefore includes any overlapping activity — there is no per-invocation filtering. This is an intentional "full copy always" trade for forensic completeness.

### Persisting the Output

The crawler never picks a sink. This function writes a result, or a failed
crawl's partial result, to a folder; the same fields map directly onto blob
names or database columns:

```python
import json
from pathlib import Path

from azcrawlerpy import CrawlResult

RECORDING_FIELDS = {
    "url", "final_url", "steps_completed", "extracted_data", "error_diagnostics",
    "run_info", "page_events", "network_log", "route_hits",
}


def save_crawl_output(result: CrawlResult, folder: Path) -> None:
    folder.mkdir(parents=True, exist_ok=True)
    (folder / "page.html").write_text(result.html, encoding="utf-8")
    for shot in result.screenshots:
        (folder / f"{shot.label}.jpg").write_bytes(shot.image)
        (folder / f"{shot.label}.html").write_text(shot.html, encoding="utf-8")
        form_state = [control.model_dump() for control in shot.form_state]
        (folder / f"{shot.label}.form_state.json").write_text(json.dumps(form_state, indent=2), encoding="utf-8")
    for name in ("crawler_log", "stdout_log", "stderr_log", "playwright_debug_log"):
        (folder / f"{name}.log").write_bytes(getattr(result, name))
    (folder / "recording.json").write_text(result.model_dump_json(include=RECORDING_FIELDS, indent=2), encoding="utf-8")
    for relative_path, content in result.platform_logs.items():
        target = folder / "platform_logs" / relative_path
        target.parent.mkdir(parents=True, exist_ok=True)
        target.write_bytes(content)
```

### Sensitive Data in the Output

Store the debug output like the input data itself:

- `input_data` and `instructions` are returned as passed.
- `html`, every screenshot and its `html` show the page as it was, typed
  values included. `form_state` never holds a password input's value.
- With the `azcrawlerpy` loggers at `INFO`, `crawler_log` records the values
  the field handlers fill, and `playwright_debug_log` records the calls
  Playwright received.
- Request bodies are the submitted form data, and captured headers carry
  cookies and tokens.
- `run_info.crawler_browser_config` masks the proxy password, and the
  telemetry spans carry no input data.

### Reading a Failed Crawl

Start with the exception message: it names the step, the selector and the
error. A budget error ends with `Time spent: ...`, and with diagnostics the
`=== AI DEBUG INFO ===` block follows. Then look where the symptom points:

| Symptom | Where to look |
|---------|---------------|
| A `wait_for` or field selector never matched | The error screenshot; `error_diagnostics["diagnostics"]` for the visible inputs, buttons and suggested selectors; the `form_state` of the last screenshots |
| The site ended on its own error page | `page_events` of kind `navigation`, for when the URL changed; then the `network_log` entries of that attempt with a 4xx or 5xx status, whose bodies hold the server's answer |
| A step ran out of budget | `Time spent: ...` in the error message, or `time_spent=[...]` in `crawler_log`, compared with a successful run |
| A restart recovered, and you want to know why the first attempt failed | Screenshots, `page_events` and `network_log` entries with `attempt == 1` |
| A click never landed | `Click failed via ...` lines in `crawler_log`. Raise the field's `click_timeout_ms` when a click times out |
| Blocked or rewritten URLs do not behave | `route_hits`, and the warnings for rules that matched nothing |
| It works locally and fails in the cloud | `run_info` of both runs: `versions`, `headless`, `platform`, `egress_checks` and `geoip_ip` |
| The browser does not start | `stderr_log` and `playwright_debug_log`, and on App Service `platform_logs` |
| Something the page logged | `page_events` of kind `console` and `pageerror` |

For crawl-level telemetry in Application Insights, such as spans per crawl and
per attempt, see [Logging & Telemetry](#logging--telemetry).

## Browser Profile Building (Profiler)

The profiler visits random sites before the main crawl to accumulate cookies and storage, making the browser appear more natural. This is configured via the `profiler` field in instructions.json.

### Profiler Modes

The profiler supports two storage modes:

**Disk mode** (`storage_path` set): Persists the browser profile to disk.
- Camoufox: saves as a Firefox user data directory (`{storage_path}.camoufox/`)
- Chromium: saves as a JSON file (`{storage_path}.chromium.json`)

**In-memory mode** (`storage_path` omitted or `null`): No permanent files written.
- Chromium: returns storage state as a dict, passed directly to `browser.new_context(storage_state=dict)`
- Camoufox: uses a temporary directory, passed as `user_data_dir` during the crawl and deleted when the crawl ends, whether it succeeds or fails

### Profiler Configuration

```json
{
  "profiler": {
    "sites": [
      {
        "url": "https://www.google.de",
        "cookie_consent": {
          "banner_selector": "[role='dialog'], .cookie-banner",
          "accept_selector": "button[id*='accept'], button[id*='agree']",
          "js_fallback_texts": ["akzept", "accept", "zustimm"]
        },
        "browse_delay_ms": 2000
      },
      {
        "url": "https://en.wikipedia.org",
        "browse_delay_ms": 3000
      }
    ],
    "visit_count": 2,
    "ignore_errors": true,
    "storage_path": null,
    "inter_site_delay_ms": 1000
  }
}
```

### ProfilerConfig Parameters

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `sites` | array | Yes | List of `ProfilerSiteConfig` objects to randomly select from |
| `visit_count` | integer | Yes | Number of sites to randomly visit (must be <= length of `sites`) |
| `ignore_errors` | boolean | Yes | If true, continue profiling when a single site visit fails |
| `storage_path` | string or null | No | Base path for profile storage. When null (default), runs in-memory mode |
| `bootstrap_path` | string or null | No | Existing profile to seed each run from (see [Bootstrap Profiles](#bootstrap-profiles)); in-memory mode only, so it cannot be combined with `storage_path` |
| `inter_site_delay_ms` | integer | No | Delay in ms between visiting each site |

### Bootstrap Profiles

A profile with no history is a signal in itself: reputation systems score a browser that has never
existed before far lower than one with aged cookies. `storage_path` solves that by persisting the
profile, but it makes runs stateful — concurrent crawls then share and mutate one directory.

`bootstrap_path` keeps runs stateless. The seed is treated as **read-only**: it is copied into the
run's throw-away profile before the warmup, and every change the crawl makes stays in the copy.

```json
{
  "profiler": {
    "visit_count": 1,
    "ignore_errors": true,
    "bootstrap_path": "./browser_profiles/it-seed",
    "sites": [{ "url": "https://www.youtube.com/", "browse_delay_ms": 15000 }]
  }
}
```

- **Camoufox**: point it at a `user_data_dir` directory; it is copied with `copytree`.
- **Chromium**: point it at a storage state JSON file; it is passed as the context's `storage_state`.
- A missing or wrong-typed path logs a warning and is skipped, so a bad seed never fails a crawl.

Create a seed once, by hand, then reuse it: launch a browser with `persistent_context=True` and a
`user_data_dir`, browse normally for a while (signing in to accounts whose cookies you want the seed
to carry), and close it. Point `bootstrap_path` at that directory. Refresh it periodically — cookies
expire, and a stale seed loses its value.

Note the seed is shared *state*, not shared *identity*: each run still gets its own device, so a
seed used across many concurrent runs presents the same cookies from different machines. Keep seeds
per-target, and prefer several seeds over one when running at volume.

### Fingerprint Consistency

Camoufox picks a new fingerprint on every launch. The profiler and the crawl are two launches, so
without intervention the crawl would present the profiler's cookies from a subtly different device —
`navigator.hardwareConcurrency` changing between the warmup and the crawl, for example.

Every camoufox crawl pins one `fingerprint_preset` and passes it to all its launches, so the
warmup, the crawl and every restart are the same machine. Presets are real captured devices and
are the mechanism camoufox documents for this; hand-built fingerprints trigger a leak warning and
are avoided deliberately. The preset follows `CrawlerBrowserConfig.os`, so it stays consistent with the
spoofed platform.

Nothing to configure — it applies automatically whenever `browser_type` is `camoufox`. The canvas,
audio and font noise seeds, font list and voices, which camoufox would redraw on every launch, are
fixed per crawl as well.

### ProfilerSiteConfig Parameters

Each site in the `sites` array supports per-site cookie consent handling:

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `url` | string | Yes | URL to visit for profile building |
| `cookie_consent` | object | No | Cookie consent config for this specific site (same schema as top-level `cookie_consent`) |
| `browse_delay_ms` | integer | No | Time in ms to linger on the page after cookie consent handling |

### ProfilerResult

The profiler returns a `ProfilerResult` with:

| Field | Type | Description |
|-------|------|-------------|
| `visited_urls` | list[str] | URLs that were visited during profiling |
| `storage_state_path` | Path or None | Path to persisted profile (disk mode), or temp dir path (Camoufox in-memory) |
| `storage_state` | dict or None | In-memory storage state dict (Chromium in-memory mode only) |
| `is_temporary` | bool | True when `storage_state_path` is a throw-away directory the consuming crawl deletes |

The crawler automatically passes the profiler result to the browser context creation, so no manual wiring is needed.

## Examples

### Insurance Quote Form

**instructions.json**:
```json
{
  "url": "https://insurance.example.com/quote",
  "browser_config": {
    "browser_type": "camoufox",
    "viewport_width": 1920,
    "viewport_height": 1080
  },
  "cookie_consent": {
    "banner_selector": "#cookie-banner",
    "accept_selector": "button:has-text('Accept')"
  },
  "steps": [
    {
      "name": "vehicle_info",
      "wait_for": "input[name='hsn']",
      "timeout_ms": 15000,
      "fields": [
        {
          "type": "text",
          "selector": "input[name='hsn']",
          "data_key": "vehicle_hsn"
        },
        {
          "type": "text",
          "selector": "input[name='tsn']",
          "data_key": "vehicle_tsn"
        },
        {
          "type": "date",
          "selector": "input[name='registration_date']",
          "data_key": "first_registration",
          "type_config": {
            "format": "DD.MM.YYYY"
          }
        }
      ],
      "next_action": {
        "type": "click",
        "selector": "button:has-text('Continue')"
      }
    },
    {
      "name": "personal_info",
      "wait_for": "input[name='birthdate']",
      "timeout_ms": 15000,
      "fields": [
        {
          "type": "date",
          "selector": "input[name='birthdate']",
          "data_key": "birthdate",
          "type_config": {
            "format": "DD.MM.YYYY"
          }
        },
        {
          "type": "text",
          "selector": "input[name='zipcode']",
          "data_key": "postal_code"
        }
      ],
      "next_action": {
        "type": "click",
        "selector": "button:has-text('Get Quote')"
      }
    }
  ],
  "final_page": {
    "wait_for": ".quote-result",
    "timeout_ms": 60000,
    "screenshot_selector": ".quote-panel"
  }
}
```

**data_row.json**:
```json
{
  "vehicle_hsn": "0603",
  "vehicle_tsn": "AKZ",
  "first_registration": "2020-03-15",
  "birthdate": "1985-06-20",
  "postal_code": "80331"
}
```

### Form with Iframes

```json
{
  "steps": [
    {
      "name": "embedded_form",
      "wait_for": "iframe#form-frame",
      "timeout_ms": 15000,
      "fields": [
        {
          "type": "text",
          "selector": "input[name='email']",
          "iframe_selector": "iframe#form-frame",
          "data_key": "email"
        },
        {
          "type": "dropdown",
          "selector": "select[name='plan']",
          "iframe_selector": "iframe#form-frame",
          "data_key": "selected_plan",
          "type_config": {
            "select_by": "text"
          }
        }
      ],
      "next_action": {
        "type": "click",
        "selector": "button:has-text('Submit')",
        "iframe_selector": "iframe#form-frame"
      }
    }
  ]
}
```

## License

MIT
