Metadata-Version: 2.4
Name: suparse
Version: 1.3.0
Summary: Official Python SDK and CLI for the Suparse Document Processing API
License-Expression: MIT
License-File: LICENSE
Keywords: document-processing,ocr,invoice,receipt,api,sdk
Author: Suparse
Author-email: support@suparse.com
Requires-Python: >=3.10
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Typing :: Typed
Requires-Dist: httpx (>=0.27.0)
Requires-Dist: pydantic (>=2.6.0)
Requires-Dist: pydantic-settings (>=2.2.0)
Requires-Dist: tenacity (>=8.2.0)
Project-URL: Homepage, https://suparse.com
Description-Content-Type: text/markdown

# Suparse Python SDK

[![PyPI version](https://img.shields.io/pypi/v/suparse)](https://pypi.org/project/suparse/)
[![Python](https://img.shields.io/pypi/pyversions/suparse)](https://pypi.org/project/suparse/)

Official Python SDK and CLI for the [Suparse](https://suparse.com) Document Processing API.

Suparse is an AI-powered document processing API for extracting structured data from any document type, including invoices, receipts, bank statements, purchase orders and many more. This SDK wraps the REST API with a Python client and CLI.

## Requirements

- Python 3.10+
- Dependencies: `httpx`, `pydantic`, `pydantic-settings`, `tenacity`

## Installation

```bash
pip install suparse
```

## Authentication

You'll need an API key to use the SDK or CLI. To obtain one:

1. Sign in at [suparse.com](https://suparse.com)
2. Go to the **API Keys** tab
3. Enter a key name and click **Generate New Key**
4. Copy the key value — it will be shown only once

Set it as an environment variable:

```bash
export SUPARSE_API_KEY="your_api_key_here"
```

For the CLI, you can also store it in `~/.config/suparse/config.json`:

```json
{
  "apiKey": "your_api_key_here"
}
```

Or pass it directly to the SDK constructor (see below).

## Quick Start

```bash
# CLI
suparse process invoice.pdf -o results.json

# Python — synchronous (with SUPARSE_API_KEY env var set)
python3 -c "
from suparse import SuparseClient

with SuparseClient() as client:
    result = client.extract('invoice.pdf')
    for r in result.succeeded:
        print(r.original_file, r.data)
"

# Python — asynchronous
python3 -c "
import asyncio
from suparse import AsyncSuparseClient

async def main():
    async with AsyncSuparseClient() as client:
        result = await client.extract('invoice.pdf')
        for r in result.succeeded:
            print(r.original_file, r.data)

asyncio.run(main())
"
```

## CLI Usage

Run `suparse --help` or `suparse process --help` for full usage information.

Set your API key and process a file directly from the terminal. The SDK will auto-upload, poll, and download the resulting JSON.

### Process a Document

```bash
export SUPARSE_API_KEY="your_api_key_here"

# Auto-detect template
suparse process path/to/invoice.pdf -o results.json

# Use specific template without auto-splitting
suparse process path/to/invoice.pdf --template-id 276a0aa8-84bc-4491-a2e7-1ea13381790c

# Auto-split a multi-page PDF containing mixed document types (e.g. receipts + bank statements)
suparse process path/to/merged.pdf --with-split

# Use specific template with auto-splitting
suparse process path/to/invoice.pdf --template-id 276a0aa8-84bc-4491-a2e7-1ea13381790c --with-split

# Process and auto-delete documents from server after download
suparse process path/to/invoice.pdf --cleanup

# Export directly to XLSX
suparse process path/to/invoice.pdf --format xlsx

# Choose the flat XLSX layout
suparse process path/to/invoice.pdf --format xlsx --xlsx-layout flat

# Export CSV using original template columns
suparse process path/to/invoice.pdf --format csv --export-type original

# Process an encrypted PDF without putting the password in shell history
suparse process path/to/encrypted.pdf --prompt-pdf-password

# Assign uploaded documents to a folder
suparse process path/to/invoice.pdf --folder-id <folder-id>

# Start automatic schema creation and wait for the generated template
suparse quick-create path/to/example.pdf --wait

# Inspect a previously started schema-creation run
suparse quick-create-status <run-id>
```

### Process a Folder

Process all supported files (`.pdf`, `.jpg`, `.jpeg`, `.png`, `.heic`, `.heif`) in a folder. All files are uploaded and polled individually, then results are exported to a single JSON file.

```bash
# Process all supported files in a folder
suparse process --folder path/to/receipts/

# Output to a specific file (default: {folder_name}_results.json)
suparse process --folder path/to/receipts/ -o all_results.json

# Process a folder with a specific template
suparse process --folder path/to/receipts/ --template-id 276a0aa8-84bc-4491-a2e7-1ea13381790c

# Process a folder with auto-splitting enabled
suparse process --folder path/to/receipts/ --with-split

# Process a folder and auto-delete documents from server after download
suparse process --folder path/to/receipts/ --cleanup

# Export a folder to CSV, XLSX, or Google Sheets
suparse process --folder path/to/receipts/ --format csv
suparse process --folder path/to/receipts/ --format xlsx -o ./exports
suparse process --folder path/to/receipts/ --format google_sheets
```

### Delete Documents

```bash
# Delete one or more documents by ID (prompts for confirmation)
suparse delete <document_id>
suparse delete <id1> <id2> <id3>

# Skip confirmation prompt
suparse delete <id1> <id2> -y
```

Deleting a parent document automatically deletes all its child documents (server-side cascade).

### List Available Templates

Templates define how a document type (invoice, receipt, bank statement, etc.) is parsed. Before processing a document, check which templates are already assigned to your account:

```bash
# List templates assigned to your account (table format)
suparse templates

# List templates in JSON format
suparse templates --format json

# Include all system templates (not yet assigned to your account)
suparse templates --include-system

# Request complete template records, including schema and saved document IDs
suparse templates --view full --format json
```

The recommended way to process documents is with auto-split enabled (`--with-split`), which handles both single-type and mixed document types automatically:

```bash
suparse process path/to/documents.pdf --with-split -o results.json
```

If you consistently process one document type, look up the template ID and pass it directly:

1. Run `suparse templates` to see templates assigned to your account.
2. Use the matching template ID:
   ```bash
   suparse process invoice.pdf --template-id <id> -o results.json
   ```
3. If no template matches, run `suparse templates --include-system` to browse all system templates. Assign one to your account via the Suparse UI.
4. If no system template fits, create a custom template using the template creator in the Suparse UI.

### Automatic Schema Creation

Quick schema creation accepts a path, string path, or open binary stream and
returns typed result models. `wait_for_quick_schema_creation()` preserves
completed-with-errors and failed-document details instead of converting them to
generic exceptions.

```python
from suparse import SuparseClient

with SuparseClient() as client:
    started = client.start_quick_schema_creation("example.pdf")
    result = client.wait_for_quick_schema_creation(started.run_id)
    print(result.model_dump(mode="json"))
```

Use `start_quick_schema_creation(..., pdf_password=...)` for encrypted PDFs.
The multipart upload is buffered so an idempotent retry can resend the complete
body, and caller-owned streams are not closed. Automatic transient retries are
enabled only when both `upload_batch_id` and `source_upload_id` are supplied;
without that identity the start request is attempted once to avoid creating
duplicate runs after an ambiguous network failure.

The `quick-create` and `quick-create-status` CLI commands return exit code 1 for
failed runs and for `completed_with_errors`; queued, processing, and fully
completed results return 0.

### Detailed Upload Confirmations

`upload_file()` remains the compatibility API and always returns a task ID
string. Advanced callers that need replay or batched-upload metadata can use
`upload_file_detailed()`:

```python
from pathlib import Path

with SuparseClient() as client:
    confirmation = client.upload_file_detailed(
        Path("invoice.pdf"),
        upload_batch_id="batch-id-from-your-upload-batch-workflow",
        source_upload_id="source-id-from-your-upload-batch-workflow",
    )
    print(confirmation.model_dump(mode="json"))
```

The public SDK does not yet create upload batches, so `upload_batch_id` must be
obtained from the surrounding batch workflow. Existing `extract()` and
`extract_folder()` calls remain task-based; one `pdf_password` or `folder_id`
provided to a multi-file call is applied to every input.

Upload confirmation creates server-side processing state. The compatibility
`upload_file()` path, and detailed confirmations without both replay identities,
are therefore attempted once; the SDK does not automatically repeat an
ambiguous confirmation. A network failure or 5xx response on that path raises
`SuparseUploadConfirmationUnknownError`, whose `upload_reference_id` identifies
the upload that may already have been accepted. Do not automatically submit the
same confirmation again. Detailed confirmations with both `upload_batch_id` and
`source_upload_id` use the server's batched replay protection for bounded
transient retries; `503` responses are surfaced without retry because they may
indicate that the document was committed but queue dispatch failed.

### Structured API Errors

HTTP API exceptions retain their existing subclasses, `status_code`, and raw
`response_body`. When the server returns a structured error, `SuparseAPIError`
and its subclasses also expose `.code` and `.meta`:

```python
from suparse import SuparseAPIError

try:
    client.list_templates()
except SuparseAPIError as exc:
    print(exc.code, exc.meta, exc.response_body)
```

### Configuration

The CLI reads settings from environment variables or a `.env` file in the working directory:

| Variable | Default | Description |
|----------|---------|-------------|
| `SUPARSE_API_URL` | `https://api.suparse.com/api/v1/` | API base URL |
| `SUPARSE_API_KEY` | — | Your API key (required) |
| `POLL_INTERVAL` | `5` | Seconds between polling attempts |
| `MAX_POLL_ATTEMPTS` | `300` | Max polling attempts before timeout |
| `LOG_LEVEL` | `INFO` | Logging level |

Priority: CLI flags > environment variables > `.env` file > defaults.

For API keys, the CLI also checks `~/.config/suparse/config.json` after `SUPARSE_API_KEY` and `.env`. The config file should contain a JSON object with an `apiKey` string.

### Global Options

These options apply to all subcommands.

| Option | Description |
|--------|-------------|
| `--api-url` | API URL (default: from `SUPARSE_API_URL` env var) |
| `--api-key` | API Key (default: from `SUPARSE_API_KEY` env var) |
| `-v`, `--verbose` | Enable verbose (DEBUG) output |

### CLI Options

| Command | Option | Description |
|---------|----------|-------------|
| `process` | `--folder` | Process all supported files in a folder (mutually exclusive with file path) |
| `process` | `-o`, `--output` | Output file path, or output directory for file exports when omitted |
| `process` | `--template-id` | Template ID to use (default: auto-detect) |
| `process` | `--with-split` | Auto-split multi-page PDFs containing mixed document types (e.g. receipts + bank statements) so each is processed with the correct template (default: off) |
| `process` | `--cleanup` | Delete documents from server after download so files are stored only during processing and cannot be accessed later (default: off) |
| `process` | `--format` | Export format: `json`, `csv`, `xlsx`, or `google_sheets` (default: `json`) |
| `process` | `--export-type` | Export mode for CSV, XLSX, and Google Sheets: `unified` or `original` (default: `unified`) |
| `process`, `export` | `--xlsx-layout` | XLSX layout: `standard` or `flat` (default: `standard`) |
| `process` | `--folder-id` | Folder ID applied to every uploaded input |
| `process` | `--prompt-pdf-password` | Prompt securely for an encrypted PDF password |
| `templates` | `--format` | Output format: `table` or `json` (default: `table`) |
| `templates` | `--include-system` | Include system templates not yet assigned to your account (default: off) |
| `templates` | `--view` | Template response view: `summary` or `full` (default: `summary`) |
| `quick-create` | `--wait` | Wait for automatic schema creation to reach a terminal state |
| `quick-create` | `--format` | Output format: `table` or `json` (default: `table`) |
| `quick-create-status` | `--format` | Output format: `table` or `json` (default: `table`) |
| `delete` | `-y`, `--yes` | Skip confirmation prompt |

## Python SDK Usage

The SDK provides both synchronous and asynchronous clients. Both handle API rate limits, parallel processing, and connection pooling automatically.

### Synchronous (Recommended for Scripts, Pandas, Jupyter Notebooks)

```python
from suparse import SuparseClient

with SuparseClient() as client:
    result = client.extract(["invoice1.pdf", "invoice2.pdf"])

    for r in result.succeeded:
        print(r.original_file, r.data)
```

### Asynchronous (Recommended for FastAPI, Aiohttp, Async Pipelines)

```python
import asyncio
from suparse import AsyncSuparseClient

async def main():
    async with AsyncSuparseClient() as client:
        result = await client.extract("invoice.pdf")
        for r in result.succeeded:
            print(r.original_file, r.data)

asyncio.run(main())
```

The client reads `SUPARSE_API_KEY` and `SUPARSE_API_URL` from environment variables by default, so you can simply use `SuparseClient()` or `AsyncSuparseClient()` if those are set.

To preserve exact bulk export attachment bytes and its server-provided filename, use
`download_documents()`; `export_documents()` remains the parsed convenience API.

### Extract One or More Documents

`extract()` is the primary API. It accepts a single file, a list of files, or any iterable (like `Path.glob`). Each input can be a string path, a `Path` object, or an open file handle.

`BatchResult.total` counts input sources. `BatchResult.succeeded` contains the
task export groups returned by the API, so one input can produce multiple
successful entries when splitting or template grouping is involved. Progress
callbacks run once per returned export group (or once for a failed input).
Callbacks preserve API order within one input, but callbacks from concurrently
processed inputs may be interleaved.

#### Synchronous

```python
from pathlib import Path
from suparse import SuparseClient

with SuparseClient() as client:
    # Single file (string or Path)
    result = client.extract("invoice.pdf")
    for r in result.succeeded:
        print(r.original_file, r.data)

    # List of files from different locations
    result = client.extract([
        "./receipts/jan.pdf",
        Path("./receipts/feb.pdf"),
    ])

    # Glob generator
    result = client.extract(
        Path("./receipts/").glob("*.pdf")
    )

    # With progress callback
    def print_progress(r):
        if isinstance(r, FailedResult):
            print(f"  Failed: {r.file} - {r.error}")
        else:
            print(f"  Done: {r.original_file}")

    result = client.extract(
        files=Path("./receipts/").glob("*.pdf"),
        on_progress=print_progress,
    )

    # Access results
    for r in result.succeeded:
        print(r.original_file, r.data, r.document_ids)

    for f in result.failed:
        print(f.file, f.error)

    print(f"Total: {result.total}")
```

#### Asynchronous

```python
import asyncio
from pathlib import Path
from suparse import AsyncSuparseClient

async def main():
    async with AsyncSuparseClient() as client:
        # Single file (string or Path)
        result = await client.extract("invoice.pdf")
        for r in result.succeeded:
            print(r.original_file, r.data)

        # List of files from different locations
        result = await client.extract([
            "./receipts/jan.pdf",
            Path("./receipts/feb.pdf"),
        ])

        # Glob generator
        result = await client.extract(
            Path("./receipts/").glob("*.pdf")
        )

        # With progress callback
        def print_progress(r):
            if isinstance(r, FailedResult):
                print(f"  Failed: {r.file} - {r.error}")
            else:
                print(f"  Done: {r.original_file}")

        result = await client.extract(
            files=Path("./receipts/").glob("*.pdf"),
            on_progress=print_progress,
        )

        # Access results
        for r in result.succeeded:
            print(r.original_file, r.data, r.document_ids)

        for f in result.failed:
            print(f.file, f.error)

        print(f"Total: {result.total}")
```

### Extract a Folder

`extract_folder()` is a convenience wrapper that discovers supported files and delegates to `extract()`.

#### Synchronous

```python
from suparse import SuparseClient

with SuparseClient() as client:
    result = client.extract_folder(
        "./receipts/",
        template_id="276a0aa8-84bc-4491-a2e7-1ea13381790c",
        split=True,
        cleanup=True,
    )
    for r in result.succeeded:
        print(r.original_file, r.data)
```

#### Asynchronous

```python
import asyncio
from suparse import AsyncSuparseClient

async def main():
    async with AsyncSuparseClient() as client:
        result = await client.extract_folder(
            "./receipts/",
            template_id="276a0aa8-84bc-4491-a2e7-1ea13381790c",
            split=True,
            cleanup=True,
        )
        for r in result.succeeded:
            print(r.original_file, r.data)

asyncio.run(main())
```

> **Note:** `r.data` contains only the extracted document fields. To get all details — including `credits_used`, `template_id`, `document_id`, `file_name`, `page_start`, and `page_end` — use `r.documents` (a list of `DocumentExport` objects). To serialize everything, use `r.model_dump(mode="json")`.

### Error Handling

All exceptions inherit from `SuparseError`. Handle them at whatever granularity makes sense:

#### Synchronous

```python
from suparse import SuparseClient
from suparse.exceptions import (
    SuparseError,
    SuparseAuthError,
    SuparseNetworkError,
    SuparsePollingTimeoutError,
)

with SuparseClient(api_key="your_api_key_here") as client:
    try:
        result = client.extract("invoice.pdf")
    except SuparseAuthError:
        print("Invalid API key")
    except SuparseNetworkError:
        print("Connection failed")
    except SuparsePollingTimeoutError:
        print("Processing timed out")
    except SuparseError as e:
        print(f"Unexpected error: {e}")
```

#### Asynchronous

```python
import asyncio
from suparse import AsyncSuparseClient
from suparse.exceptions import (
    SuparseError,
    SuparseAuthError,
    SuparseNetworkError,
    SuparsePollingTimeoutError,
)

async def main():
    async with AsyncSuparseClient(api_key="your_api_key_here") as client:
        try:
            result = await client.extract("invoice.pdf")
        except SuparseAuthError:
            print("Invalid API key")
        except SuparseNetworkError:
            print("Connection failed")
        except SuparsePollingTimeoutError:
            print("Processing timed out")
        except SuparseError as e:
            print(f"Unexpected error: {e}")

asyncio.run(main())
```

Pass file objects or in-memory streams directly (no `Path` needed):

```python
    with open("invoice.pdf", "rb") as f:
        result = client.extract(f)
```

For batch operations, check `result.failed` to handle per-file errors without catching exceptions:

```python
result = client.extract(["a.pdf", "b.pdf", "c.pdf"])

for r in result.succeeded:
    print(f"OK: {r.original_file} -> {r.document_ids}")

for f in result.failed:
    print(f"FAIL: {f.file} -> {f.error}")
```

### Low-Level API

For cases where you need file-based output or direct control over the upload/poll/download cycle:

#### Synchronous

```python
from pathlib import Path
from suparse import ExportFormat, SuparseClient

with SuparseClient() as client:
    # Upload a file and get back a task ID
    task_id = client.upload_file(
        Path("invoice.pdf"),
        template_id="276a0aa8-84bc-4491-a2e7-1ea13381790c",
        split=False,
        auto_approve=True,  # Set to False to require human review in the Suparse UI
    )

    # Poll until processing completes (returns status + document IDs)
    status, doc_ids = client.poll_task_status(task_id)

    # Export results by document IDs to a JSON file. Omitting the output path
    # uses the API filename or a timestamped fallback in the current directory.
    saved_path = client.download_results(
        ["doc-id-1", "doc-id-2"],
        Path("output.json"),
    )

    # Fetch non-JSON exports in memory
    csv_export = client.fetch_results(
        ["doc-id-1", "doc-id-2"],
        format=ExportFormat.CSV,
    )
    print(csv_export.filename, csv_export.content_type, csv_export.is_zip)

    # Delete documents by ID
    client.delete_documents([
        "550e8400-e29b-41d4-a716-446655440000",
    ])

    # List available templates
    templates = client.list_templates()
    for t in templates:
        print(f"{t.name} ({t.template_language})")
```

#### Asynchronous

```python
import asyncio
from pathlib import Path
from suparse import AsyncSuparseClient, ExcelLayout, ExportFormat, ExportType

async def main():
    async with AsyncSuparseClient() as client:
        # Process a single document to a JSON file
        success = await client.process_document(
            file_path=Path("invoice.pdf"),
            output_path=Path("results.json"),
            cleanup=True,
        )

        # Process multiple files in parallel (returns raw tuples)
        succeeded, failed = await client.process_batch(
            [Path("a.pdf"), Path("b.pdf")],
            template_id="276a0aa8-84bc-4491-a2e7-1ea13381790c",
        )
        # succeeded: list of (Path, task_id, [document_ids])
        # failed: list of (Path, task_id or None, exception)

        # Upload a file and get back a task ID
        task_id = await client.upload_file(
            Path("invoice.pdf"),
            template_id="276a0aa8-84bc-4491-a2e7-1ea13381790c",
            split=False,
            auto_approve=True,  # Set to False to require human review in the Suparse UI
        )

        # Poll until processing completes (returns status + document IDs)
        status, doc_ids = await client.poll_task_status(task_id)

        # Export results by document IDs to a JSON file
        saved_path = await client.download_results(
            ["doc-id-1", "doc-id-2"],
            Path("output.json"),
        )

        # Save an XLSX export to a directory using the API filename if present
        saved_path = await client.download_results(
            ["doc-id-1", "doc-id-2"],
            Path("./exports"),
            format=ExportFormat.XLSX,
            export_type=ExportType.UNIFIED,
            xlsx_layout=ExcelLayout.FLAT,
        )

        # Delete documents by ID
        await client.delete_documents([
            "550e8400-e29b-41d4-a716-446655440000",
        ])

        # List available templates
        templates = await client.list_templates()
        for t in templates:
            print(f"{t.name} ({t.template_language})")

asyncio.run(main())
```

### Constructor Parameters

Both `SuparseClient` and `AsyncSuparseClient` accept the same parameters:

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `api_url` | `str` | `SUPARSE_API_URL` env var or `https://api.suparse.com/api/v1/` | API base URL |
| `api_key` | `str` | `SUPARSE_API_KEY` env var | Your API key (required) |
| `poll_interval` | `int` | `5` | Seconds between polling attempts |
| `max_poll_attempts` | `int` | `300` | Max polling attempts before timeout |

### Result Objects

| Object | Properties | Description |
|--------|-----------|-------------|
| `TaskExport` | `task_id`, `original_file`, `total_documents_extracted`, `documents`, `data`, `document_ids` | Successfully extracted file; `documents` contains `DocumentExport` objects with `credits_used`, `template_id`, etc. |
| `FailedResult` | `file`, `error` | File that failed during extraction |
| `BatchResult` | `succeeded`, `failed`, `total` | Container for batch results; `total` counts inputs while `succeeded` contains returned task export groups |
| `ExportResult` | `format`, `content_type`, `filename`, `is_zip`, `data`, `task_exports`, `google_sheets_export` | Low-level export result for JSON, CSV, XLSX, and Google Sheets |
| `QuickSchema...` | Typed start, processing, completed, and failed result fields | Automatic schema-creation lifecycle |

### Export Formats

| Format | Result | Notes |
|--------|--------|-------|
| `json` | `TaskExport` list | Default structured extraction result |
| `csv` | Binary export bytes | Uses API filename when provided; may be a ZIP |
| `xlsx` | Binary export bytes | Uses API filename when provided; may be a ZIP |
| `google_sheets` | `GoogleSheetsExport` | Requires Google Sheets integration |

`export_type` accepts `unified` or `original` and defaults to `unified`. It affects CSV, XLSX, and Google Sheets exports; JSON remains the task-oriented extraction result.

`xlsx_layout` accepts `ExcelLayout.STANDARD` (the default) or `ExcelLayout.FLAT`. It only affects XLSX exports.

### Exceptions

All exceptions inherit from `SuparseError`.

| Exception | Raised When |
|-----------|-------------|
| `SuparseError` | Base exception for all SDK errors |
| `SuparseNetworkError` | Network connection fails or times out |
| `SuparsePollingTimeoutError` | Polling exceeds `max_poll_attempts` |
| `SuparseProcessingError` | Document fails to process on the server |
| `SuparseAPIError` | Base for HTTP error responses (has `status_code`, `response_body`, `code`, `meta`) |
| `SuparseAuthError` | 401/403 authentication or authorization error |
| `SuparseNotFoundError` | 404 resource not found |
| `SuparseRateLimitError` | 429 too many requests |
| `SuparseServerError` | 5xx server error |
| `SuparseUploadConfirmationUnknownError` | Upload confirmation outcome is unknown; inspect `upload_reference_id` before taking action |
| `SuparseSDKError` | SDK fails to parse the API response into the expected data model |

### `extract()` Parameters

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `files` | `str`, `Path`, `IO[bytes]`, or iterable of these | required | One or more files to process |
| `template_id` | `str` | `None` | Template ID (auto-detect if omitted) |
| `split` | `bool` | `False` | Auto-split multi-page documents |
| `auto_approve` | `bool` | `True` | Set to `False` to require human review in the Suparse UI |
| `pdf_password` | `str` | `None` | Password applied to every input encrypted PDF |
| `folder_id` | `str` | `None` | Folder ID applied to every uploaded input |
| `cleanup` | `bool` | `False` | Delete documents from server after extraction |
| `on_progress` | `callable` | `None` | Called with each `TaskExport` or `FailedResult` as it completes |

### `extract_folder()` Parameters

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `folder` | `str` or `Path` | required | Directory to scan |
| `pattern` | `str` | `"*"` | Glob pattern for file discovery (filtered by supported extensions) |
| `template_id` | `str` | `None` | Template ID (auto-detect if omitted) |
| `split` | `bool` | `False` | Auto-split multi-page documents |
| `auto_approve` | `bool` | `True` | Set to `False` to require human review in the Suparse UI |
| `pdf_password` | `str` | `None` | Password applied to every input encrypted PDF |
| `folder_id` | `str` | `None` | Folder ID applied to every uploaded input |
| `cleanup` | `bool` | `False` | Delete documents from server after extraction |
| `on_progress` | `callable` | `None` | Called with each result as it completes |

## Documentation

Full API documentation is available at [suparse.com/docs](https://suparse.com/docs).

## License

[MIT](LICENSE)

