Metadata-Version: 2.5
Name: pdf4me
Version: 1.0.1
Summary: PDF4me PDF suite SDK for Python: convert HTML, URL, Markdown, Word, Excel and images to and from PDF; merge, split, compress, rotate, OCR, PDF/A; e-sign, protect, unlock; create and fill PDF forms; find and replace text; hyperlink annotations, text coordinates, highlighting; Word tracked changes; image resize, crop, watermark, EXIF; barcodes, SwissQR and EPC QR codes; ZUGFeRD e-invoices; document generation; AI extraction from invoices, receipts, contracts, bank statements, tax documents and pay stubs.
Project-URL: Homepage, https://pdf4me.com/
Project-URL: Documentation, https://docs.pdf4me.com/pdf4me-api/
Project-URL: Source, https://github.com/pdf4me/pdf4me-clientapi-python
Project-URL: Bug Reports, https://github.com/pdf4me/pdf4me-clientapi-python/issues
Author-email: Pdf4me <support-dev@pdf4me.com>
License-Expression: MIT
License-File: LICENSE
Keywords: async,asyncio,bank-statement,barcode,compress-pdf,contract,convert,convert-to-pdf,data-extraction,digital-signature,document-ai,document-generation,document-parser,docx,e-invoicing,encrypt-pdf,epc-qr,esign,exif,fill-pdf-form,find-and-replace,html-to-pdf,image-processing,image-to-pdf,invoice,markdown-to-pdf,merge-pdf,ocr,optimize,pdf,pdf-annotations,pdf-api,pdf-forms,pdf-metadata,pdf-sdk,pdf-to-excel,pdf-to-word,pdf4me,pdfa,qrcode,receipt,resize-image,rotate-pdf,split-pdf,stamp,swissqr,tax-document,text-extraction,tracked-changes,unlock-pdf,url-to-pdf,watermark,word-to-pdf,zugferd
Classifier: Development Status :: 5 - Production/Stable
Classifier: Framework :: AsyncIO
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Information Technology
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Multimedia :: Graphics :: Graphics Conversion
Classifier: Topic :: Office/Business
Classifier: Topic :: Office/Business :: Financial :: Accounting
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Text Processing
Classifier: Topic :: Utilities
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: httpx[http2]<1.0.0,>=0.25
Requires-Dist: microsoft-kiota-abstractions>=1.12.0
Requires-Dist: microsoft-kiota-http>=1.12.0
Requires-Dist: microsoft-kiota-serialization-form>=1.12.0
Requires-Dist: microsoft-kiota-serialization-json>=1.12.0
Requires-Dist: microsoft-kiota-serialization-multipart>=1.12.0
Requires-Dist: microsoft-kiota-serialization-text>=1.12.0
Description-Content-Type: text/markdown

# pdf4me

[![PyPI](https://img.shields.io/pypi/v/pdf4me.svg)](https://pypi.org/project/pdf4me/)
[![Python versions](https://img.shields.io/pypi/pyversions/pdf4me.svg)](https://pypi.org/project/pdf4me/)
[![License](https://img.shields.io/pypi/l/pdf4me.svg)](https://github.com/pdf4me/pdf4me-clientapi-python/blob/master/LICENSE)

Python client for the PDF4me API. Convert, optimize, merge, split, stamp, and
extract content from documents with async functions and typed results.

Requires Python 3.11 or newer and a PDF4me API key.

> **Upgrading from 0.8.x?** Version 1.0 is a rewrite and nothing from 0.8.x
> keeps working. See [Migrating from 0.8.x](#migrating-from-08x).

## Installation

```sh
python -m pip install pdf4me
```

To install from a source checkout, run `python -m pip install .` in this
package directory. The import name is `pdf4me`.

## Quick start

Set your API key in the environment:

```sh
export PDF4ME_API_KEY="your-api-key"
```

Save this as `optimize_pdf.py` and put an `input.pdf` in the same working
directory. The example uploads the PDF for optimization and saves the result
as `optimized.pdf`.

```python
import asyncio
import os
from pathlib import Path

from pdf4me import Pdf4meClient
from pdf4me.optimize import optimize


async def main() -> None:
    api_key = os.environ.get("PDF4ME_API_KEY")
    if not api_key:
        raise SystemExit("Set PDF4ME_API_KEY before running this example.")

    source = Path("input.pdf")
    async with Pdf4meClient(api_key) as client:
        result = await optimize(
            client, source.read_bytes(), doc_name=source.name
        )
    Path("optimized.pdf").write_bytes(result)
    print("Saved optimized.pdf")


if __name__ == "__main__":
    asyncio.run(main())
```

Run it with:

```sh
python optimize_pdf.py
```

## Using actions

Import actions from their module. Pass the client first, document content next,
and action options as keyword arguments. Reuse one client for multiple calls;
the async context manager closes its connections when finished.

For example, inside an async function with an open `client`:

```python
from pdf4me.pdf import get_pdf_metadata
from pdf4me.edit import StampAlignX, stamp

metadata = await get_pdf_metadata(client, data, doc_name="input.pdf")
print(metadata.page_count)

stamped = await stamp(client, data, text="DRAFT", align_x=StampAlignX.Center)
Path("stamped.pdf").write_bytes(stamped)
```

Here, `data` contains the PDF bytes, for example `Path("input.pdf").read_bytes()`.
Enums such as `OptimizeProfile` and `StampAlignX` are exported by the action's
module.

Source documents accept `bytes`, base64 strings, `pdf4meblobid://<id>` references,
or HTTPS URLs. Bytes are encoded automatically. File actions return `bytes`;
JSON actions return typed models, such as `DocMetadata` or `SplitPdfRes`.
Queued jobs are polled automatically until the result is available.

Every action accepts `extra`, a dictionary of additional API request fields.
Use the API's field names for these values.

## Client configuration and errors

```python
async with Pdf4meClient(
    api_key,
    base_url="https://api.pdf4me.com",
    timeout=100,
    max_wait=300,
) as client:
    result = await optimize(client, data)
```

Durations are in **seconds**. `timeout` controls the HTTP request timeout;
`max_wait` limits waiting for a queued job.

API failures raise `Pdf4meException`, importable from `pdf4me`, with the status
code, API message, and trace ID. A queued job exceeding `max_wait` raises
`TimeoutError`.

## Available modules

The client provides 106 actions across 18 modules.

| Module                          | Docs section                                                                                         | Actions |
| ------------------------------- | ---------------------------------------------------------------------------------------------------- | ------- |
| `pdf4me.ai_document_extraction` | [AI Document Extraction](https://docs.pdf4me.com/general-guidelines/ai-document-parser-using-parse/) | 14      |
| `pdf4me.barcode`                | [Barcode](https://docs.pdf4me.com/pdf4me-api/barcode/add-barcode-to-pdf/)                            | 7       |
| `pdf4me.convert`                | [Convert](https://docs.pdf4me.com/pdf4me-api/convert/convert-json-to-excel/)                         | 13      |
| `pdf4me.edit`                   | [Edit](https://docs.pdf4me.com/pdf4me-api/edit/add-attachment-to-pdf/)                               | 7       |
| `pdf4me.excel`                  | [Excel](https://docs.pdf4me.com/pdf4me-api/excel/add-image-header-footer/)                           | 1       |
| `pdf4me.extract`                | [Extract](https://docs.pdf4me.com/pdf4me-api/extract/classify-document/)                             | 8       |
| `pdf4me.find_search`            | [Find Search](https://docs.pdf4me.com/pdf4me-api/find-search/find-and-replace-text/)                 | 2       |
| `pdf4me.forms`                  | [Forms](https://docs.pdf4me.com/pdf4me-api/forms/add-form-fields-to-pdf/)                            | 2       |
| `pdf4me.generate`               | [Generate](https://docs.pdf4me.com/pdf4me-api/generate/enable-tracking-changes-in-word/)             | 6       |
| `pdf4me.image`                  | [Image](https://docs.pdf4me.com/pdf4me-api/image/add-image-watermark-to-image/)                      | 13      |
| `pdf4me.merge_split`            | [Merge & Split](https://docs.pdf4me.com/pdf4me-api/merge-split/merge/)                               | 5       |
| `pdf4me.optimize`               | [Optimize](https://docs.pdf4me.com/pdf4me-api/optimize/compress-pdf/)                                | 1       |
| `pdf4me.organize`               | [Organize](https://docs.pdf4me.com/pdf4me-api/organize/delete-blank-pages-from-pdf/)                 | 5       |
| `pdf4me.pdf`                    | [PDF](https://docs.pdf4me.com/pdf4me-api/pdf/get-pdf-information/)                                   | 16      |
| `pdf4me.pdf4me`                 | [PDF4me](https://docs.pdf4me.com/pdf4me-api/pdf4me/update-hyperlinks-annotation/)                    | 1       |
| `pdf4me.security`               | [Security](https://docs.pdf4me.com/pdf4me-api/security/protect-document/)                            | 2       |
| `pdf4me.word`                   | [Word](https://docs.pdf4me.com/pdf4me-api/word/disable-tracking-changes-in-word/)                    | 2       |
| `pdf4me.zugferd`                | ZUGFeRD                                                                                              | 1       |

## Migrating from 0.8.x

Version 1.0 replaces the whole surface of the library. Upgrading will not be a
drop-in change; the four things that break every 0.8.x program are:

- **Calls are async.** Every action is a coroutine and must be awaited from
  inside an event loop.
- **The client classes are gone.** `MergeClient`, `OptimizeClient`,
  `SplitClient` and the rest are replaced by plain functions, imported from the
  module for their API section, that take the client as their first argument.
- **`config.properties` is no longer read.** Pass the API key to
  `Pdf4meClient` yourself. `jprops`, `requests` and `keyring` are no longer
  dependencies.
- **Python 2.7 and 3.6 are no longer supported.** The minimum is 3.11.

Files are passed as `bytes` rather than as the file handles
`FileReader().get_file_handler(...)` used to return, so `Path(...).read_bytes()`
replaces it. Actions that produce a document return `bytes`; actions that
answer with data return a typed model.

| 0.8.x                                               | 1.0                                                    |
| --------------------------------------------------- | ------------------------------------------------------ |
| `Pdf4meClient(token=token)`                         | `Pdf4meClient(api_key)`                                |
| `MergeClient(c).merge_2_pdfs(f1, f2)`               | `await merge(client, [a, b])`                          |
| `OptimizeClient(c).optimize_by_profile(profile, f)` | `await optimize(client, data, profile=...)`            |
| `SplitClient(c).split_by_page_nr(n, f)`             | `await split_pdf(client, data, split_action_number=n)` |
| `ExtractClient(c).extract_pages(page_nrs, f)`       | `await extract_pages(client, data, page_numbers=...)`  |
| `StampClient(c).text_stamp(text, pages, ax, ay, f)` | `await stamp(client, data, text=..., pages=...)`       |
| `BarcodeClient(c).read_barcodes_by_type(t, f)`      | `await read_barcodes(client, data, barcode_type=[t])`  |

The table covers the common calls; the module table above says where to import
everything else from, and each action's docstring links its documentation page.

## Contributing

Issues and pull requests are welcome on
[GitHub](https://github.com/pdf4me/pdf4me-clientapi-python).

`src/pdf4me/_generated/` is generated from the PDF4me OpenAPI description.
Edits made there are overwritten the next time the client is regenerated, so
please open an issue for anything wrong in that tree - a missing field, a wrong
type - rather than patching it. The client and the action modules around it are
hand-written and are the right place for a pull request.

The tests run against a fake transport, so they need no API key and make no
network calls:

```sh
uv run pytest
```

## Support

- Documentation: [docs.pdf4me.com](https://docs.pdf4me.com/pdf4me-api/)
- API keys and account: [dev.pdf4me.com](https://dev.pdf4me.com/)
- Bugs and feature requests:
  [GitHub issues](https://github.com/pdf4me/pdf4me-clientapi-python/issues)
- Email: [support-dev@pdf4me.com](mailto:support-dev@pdf4me.com)

Quote the `trace_id` from a `Pdf4meException` when reporting an API failure.

Looking for the PDF4me online tools rather than the API?
See [pdf4me.com](https://pdf4me.com/).

## License

Released under the [MIT License](https://github.com/pdf4me/pdf4me-clientapi-python/blob/master/LICENSE).
