Metadata-Version: 2.4
Name: llm-api-adapter
Version: 0.9.8
Summary: Lightweight multi-provider LLM API adapter for Python — OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai.
Author: Sergey Inozemtsev
License: MIT License
        
        Copyright (c) 2025 Sergey Inozemtsev
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Repository, https://github.com/Inozem/llm_api_adapter/
Keywords: llm,adapter,lightweight,api,openai,anthropic,google,mistral,xai,grok,qwen,kimi,deepseek,zai
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Typing :: Typed
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: requests>=2.32
Provides-Extra: async
Requires-Dist: httpx>=0.28; extra == "async"
Provides-Extra: httpx
Requires-Dist: httpx>=0.28; extra == "httpx"
Provides-Extra: kimi
Requires-Dist: llm-api-adapter-kimi<0.2.0,>=0.1.0; extra == "kimi"
Provides-Extra: mistral
Requires-Dist: llm-api-adapter-mistral<0.2.0,>=0.1.1; extra == "mistral"
Provides-Extra: qwen
Requires-Dist: llm-api-adapter-qwen<0.2.0,>=0.1.0; extra == "qwen"
Provides-Extra: xai
Requires-Dist: llm-api-adapter-xai<0.2.0,>=0.1.1; extra == "xai"
Provides-Extra: deepseek
Requires-Dist: llm-api-adapter-deepseek<0.2.0,>=0.1.0; extra == "deepseek"
Provides-Extra: zai
Requires-Dist: llm-api-adapter-zai<0.2.0,>=0.1.0; extra == "zai"
Dynamic: license-file

# LLM API Adapter SDK for Python
[![PyPI Downloads](https://static.pepy.tech/personalized-badge/llm-api-adapter?period=total&units=INTERNATIONAL_SYSTEM&left_color=BLACK&right_color=GREEN&left_text=downloads)](https://pepy.tech/projects/llm-api-adapter)
[![Tests](https://github.com/Inozem/llm_api_adapter/actions/workflows/ci-main.yml/badge.svg)](https://github.com/Inozem/llm_api_adapter/actions/workflows/ci-main.yml)
[![Python versions](https://img.shields.io/pypi/pyversions/llm-api-adapter.svg)](https://pypi.org/project/llm-api-adapter/)
[![codecov](https://codecov.io/github/Inozem/llm_api_adapter/branch/dev/graph/badge.svg?token=T83ZPH1F7Z)](https://codecov.io/github/Inozem/llm_api_adapter)

## Overview

**llm-api-adapter** is a minimal, typed adapter for nine organizations: OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai. It provides one provider-neutral contract for messages, tools, structured output, multimodal input, errors, usage, cost, and streaming — with one runtime dependency and no provider SDKs or orchestration framework. Switching organizations means changing two arguments.

**Note:** Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai are installed separately with their optional extras.

Supports Python 3.10–3.14.

## Contents

- [Overview](#overview)
- [Why this library?](#why-this-library)
- [Features](#features)
- [Installation](#installation)
- [Getting Started](#getting-started)
- [Model-specific request compatibility](#model-specific-request-compatibility)
- [Built-in model exceptions](#built-in-model-exceptions)
- [Streaming](#streaming)
- [Async API](#async-api)
- [Optional HTTPX Sync Transport](#optional-httpx-sync-transport)
- [Handling Errors](#handling-errors)
- [Configuration and Management](#configuration-and-management)
- [Example Use Case](#example-use-case)
- [Timeout Support](#timeout-support)
- [Reasoning Support](#reasoning-support)
- [Tool / Function Calling](#tool--function-calling)
- [Structured Output](#structured-output)
- [Vision Input](#vision-input)
- [Document Input](#document-input)
- [Token Usage and Pricing](#token-usage-and-pricing)
- [Logging](#logging)
- [Adoption](#adoption)
- [Related project](#related-project)
- [Development & Testing](#development--testing)
- [License](#license)

## Why this library?

`llm-api-adapter` is for applications that call several supported models directly. `llm-api-adapter` provides typed messages and responses, registered model-specific request rules, normalized errors, and available usage and cost estimates on `ChatResponse`. `llm-api-adapter` does not provide a router, gateway, agent framework, or broad model coverage.

### Which library should you choose?

- **[LiteLLM](https://docs.litellm.ai/docs/):** LiteLLM provides much broader model coverage, an OpenAI-style interface, routing, and a gateway. `llm-api-adapter` focuses on direct calls to a smaller model catalog and does not include routing or a gateway. For the compared releases, the [`llm-api-adapter`](https://pypi.org/project/llm-api-adapter/0.9.7/) universal wheel is 110 kB with one direct base dependency (`requests`); the [`litellm`](https://pypi.org/project/litellm/1.102.1/) Linux x86-64 wheel is 27.4 MB (about 249 times larger) with 14 direct base dependencies. LiteLLM uses the [OpenAI SDK](https://docs.litellm.ai/docs/providers/openai_compatible) to call OpenAI and OpenAI-compatible endpoints, but installs it as a [base dependency](https://pypi.org/pypi/litellm/1.102.1/json) even when you only call another provider; `llm-api-adapter` does not require provider SDKs. Wheel sizes exclude dependencies, direct counts exclude transitive dependencies and extras, and neither figure measures cold-start time. Choose LiteLLM when provider breadth or routing infrastructure matters more than a small direct-call client.
- **[AISuite](https://github.com/andrewyng/aisuite):** AISuite provides an OpenAI-style interface, agents, and MCP integration. Some of its provider integrations rely on vendor SDKs (for example, [Anthropic](https://github.com/andrewyng/aisuite/blob/main/aisuite/providers/anthropic_provider.py) and [Mistral](https://github.com/andrewyng/aisuite/blob/main/aisuite/providers/mistral_provider.py)). `llm-api-adapter` provides its own typed messages and `ChatResponse`, plus registered model-specific request rules and cost fields on the response, without provider SDKs; it does not include agents or MCP integration. Choose AISuite when its OpenAI-shaped interface or agent features are more important than this model-aware direct-call contract.
- **[LangChain](https://docs.langchain.com/oss/python/learn):** LangChain combines chat-model integrations with retrieval/RAG and agent components. Some of its provider integrations rely on vendor SDKs (for example, [OpenAI](https://github.com/langchain-ai/langchain/blob/master/libs/partners/openai/pyproject.toml) and [Anthropic](https://github.com/langchain-ai/langchain/blob/master/libs/partners/anthropic/pyproject.toml)). `llm-api-adapter` calls all nine supported organizations with `requests` and no provider SDKs, but does not implement retrieval or agent execution. Choose LangChain when your application needs those higher-level components.
- **Provider SDK:** A provider's SDK gives direct access to that provider's native features. `llm-api-adapter` gives supported models a shared message, tool, response, error, and `reasoning_level` interface, but does not expose every native feature. Choose the provider SDK when you need an unsupported or newly released native feature.

## Features

- **Synchronous Streaming**: Iterate normalized text with `stream_chat()` across supported organizations. Optionally coalesce it into bounded chunks and observe per-chunk metadata without changing the yielded `str` contract.
- **Optional HTTPX Sync Transport**: Keep `requests` as the default synchronous transport or opt into HTTPX with `transport="httpx"`. See the [HTTPX sync pilot guide](HTTPX_SYNC_PILOT.md).
- **Asynchronous API**: Use `achat()` and `astream_chat()` with the optional `[async]` installation for non-blocking HTTPX requests and awaitable callbacks. See the [Async API guide](ASYNC_API.md).
- **Reasoning Observability**: Opt in to provider-emitted reasoning summaries through `capture_reasoning=True`, `ReasoningEvent`, and the `on_reasoning` callback without mixing reasoning into visible text.
- **Provider-Neutral Messages and Responses**: Use the same typed messages, `ChatResponse`, usage, pricing, parsed output, and tool-call fields regardless of the provider.
- **Vision Input**: Send images alongside text via `ImagePart` as a URL, raw bytes, or data URI when the selected organization supports that form.
- **PDF Documents**: Send PDF URLs or bytes via `DocumentPart`; provider-specific file/document payloads are generated automatically.
- **Tool / Function Calling**: Provider-agnostic tool definitions and normalized tool calls in `ChatResponse.tool_calls`.
- **Portable Structured Output**: Pass a documented portable JSON Schema to `chat()` and get a parsed object in `ChatResponse.parsed_json` across all built-in adapters (the optional Z.ai package does not support portable structured output).
- **Pydantic Integration**: Pass a portable Pydantic model as `response_model` and get a typed instance back in `ChatResponse.parsed_model` — no manual schema writing required.
- **Request Timeouts**: Per-request timeout control via `timeout_s`; raises `LLMAPITimeoutError` on expiry.
- **Flexible Configuration**: `temperature`, `max_tokens`, `top_p`, and other parameters are translated to a verified provider payload; known unsupported fields are omitted predictably.
- **Pricing Registry**: Bundled model tiers contain ordinary input/output rates and verified automatic cache rates; ordinary rates are overridable per instance.

## Installation

To install the SDK, you can use pip:

```bash
pip install llm-api-adapter
```

To use Mistral, install its optional organization package:

```bash
pip install "llm-api-adapter[mistral]"
```

To use xAI, install its optional organization package:

```bash
pip install "llm-api-adapter[xai]"
```

To use Qwen Model Studio, install its optional organization package:

```bash
pip install "llm-api-adapter[qwen]"
```

To use Kimi / Moonshot, install its optional organization package:

```bash
pip install "llm-api-adapter[kimi]"
```

To use DeepSeek, install its optional organization package:

```bash
pip install "llm-api-adapter[deepseek]"
```

To use Z.ai / GLM, install its optional organization package:

```bash
pip install "llm-api-adapter[zai]"
```

The [Mistral package README](packages/organizations/mistral/README.md),
[xAI package README](packages/organizations/xai/README.md),
[Qwen package README](packages/organizations/qwen/README.md),
[Kimi package README](packages/organizations/kimi/README.md),
[DeepSeek package README](packages/organizations/deepseek/README.md), and
[Z.ai package README](packages/organizations/zai/README.md) list their
supported models and organization-specific behaviour. Direct installation of
`llm-api-adapter-mistral`, `llm-api-adapter-xai`,
`llm-api-adapter-qwen`, `llm-api-adapter-kimi`,
`llm-api-adapter-deepseek`, or `llm-api-adapter-zai` remains supported.

**Core baseline and organization packages.** An organization may be included
in Core only when every included model and its adapter implement the complete
provider-neutral baseline: typed multi-turn messages and roles; sync/async chat
and streaming; image and PDF input; tools; portable structured output; common
sampling, output-limit, and timeout parameters; and normalized `ChatResponse`,
errors, usage, and token pricing. The shared conformance suite verifies this
contract.

Passing the baseline makes Core inclusion possible, not automatic. An
organization may remain an optional package to keep the base installation
small and its model updates and releases independent. Package placement alone
does not imply a missing capability; each model's documented capability
profile and tested exceptions describe its actual support.

Optional organization packages own their adapters, registries, and
provider-specific preprocessing. Mistral PDF input is one such extension: the
package calls Mistral OCR before chat because OCR is not part of the selected
chat model, then passes the resulting Markdown to it. That page-based OCR
charge is a `CostLineItem`, not a token cost. See the
[Mistral package README](packages/organizations/mistral/README.md#pdf-input)
for its exact behaviour and pricing.

For non-blocking requests, install the optional `[async]` extra and follow the
[Async API guide](ASYNC_API.md).

To opt into HTTPX for synchronous `chat()` and `stream_chat()` calls, install
the optional `[httpx]` extra and pass `transport="httpx"`. The default remains
`requests`; see the [HTTPX sync pilot guide](HTTPX_SYNC_PILOT.md).

**Note:** You need an API key from each LLM provider you use, including
Mistral, xAI, Qwen, Kimi, DeepSeek, or Z.ai when their optional packages are
installed. Refer to the provider's documentation for API-key instructions.


## Getting Started

### Importing and Setting Up the Adapter

To start using the adapter, you need to import the necessary components:

```python
from llm_api_adapter.models.messages.chat_message import (
    AIMessage, Prompt, UserMessage
)
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter
```

### Sending a Simple Request

The SDK supports three types of messages for interacting with the LLM:

- **Prompt**: Use `Prompt` to set the context or initial prompt for the model.
- **UserMessage**: Use `UserMessage` to send messages from the user during a conversation.
- **AIMessage**: Use `AIMessage` to simulate responses from the assistant during a conversation.

Here is an example of how to send a simple request to the adapter:

```python
messages = [
    UserMessage("Hi! Can you explain how artificial intelligence works?")
]

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)

response = adapter.chat(
    messages=messages,
    max_tokens=max_tokens,
    temperature=temperature,
    top_p=top_p
)
print(response.content)
```

### Parameters

- **max\_tokens**: The maximum number of tokens to generate in the response. This limits the length of the output.

- **temperature**: Controls the randomness of the response. Higher values (e.g., 0.8) make the output more random, while lower values (e.g., 0.2) make it more focused and deterministic. Default value: `1.0` (range: 0 to 2).

- **top\_p**: Limits the response to a certain cumulative probability. This is used to create more focused and coherent responses by considering only the highest probability options. Default value: `1.0` (range: 0 to 1).

### Model-specific request compatibility

The public `chat()`, `achat()`, `stream_chat()`, and `astream_chat()` methods
keep the same arguments across providers. For a verified model, the adapter
applies its registered payload compatibility rules before making the request;
sync and async calls use the same rules.

If a rule omits a parameter, its documented default is removed silently. A
different supplied value is also removed, but emits one `UserWarning` and one
`WARNING` log record per ignored parameter. Treat that warning as a prompt to
remove the setting or select a model that supports it.

Rules match the **exact** model names in the bundled registry, with two narrow
snapshot exceptions for direct provider APIs: Anthropic
`claude-...-YYYYMMDD` and OpenAI `gpt-...-YYYY-MM-DD` IDs inherit a registered
unsuffixed base model's metadata when the date is valid. The requested snapshot
ID is still sent to the provider. The adapter does not infer behavior from a
name prefix; Google aliases and preview IDs, fine-tuned IDs, and any snapshot
whose base model is unregistered receive no compatibility transformation. An
unknown OpenAI model uses Chat Completions rather than the Responses API. Use a
listed model or request a verified registry addition when a provider adds a new
alias or model.

For `gemini-3.8-flash`, `gemini-3.7-flash`, `gemini-3.6-flash`, and
`gemini-3.5-flash-lite`, Google does not support sampling controls. The adapter
omits `temperature` and `top_p` from these model requests; non-default values
produce the compatibility warning described above.

`gpt-6-astra` and `gpt-6.1-sol` use the OpenAI Responses API and do not support
`temperature` or `top_p`; the adapter omits both according to the same warning
policy. Neither model can disable reasoning: `reasoning_level="none"` resolves
to its lowest supported effort, `low`, with a `UserWarning`.

`gpt-6-sol` and `gpt-6-luna` use the Responses API and support
`reasoning_level="none"`. With any higher reasoning effort, OpenAI does not
accept `temperature` or `top_p`; the adapter omits them using the same warning
policy.

### Built-in model exceptions

Each registered model declares a `capability_exceptions` list; an empty list
means no exception. Each entry identifies the affected capability and behavior
and describes the exception in `behavior`. Exact supported values and payload
transformations remain in `reasoning_capability` and `request_rules`.

See the [OpenAI](src/llm_api_adapter/llm_registry/organizations/openai.json),
[Anthropic](src/llm_api_adapter/llm_registry/organizations/anthropic.json), and
[Google](src/llm_api_adapter/llm_registry/organizations/google.json) registries
for current model details. The [project constitution](.specify/memory/constitution.md)
defines the baseline and exception policy; optional organization packages have
their own [package READMEs](#installation).

### Alternative Message Format

In addition to the built-in message classes, the SDK also supports the standard OpenAI-style message format for quick adoption and compatibility:

```python
messages = [
    {"role": "system", "content": "You are a friendly assistant who answers only yes or no."},
    {"role": "user", "content": "Do you know how AI learns?"},
    {"role": "assistant", "content": "Yes."},
    {"role": "user", "content": "Can you explain it in one sentence?"}
]

response = adapter.chat(messages=messages, max_tokens=50)
print(response.content)
```
> **Note**  
> The adapter automatically normalizes message input — you can mix custom message classes and OpenAI-style dicts in one list.

## Streaming

`stream_chat()` is the synchronous streaming counterpart to `chat()`. It always yields normalized visible text as `str`, independently of the provider's native streaming event format.

By default, `buffer_chars=None` preserves provider text-delta behavior. Set a positive `buffer_chars` value when a consumer prefers bounded, coalesced text; `on_chunk` then receives a `StreamChunk` with the same text and observability metadata.

```python
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5.6-sol",
    api_key="...",
)

chunk_metadata = []
for text in adapter.stream_chat(
    messages=[{"role": "user", "content": "Explain SSE in one sentence."}],
    buffer_chars=80,
    on_chunk=chunk_metadata.append,
    on_tool_call=lambda call: print(call.name, call.arguments),
    on_done=lambda response: print(response.usage),
):
    print(text, end="", flush=True)  # `text` is still a str.

for chunk in chunk_metadata:
    print(chunk.index, chunk.elapsed_s, chunk.usage, chunk.output_tokens_delta)
```

- `buffer_chars` accepts `None` (the default) or a positive integer. Buffered chunks never exceed the configured size; any remaining text is emitted during normal completion.
- `on_chunk(chunk)` receives a `StreamChunk` with `text`, monotonic `index`, local `elapsed_s` / `delta_s`, and optional `usage` / `output_tokens_delta` fields.
- `on_delta(text)` is called for every yielded visible text chunk. The order is always `on_chunk` → `on_delta` → `yield`.
- `on_tool_call(tool_call)` receives only completed, normalized `ToolCall` objects after the provider stream finishes.
- `on_done(response)` receives the finalized `ChatResponse`, including usage, pricing, parsed structured output, and any tool calls, after the final buffer flush.
- Token metadata is optional and comes only from provider usage payloads. When a provider reports cumulative output usage, `output_tokens_delta` is the local increment; no token estimation is performed.

Buffering is pull-based and has no background worker or time-based flush. If the stream fails or a caller closes the iterator early, pending text is not emitted as a successful final chunk and `on_done` is not called.

`stream_chat()` uses the synchronous `requests` transport by default. To opt
into HTTPX for both `chat()` and `stream_chat()`, see the
[HTTPX sync pilot guide](HTTPX_SYNC_PILOT.md). For the matching
`astream_chat()` contract and its `httpx.AsyncClient` transport, see the
[Async API guide](ASYNC_API.md#async-streaming).

Provider event references: [OpenAI Responses streaming](https://platform.openai.com/docs/api-reference/responses-streaming), [Anthropic streaming](https://platform.claude.com/docs/en/build-with-claude/streaming), and [Google `streamGenerateContent`](https://ai.google.dev/api/generate-content).

## Async API

The optional asynchronous client, installation command, `achat()` and
`astream_chat()` examples, callback ordering, cancellation behavior, and
reasoning observability are documented in the [Async API guide](ASYNC_API.md).

## Optional HTTPX Sync Transport

The default synchronous transport is `requests` and requires no additional
dependency. HTTPX is available as an opt-in pilot for synchronous `chat()` and
`stream_chat()` calls:

```bash
pip install "llm-api-adapter[httpx]"
```

```python
adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key="...",
    transport="httpx",
)
response = adapter.chat([UserMessage("Say hello")])
```

See the [HTTPX sync pilot guide](HTTPX_SYNC_PILOT.md) for compatibility,
limitations, and E2E verification details.

## Handling Errors

### Common Errors

The SDK provides a set of standardized errors for easier debugging and integration:

### API Errors

- **LLMAPIError**: Base class for all API-related errors. This error is also used for any unexpected LLM API errors.

- **LLMAPIAuthorizationError**: Raised when authentication or authorization fails.

- **LLMAPIRateLimitError**: Raised when rate limits are exceeded.

- **LLMAPITokenLimitError**: Raised when token limits are exceeded.

- **LLMAPIClientError**: Raised when the client makes an invalid request.

- **LLMAPIServerError**: Raised when the server encounters an error.

- **LLMAPITimeoutError**: Raised when a request times out.

- **LLMAPIUsageLimitError**: Raised when usage limits are exceeded.

- **InvalidToolSchemaError**: Raised when a provided tool schema is invalid.

- **InvalidToolArgumentsError**: Raised when tool arguments cannot be parsed or validated.

- **ToolChoiceError**: Raised when `tool_choice` is invalid or references an unknown tool.

- **JSONSchemaError**: Raised for incompatible structured-output arguments, a non-portable or non-recursive schema, an invalid Pydantic response model, invalid JSON, or failed Pydantic response validation.

### Config Errors

- **LLMConfigError**: Raised when the request configuration is invalid or incompatible.

- **LLMReasoningLevelError**: Raised only for Anthropic models when max_tokens is less than reasoning_level.

### Provider Error Mapping

| Exception | OpenAI | Anthropic | Google | Kimi |
|---|---|---|---|---|
| `LLMAPIAuthorizationError` | HTTP 401; `InvalidAuthenticationError`, `AuthenticationError` | HTTP 401; `AuthenticationError`, `PermissionError` | HTTP 401/403; `PERMISSION_DENIED` | HTTP 401/403; authentication or permission error type |
| `LLMAPIRateLimitError` | HTTP 429; `RateLimitError` | HTTP 429; `RateLimitError` | HTTP 429; `RESOURCE_EXHAUSTED` | HTTP 429; `rate_limit_error`, `rate_limit_exceeded` |
| `LLMAPITokenLimitError` | `MaxTokensExceededError`, `TokenLimitError` | — | — | `context_length_exceeded`, `input_too_long`, output-token limit types |
| `LLMAPIClientError` | HTTP 4xx; `InvalidRequestError`, `BadRequestError` | HTTP 4xx; `InvalidRequestError`, `RequestTooLargeError`, `NotFoundError` | HTTP 4xx; `INVALID_ARGUMENT`, `FAILED_PRECONDITION`, `NOT_FOUND` | Unclassified HTTP/client or SSE error |
| `LLMAPIServerError` | HTTP 5xx; `InternalServerError`, `ServiceUnavailableError` | HTTP 5xx; `APIError`, `OverloadedError` | HTTP 5xx; `INTERNAL`, `UNAVAILABLE` | HTTP 5xx; `api_error`, `internal_error`, `overloaded_error` |
| `LLMAPITimeoutError` | `requests.Timeout`, `httpx.TimeoutException`; `TimeoutError` | `requests.Timeout`, `httpx.TimeoutException` | `requests.Timeout`, `httpx.TimeoutException`; `DEADLINE_EXCEEDED` | HTTP 408/504; `timeout`, `timeout_error` |
| `LLMAPIUsageLimitError` | `UsageLimitError`, `QuotaExceededError` | — | — | `insufficient_quota`, `quota_exceeded`, `usage_limit_exceeded` |

> - Kimi maps its documented token-limit and quota error types to `LLMAPITokenLimitError` and `LLMAPIUsageLimitError`. Equivalent cases for Anthropic and Google fall into `LLMAPIClientError` (HTTP 4xx).
> - `LLMAPIClientError` is the default fallback for all unhandled HTTP 4xx responses.
> - `LLMAPIServerError` is the default fallback for all HTTP 5xx responses.
> - `InvalidToolSchemaError`, `InvalidToolArgumentsError`, `ToolChoiceError`, `JSONSchemaError` — client-side errors (validated before the request is sent); inherit from `LLMAPIClientError`.
> - `LLMConfigError`, `LLMReasoningLevelError` — configuration errors (parameter validation before the request); `LLMReasoningLevelError` is only raised for Anthropic models with `budget_tokens`.

## Configuration and Management

### Using Different Providers and Models

The SDK allows you to easily switch between LLM providers and specify the model you want to use. Currently supported providers are OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai. Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai require their corresponding optional extras.

- **OpenAI**: You can use models like `gpt-6-astra`, `gpt-6.1-sol`, `gpt-6-sol`, `gpt-6-luna`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`, `gpt-5.2`, `gpt-5.1`, `gpt-5`, `gpt-5-mini`, `gpt-5-nano`, `gpt-4.1`, `gpt-4.1-mini`, `gpt-4.1-nano`, `gpt-4o`, `gpt-4o-mini`.
- **Anthropic**: Available models include `claude-fable-5-1`, `claude-fable-5`, `claude-opus-5-5`, `claude-sonnet-5-5`, `claude-sonnet-5`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-opus-4-6`, `claude-sonnet-4-6`, `claude-opus-4-5`, `claude-sonnet-4-5`, `claude-haiku-4-5`. For `claude-sonnet-5-5`, tool choice is limited to `auto`/`none`; non-default temperature is omitted with a warning.
- **Google**: Models such as `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.5-flash`, `gemini-3.5-flash-lite`, `gemini-3.1-pro-preview`, `gemini-3.1-flash-lite`, `gemini-3-flash-preview`, `gemini-2.5-pro`, `gemini-2.5-flash`, and `gemini-2.5-flash-lite` can be used.
- **Mistral**: Install with `pip install "llm-api-adapter[mistral]"`. Available models are `mistral-small-2603`, `mistral-medium-3-5`, and `mistral-large-2512`; see the [Mistral package README](packages/organizations/mistral/README.md) for Mistral-specific behaviour.
- **xAI**: Install with `pip install "llm-api-adapter[xai]"`. Fixed model IDs are `grok-4.7`, `grok-4.6`, and `grok-4.5`; see the [xAI package README](packages/organizations/xai/README.md) for its capability matrix and data-handling notes.
- **Qwen**: Install with `pip install "llm-api-adapter[qwen]"`. Fixed model IDs are `qwen3.8-max`, `qwen3.8-flash`, `qwen3.7-plus`, and `qwen3.7-flash`; every operation requires an explicit Frankfurt `workspace_id`. See the [Qwen package README](packages/organizations/qwen/README.md) for its capability boundary, including unsupported PDF input.
- **Kimi**: Install with `pip install "llm-api-adapter[kimi]"`. Fixed model IDs are `kimi-k3` and `kimi-k2.6`; image bytes/data URIs are supported, while public image URLs and all PDF `DocumentPart` forms are rejected before HTTP. See the [Kimi package README](packages/organizations/kimi/README.md) for reasoning, cache-pricing, and data-handling details.
- **DeepSeek**: Install with `pip install "llm-api-adapter[deepseek]"`, or install `llm-api-adapter-deepseek` directly. The package exposes only the verified `deepseek-flash` model through the official Responses API, with text, tools, portable structured output, reasoning, streaming, and image URL/bytes/data-URI input. Documents, generic files, OCR, upload, and conversion are rejected before HTTP. See the [DeepSeek package README](packages/organizations/deepseek/README.md) for its compatibility matrix, continuation privacy, usage/cost boundary, and official references.
- **Z.ai**: Install with `pip install "llm-api-adapter[zai]"`, or install `llm-api-adapter-zai` directly. The package exposes only verified `glm-5.3-flash`, with text, tools, reasoning, streaming, and image URL/bytes/data-URI input. Portable structured output and documents are rejected before HTTP. See the [Z.ai package README](packages/organizations/zai/README.md) for its capability matrix and request boundaries.

Example:

```python
adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)
```

To switch to another provider, simply change the `organization` and `model` parameters.

### Switching Providers

Here is an example of how to switch between different LLM providers using the SDK:

**Note**: Each instance of `UniversalLLMAPIAdapter` is tied to a specific provider and model. You cannot change the `organization` parameter for an existing adapter object. To use a different provider, you must create a new instance.

```python
gpt = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)
gpt_response = gpt.chat(messages=messages)
print(gpt_response.content)

claude = UniversalLLMAPIAdapter(
    organization="anthropic",
    model="claude-sonnet-4-5",
    api_key=anthropic_api_key
)
claude_response = claude.chat(messages=messages)
print(claude_response.content)

google = UniversalLLMAPIAdapter(
    organization="google",
    model="gemini-2.5-flash",
    api_key=google_api_key
)
google_response = google.chat(messages=messages)
print(google_response.content)
```

## Example Use Case

Here is a comprehensive example that showcases all possible message types and interactions:

```python
from llm_api_adapter.models.messages.chat_message import (
    AIMessage, Prompt, UserMessage
)                                               
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

messages = [
    Prompt(
        "You are a friendly assistant who explains complex concepts "
        "in simple terms."
    ),
    UserMessage("Hi! Can you explain how artificial intelligence works?"),
    AIMessage(
        "Sure! Artificial intelligence (AI) is a system that can perform "
        "tasks requiring human-like intelligence, such as recognizing images "
        "or understanding language. It learns by analyzing large amounts of "
        "data, finding patterns, and making predictions."
    ),
    UserMessage("How does AI learn?"),
]

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)

response = adapter.chat(
    messages=messages,
    max_tokens=256,
    temperature=1.0,
    top_p=1.0
)
print(response.content)
```

The `ChatResponse` object returned by `chat` includes:

1. **model**: The model that generated the response.
2. **response\_id**: Unique identifier for the response.
3. **timestamp**: Response generation time.
4. **usage**: Object containing `input_tokens`, `output_tokens`, and `total_tokens`.
5. **currency**: The currency used for cost calculation.
6. **cost\_input**: Cost of input tokens.
7. **cost\_output**: Cost of output tokens.
8. **cost\_total**: Total combined cost.
9. **cost\_breakdown**: Optional non-token `CostLineItem` values supplied by an organization package.
10. **content**: The generated text response.
11. **finish\_reason**: Provider-native reason why generation stopped (for example, `"stop"` or `"length"`).
12. **refusal**: Provider-normalized refusal detail, or `None` when the response was not a refusal.
13. **incomplete\_reason**: Provider-normalized incomplete-generation reason, or `None` when generation completed.
14. **parsed\_json**: Parsed JSON object for a valid completed structured result, otherwise `None`.
15. **parsed\_model**: Typed Pydantic instance for a valid completed `response_model` result, otherwise `None`.

## Timeout Support

The SDK supports per-request timeouts for all providers.

### timeout\_s parameter

`timeout_s` defines the maximum time (in seconds) the SDK will wait for an LLM response. It applies to `chat()`, `achat()`, `stream_chat()`, and `astream_chat()`.

If the timeout is exceeded, the request is aborted and `LLMAPITimeoutError` is raised.

### Example

```python
chat_params = {
    "messages": messages,
    "timeout_s": 2.5
}
gpt = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5.2",
    api_key=openai_api_key
)
response = gpt.chat(**chat_params)
```

### Notes

- Timeout is applied uniformly across all providers.
- The parameter is optional; if omitted, the provider default is used.
- Timeout affects the full request lifecycle (network + model execution).

### Handling timeout errors

Timeouts raise a dedicated exception that can be handled explicitly:

```python
from llm_api_adapter.errors import LLMAPITimeoutError

try:
    response = gpt.chat(**chat_params)
except LLMAPITimeoutError:
    # retry, fallback, or abort
    print("LLM request timed out")
```

## Reasoning Support

This section describes the provider-neutral `reasoning_level` parameter for `chat()`, `achat()`, `stream_chat()`, and `astream_chat()`. It gives application code one input surface; it does not claim that every model exposes the same levels or consumes the same number of reasoning tokens.

```python
response = adapter.chat(
    messages=[UserMessage("Solve this step-by-step")],
    reasoning_level=2048,
)
```

### Default behavior

When `reasoning_level` is omitted, the adapter bypasses model-aware resolution
and preserves the provider request behavior that existed before it. Pass a
level, including `"none"`, when the request must express reasoning intent.

### `reasoning_level` parameter

`reasoning_level` is optional and provider-agnostic. The adapters translate it to the provider's native effort or thinking-budget format.

Supported forms:

- **int** — explicit reasoning-token budget
- **str** — one of the canonical levels: `"none"`, `"minimal"`, `"low"`,
  `"medium"`, `"high"`, or `"very_high"`

For a verified categorical model, an exact provider value listed for that
model is preserved. This permits native values such as `"xhigh"` or `"max"`
where the model supports them, but those values are not portable. Other
canonical strings are projected upward through the model's ordered native
values; `"none"` becomes the model minimum with a warning if it cannot disable
reasoning. An integer is treated as a fraction of the model context window,
clamped to 0–100%, then rounded upward to an available categorical value.
When the native list has no `"none"`, its first value is the first positive
step: `"minimal"` and a numeric percentage at or below that step resolve to it
rather than skipping it.

For a verified numeric-budget model, canonical values from `"minimal"` through
`"very_high"` are evenly interpolated from that model's documented minimum to
maximum budget (with upward rounding). `"none"` resolves to zero when zero is
supported, otherwise to the minimum with a warning. An integer is a literal
budget: a value below the documented minimum falls back to that minimum with a
warning, while a value above the documented maximum is forwarded so the
provider can return its native validation error.

Named levels express relative intent. Do not treat `"medium"` or a numeric
value as an identical token budget across providers or models.

### Usage examples

```python
# Named level
response = adapter.chat(
    messages=[UserMessage("Explain this")],
    reasoning_level="medium",
)

# Explicit numeric level
response = adapter.chat(
    messages=[UserMessage("Solve this step-by-step")],
    reasoning_level=2048,
)

# Explicitly disable reasoning where the model supports it
response = adapter.chat(
    messages=[UserMessage("Simple answer, no reasoning")],
    reasoning_level="none",
)
```

### Provider independence

`reasoning_level` provides the same application-level entry point for all supported providers:

- same parameter name
- no provider-specific reasoning kwargs in application code
- provider-specific translation is handled inside the adapter

The available levels, native mapping, token budget, and whether reasoning can be disabled remain model/provider-dependent. This allows switching between OpenAI, Anthropic, and Google without rewriting the request shape, while keeping provider limitations explicit.

Provider references: [OpenAI reasoning](https://developers.openai.com/api/docs/guides/reasoning), [Anthropic extended thinking](https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking), and [Google thinking](https://ai.google.dev/api/generate-content).

## Tool / Function Calling

The SDK provides a unified provider‑agnostic tool calling interface.

Tools are defined using `ToolSpec`, and tool calls are returned in normalized form through `ChatResponse.tool_calls`.

The adapter **does not execute tools**. Tool execution must be implemented by the caller.

### ToolSpec

```python
from llm_api_adapter.models.tools import ToolSpec

tool = ToolSpec(
    name="get_weather",
    description="Get current weather for a city",
    json_schema={
        "type": "object",
        "properties": {
            "city": {"type": "string"}
        },
        "required": ["city"],
        "additionalProperties": False
    }
)
```

### Tool parameters

All request methods support the same tool parameters:

- `tools`
- `tool_choice`
- `parallel_tool_calls`

Tool calls are returned as normalized `ToolCall` objects. The adapter does not
execute them, so the caller appends each resulting `ToolMessage` and makes the
follow-up request itself.

`parallel_tool_calls` is provider-dependent: Anthropic can enable or
disable parallel tool use, OpenAI supports it for Chat Completions, and Google
ignores it because Gemini has no equivalent option.

### Tool round‑trip example

```python
import json
from typing import Any, Dict

from llm_api_adapter.models.messages.chat_message import (
    Prompt,
    UserMessage,
    AIMessage,
    ToolMessage,
)
from llm_api_adapter.models.tools import ToolSpec
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter


tools = [
    ToolSpec(
        name="get_weather",
        description="Get current weather for a city",
        json_schema={
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
    )
]


def run_tool(name: str, args: Dict[str, Any]) -> Dict[str, Any]:
    if name == "get_weather":
        return {"city": args["city"], "temperature": 22, "unit": "C"}
    raise ValueError(f"Unknown tool: {name}")


adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5.2",
    api_key=openai_api_key,
)

messages = [
    Prompt("If the user asks about weather, call get_weather."),
    UserMessage("What's the weather in Tel Aviv today?")
]

first = adapter.chat(
    messages=messages,
    tools=tools,
    tool_choice="auto",
    max_tokens=1000,
)

if first.tool_calls:
    messages.append(AIMessage(content="", tool_calls=first.tool_calls))

    for tc in first.tool_calls:
        result = run_tool(tc.name, tc.arguments)
        messages.append(
            ToolMessage(
                tool_call_id=tc.call_id,
                content=json.dumps(result)
            )
        )

    final = adapter.chat(
        messages=messages,
        previous_response=first,
        max_tokens=1000
    )
    print(final.content)
```

For a complete async tool round-trip, see the [Async API guide](ASYNC_API.md#tool-calling).

### `previous_response` parameter

`previous_response` accepts the `ChatResponse` returned by an earlier `chat()`
or `achat()` call, and can also be passed to the streaming methods.

For OpenAI models that use the Responses API (o-series and newer GPT models), the adapter extracts `response_id` from the previous response and passes it to the API as `previous_response_id`. This enables stateful multi-turn conversations where the model retains context server-side, which reduces the tokens you need to send in subsequent turns.

For Anthropic, Google, and xAI, the parameter is accepted but ignored — context is carried entirely through the `messages` list regardless. The xAI package also does not expose its provider-specific `store` option; see its README for data-retention details.

DeepSeek is also stateless and does not send `previous_response_id`. Its
package uses `previous_response` only to carry matching opaque reasoning replay
metadata while the caller continues to send the complete `messages` history.
That metadata is not rendered, serialized, represented, or logged; visible
reasoning remains opt-in through `capture_reasoning=True`. See the
[DeepSeek continuation and privacy contract](packages/organizations/deepseek/README.md#history-and-continuation-privacy).

If you omit `previous_response`, the call works normally; you just won't get the stateful-session benefit on OpenAI Responses API models.

## Structured Output

The SDK supports structured output from `chat()`, `achat()`, `stream_chat()`,
and `astream_chat()` through a raw `json_schema` dict or a Pydantic model passed
as `response_model`. The documented Core portable profile below is guaranteed
across OpenAI, Anthropic, Google, Mistral, and xAI. Its explicit boundaries
define the schema vocabulary that remains portable between organizations.

For streaming methods, parsed fields and terminal structured-output state are
available on the finalized `ChatResponse` passed to `on_done`.

### Core portable JSON Schema profile

The profile is the provider-neutral intersection enforced before an HTTP
request. A portable schema has:

- an object as its root;
- only `string`, `number`, `integer`, `boolean`, `object`, and `array` types;
  nullable fields are expressed only as `[<one type>, "null"]`;
- `additionalProperties: false` on every object, and a `required` list that
  contains every object property exactly once; use a nullable property instead
  of an optional one;
- a single object schema in `items` for every array;
- a non-empty `enum` when `enum` is used; and
- only the profile's structural and metadata keywords: `type`, `properties`,
  `required`, `additionalProperties`, `items`, `enum`, `title`, `description`,
  `$id`, and `$schema`.

Direct local references of the form `#/$defs/<name>` (including
Pydantic-generated references) are inlined locally before this profile is
checked. External, unresolved, recursive, or
semantically combined `$ref` values are rejected with `JSONSchemaError` before
a provider request. Other JSON Schema composition, validation, and conditional
keywords are intentionally outside the portable profile.

OpenAI, Anthropic, and Google enforce this Core profile. The Mistral, xAI,
Qwen, and Kimi organization packages enforce the same boundary; xAI
additionally rejects the documented xAI-only invalid constructs before sending
its request.

### Pydantic Integration (`response_model`)

Pass a Pydantic `BaseModel` subclass as `response_model` — the adapter extracts
the JSON Schema automatically and returns a typed instance in
`ChatResponse.parsed_model`.

```python
from pydantic import BaseModel, ConfigDict
from llm_api_adapter.models.messages.chat_message import Prompt, UserMessage
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

class Person(BaseModel):
    model_config = ConfigDict(extra="forbid")

    name: str
    age: int

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key,
)

response = adapter.chat(
    messages=[
        Prompt("Extract structured data from the user's message."),
        UserMessage("My name is Alice and I'm 30 years old."),
    ],
    response_model=Person,
    max_tokens=200,
)

print(response.parsed_model)  # Person(name='Alice', age=30)
print(response.parsed_json)   # {"name": "Alice", "age": 30}
```

> **Note:** Pydantic is not a required dependency of this package. Install it separately: `pip install pydantic`. Every nested Pydantic model must also use `ConfigDict(extra="forbid")` so its generated object schema satisfies the portable profile.

`response_model` cannot be combined with `json_schema` or `tools` — passing both raises `JSONSchemaError`.

### Raw JSON Schema (`json_schema`)

The SDK also supports portable structured JSON output via a `json_schema`
parameter in `chat()`, `achat()`, `stream_chat()`, and `astream_chat()`. The
adapter preserves the schema's semantics: it validates the profile rather than
silently adding required fields or changing `additionalProperties`.

```python
from llm_api_adapter.models.messages.chat_message import Prompt, UserMessage
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

schema = {
    "type": "object",
    "properties": {
        "name": {"type": "string"},
        "age": {"type": "integer"},
    },
    "required": ["name", "age"],
    "additionalProperties": False,
}

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key,
)

response = adapter.chat(
    messages=[
        Prompt("Extract structured data from the user's message."),
        UserMessage("My name is Alice and I'm 30 years old."),
    ],
    json_schema=schema,
    max_tokens=200,
)

print(response.content)      # '{"name": "Alice", "age": 30}'
print(response.parsed_json)  # {"name": "Alice", "age": 30}
```

### `json_schema` parameter

- Accepts a `dict` containing a JSON Schema object.
- Cannot be combined with `tools` or `response_model` — passing both raises `JSONSchemaError`.
- Must satisfy the Core portable profile; unsupported, external, or recursive references raise `JSONSchemaError` before the request.
- Raw schemas are parsed as JSON only. The SDK does not run a local JSON Schema validator against the response.
- If a completed, non-refusal response is not valid JSON, `JSONSchemaError` is raised.

### `parsed_json` field

`ChatResponse.parsed_json` contains the parsed `dict` for a valid completed
structured result. It remains `None` for a refusal or incomplete generation,
and `parsed_model` remains `None` in those cases as well.

### Provider behavior

| Provider | Implementation |
|----------|----------------|
| **OpenAI** (standard) | Native `response_format.type=json_schema` with `strict=true` |
| **OpenAI** (Responses API / o-series) | Native `text.format.type=json_schema` with `strict=true` |
| **Anthropic** | Native `output_config.format.type=json_schema` |
| **Google** | `generationConfig.responseMimeType="application/json"` + `responseJsonSchema` |
| **Mistral** | Native `response_format.type=json_schema` with `strict=true` (optional organization package) |
| **xAI** | Responses `text.format.type=json_schema` with `strict=true` plus xAI's additive local overlay (optional organization package) |
| **Qwen** | Messages `output_config.format.type=json_schema` (optional organization package) |

The same **portable** schema can be reused across these organizations without
provider-specific request code. Google sends the portable JSON Schema through
`responseJsonSchema` and drops only `$id` and `$schema`, which are
non-semantic metadata; no transformer silently drops a semantic constraint.

### Refusal, incomplete, and invalid results

A provider refusal is exposed as `response.refusal`; an incomplete generation
is exposed as `response.incomplete_reason`. Neither is parsed or Pydantic
validated, so both parsed fields remain `None`. A completed non-refusal result
that is invalid JSON, or fails `response_model.model_validate()`, raises
`JSONSchemaError`. Check the terminal fields before using structured output:

```python
response = adapter.chat(
    messages=[UserMessage("Give me the data.")],
    json_schema=schema,
    max_tokens=200,
)
if response.refusal is not None:
    handle_refusal(response.refusal)
elif response.incomplete_reason is not None:
    handle_incomplete(response.incomplete_reason)
else:
    data = response.parsed_json
```

### Error handling

```python
from llm_api_adapter.errors import JSONSchemaError

try:
    response = adapter.chat(
        messages=[UserMessage("Give me the data.")],
        json_schema=schema,
        max_tokens=200,
    )
except JSONSchemaError as e:
    print(f"JSON schema error: {e}")
```

## Vision Input

The SDK supports sending images alongside text using `ImagePart` and the `files` parameter on `UserMessage`. The input contract is shared across OpenAI, Anthropic, and Google; wire-format differences and provider-specific limitations are handled by the adapter.

### Import

```python
from llm_api_adapter.models.messages.chat_message import UserMessage
from llm_api_adapter.models.messages.file_parts import ImagePart
```

### Image from URL

```python
msg = UserMessage(
    "What is in this image?",
    files=[ImagePart(url="https://example.com/photo.jpg")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)
```

### Image from bytes

```python
with open("photo.png", "rb") as f:
    image_bytes = f.read()

msg = UserMessage(
    "Describe this image.",
    files=[ImagePart(data=image_bytes, media_type="image/png")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)
```

### Multiple images

```python
msg = UserMessage(
    "Compare these two images.",
    files=[
        ImagePart(url="https://example.com/before.jpg"),
        ImagePart(url="https://example.com/after.jpg"),
    ]
)
```

### Supported formats

`ImagePart` accepts any image MIME type starting with `image/` — `image/jpeg`, `image/png`, `image/gif`, `image/webp`, `image/heic`, `image/heif`.

When passing a URL, `media_type` is auto-detected from the file extension. For URLs without an extension, pass `media_type` explicitly:

```python
ImagePart(url="https://api.example.com/image?id=42", media_type="image/jpeg")
```

### OpenAI-style dict compatibility

The adapter also normalizes OpenAI-style content lists, so existing code that uses dicts works without changes:

```python
messages = [{
    "role": "user",
    "content": [
        {"type": "text", "text": "What is this?"},
        {"type": "image_url", "image_url": {"url": "https://example.com/img.jpg"}},
    ]
}]
response = adapter.chat(messages=messages, max_tokens=200)
```

Kimi accepts image bytes and data URIs for its supported models, but
does not fetch public image URLs. `ImagePart(url=...)` is rejected before HTTP;
use `ImagePart(data=..., media_type="image/...")` instead. See the
[Kimi package README](packages/organizations/kimi/README.md#history-images-files-and-data-handling).

DeepSeek `deepseek-flash` accepts user-message image URLs, bytes, and base64
data URIs for JPEG, PNG, GIF, and WebP images. Its adapter rejects unsupported
image forms before HTTP and enforces the provider's URL, inline-size, and
per-request image-count limits. See the [DeepSeek image and file boundary](packages/organizations/deepseek/README.md#images-and-the-file-boundary)
and the [official Vision guide](https://api-docs.deepseek.com/guides/vision/).

> **Note:** Google already supports audio input, but `AudioPart` is postponed because Anthropic does not support audio and OpenAI uses a separate audio API, so there is no common provider-neutral contract yet.

## Document Input

The SDK supports PDF documents alongside text using `DocumentPart` and the `files` parameter on `UserMessage` when the selected organization supports document input. Provider-specific wire formats are handled automatically.

### Import

```python
from llm_api_adapter.models.messages.chat_message import UserMessage
from llm_api_adapter.models.messages.file_parts import DocumentPart
```

### PDF from URL

```python
msg = UserMessage(
    "Summarize this document in one sentence.",
    files=[DocumentPart(url="https://example.com/report.pdf")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)
```

The PDF MIME type is detected from the `.pdf` extension. Anthropic, Google, and the OpenAI Responses API accept a document URL. OpenAI Chat Completions models below `gpt-5` do not accept a document URL; use bytes instead.

### PDF from bytes

```python
with open("report.pdf", "rb") as f:
    pdf_bytes = f.read()

msg = UserMessage(
    "Summarize this document in one sentence.",
    files=[DocumentPart(data=pdf_bytes, media_type="application/pdf")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)
```

For bytes, the adapter sends the PDF as base64 data in the provider-specific request format. The same `DocumentPart` works with Anthropic, Google, OpenAI Chat Completions, and the OpenAI Responses API.

### File type support

| File type | Anthropic | OpenAI (< gpt-5) | OpenAI (gpt-5+) | Google | Qwen | Kimi | DeepSeek |
|-----------|-----------|------------------|-----------------|--------|------|------|----------|
| ImagePart (URL) | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ |
| ImagePart (bytes) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| DocumentPart (URL) | ✅ | ❌ | ✅ | ✅ | ❌ | ❌ | ❌ |
| DocumentPart (bytes) | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | ❌ |

Qwen supports images, but rejects every `DocumentPart` URL or byte before
HTTP: PDF and OCR input are outside its package contract. See the
[Qwen package README](packages/organizations/qwen/README.md#pdf-input).

Kimi also rejects every `DocumentPart` URL or byte before HTTP. Kimi's
Files API exposes extracted text rather than a Chat Completions attachment, so
it cannot meet the same bytes-and-URL contract without hidden URL retrieval.
The adapter does not upload or delete files for Kimi. See the
[Kimi package README](packages/organizations/kimi/README.md#history-images-files-and-data-handling).

DeepSeek rejects every `DocumentPart`, generic non-image `FilePart`,
OCR/upload/conversion route, and unsupported image form before either client
is called. It does not fetch document URLs, upload files, process PDFs locally,
or silently fall back to another endpoint or model. See the [DeepSeek package
file boundary](packages/organizations/deepseek/README.md#images-and-the-file-boundary).

## Token Usage and Pricing

`achat()` returns the same `usage`, `currency`, and cost fields as `chat()`.
For `astream_chat()`, the finalized response passed to `on_done` contains the
same fields, while `StreamChunk.usage` is populated when the provider reports
usage during streaming. `cost_input` and `cost_output` are reserved for token
costs; non-token components use `cost_breakdown`.

### Provider-reported usage

The adapter exposes provider-reported counts without estimating missing tokens.
If no usage is reported, `response.usage` is `None`. A partial usage object may
contain `None` in any unreported count; an explicitly reported zero stays `0`.

| `Usage` field | Meaning |
| --- | --- |
| `input_tokens` | Total reported input, including confirmed cache-read and cache-write subsets. |
| `output_tokens` | Reported generated tokens, including reasoning tokens where the provider exposes the required counts. |
| `total_tokens` | Reported total, or a sum of confirmed input and output counts when the provider format requires it. |
| `cached_tokens` | Confirmed input tokens read from cache. `None` means no valid count was reported. |
| `cache_write_tokens` | Confirmed input tokens charged as a separate cache write. `None` means no valid count was reported. |

These are counts, not the cached text or a list of token IDs. Output tokens are
priced at the output rate; there is no separate output-cache component.

### Tiered estimates and automatic caching

Costs are a standard-rate estimate, not a provider invoice. When a provider
reports a valid input count, the adapter selects the first pricing tier whose
`up_to_prompt_tokens` is greater than or equal to `usage.input_tokens`. The
boundary is inclusive; the final tier has no boundary. Total input, including
cache subsets, selects the tier. The selected rates apply to the entire
request, rather than pricing successive portions in different bands.

- `context_window_tokens` is the combined input and generated-output capacity.
- `max_output_tokens` is the generated-output capacity.
- A pricing-tier boundary only selects a price; it is not a request limit.
- Missing input prevents automatic tier selection. Missing usage leaves token
  cost fields unset; separately metered operations can still have entries in
  `cost_breakdown`.

An exact-model `pricing_tiers` entry may additionally contain
`cache_read_input_per_1m` and `cache_write_input_per_1m`. Each is an independent,
verified rate: zero is valid, and an absent rate has no fallback to the ordinary
input rate or the other cache rate. With complete counts and rates converted
to a per-token basis:

```text
ordinary_input = input_tokens - cached_tokens - cache_write_tokens
cost_input = ordinary_input * ordinary_rate
           + cached_tokens * cache_read_rate
           + cache_write_tokens * cache_write_rate
cost_output = output_tokens * output_rate
```

The formula applies to the cache components the provider meters; it does not
require a separate counter for a component the model does not meter. Cache
counts must be nonnegative integers with a sum no greater than total input.

| Built-in organization | Automatic cache usage and registered pricing |
| --- | --- |
| OpenAI | Responses reads `input_tokens_details.cached_tokens` and `cache_write_tokens`; Chat Completions reads the same fields from `prompt_tokens_details`. All registered models have a cache-read rate. `gpt-6-astra`, `gpt-6.1-sol`, `gpt-6-sol`, `gpt-6-luna`, `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` also have a separately priced automatic cache-write component. |
| Google | `usageMetadata.cachedContentTokenCount` becomes `cached_tokens`, a subset of `promptTokenCount`. Automatic reads have registered rates except for `gemini-3-flash-preview`, whose cache-read rate is unverified. There is no separately priced automatic cache-write component. |
| Anthropic | The adapter does not enable opt-in prompt caching, and the registry has no automatic cache-read or cache-write rates. Ordinary input/output pricing applies to its supported requests. |

The same accounting consumes automatic cache usage from installed organization
packages. See their [package READMEs](#installation) for model-specific usage
formats and rate schedules.

Only cache components returned for ordinary requests without opt-in caching
are covered. Explicit cache resources, cache-control settings, selectable TTLs,
and storage charges are outside this accounting. The adapter does not enable
those modes. Batch, flex, priority, modality-specific, provider-hosted tool,
and negotiated-volume charges are also outside the bundled token estimate.

### Incomplete costs

- If a registered cache rate requires a split that the provider omits, or the
  split is invalid or exceeds input, `cost_input` and `cost_total` are `None`.
- A positive reported cache component with no verified rate also leaves
  `cost_input` and `cost_total` unknown. For example, a cache hit on
  `gemini-3-flash-preview` cannot be priced from its bundled tier.
- A confirmed zero cache count incurs no cache charge, even without a cache
  rate. An unreported count remains `None`; it is not rewritten as zero.
- `cost_output` can remain known when only input accounting is incomplete.
  Missing output leaves `cost_output` and `cost_total` unknown.
- `cost_total` is available only when input, output, and any incurred non-token
  operations are fully priced in the same currency. `cost_breakdown` retains
  known non-token items even when the total is unknown.

An unknown cost is not a free request. Callers can use the reported usage for
their own accounting, but the adapter does not fabricate a missing split or
substitute the ordinary price for an unknown cache rate.

### Token Usage and Pricing Example

Provider-parsed `Usage.input_tokens`, `output_tokens`, and `total_tokens` may
be `None` even when `response.usage` exists. Check each count before
arithmetic, and check costs before formatting or adding them to a budget:

```python
google = UniversalLLMAPIAdapter(
    organization="google",
    model="gemini-2.5-flash",
    api_key=google_api_key
)

response = google.chat(**chat_params)

usage = response.usage
if usage is None:
    print("Provider did not report usage.")
else:
    if usage.input_tokens is not None and usage.output_tokens is not None:
        print("Input plus output:", usage.input_tokens + usage.output_tokens)
    if usage.total_tokens is not None:
        print("Total tokens:", usage.total_tokens)
    print("Cache reads:", usage.cached_tokens)
    print("Cache writes:", usage.cache_write_tokens)

if response.cost_total is not None:
    print(f"Estimated total: {response.cost_total:.6f} {response.currency}")
else:
    print("Complete cost is unavailable.")
```

Avoid `count or 0` or `cost or 0` when an unknown value must stay visible in
your accounting.

### Overriding Pricing or Currency

`set_in_per_1m` and `set_out_per_1m` replace the ordinary input or output rate in
every pricing tier for the selected model. They preserve cache rates and do
not change registered metered-operation rates. `set_currency` changes the
currency label; it does not perform a currency conversion. Bundled rates are
maintained in the registry and can become outdated as providers change prices.

```python
google = UniversalLLMAPIAdapter(
    organization="google",
    model="gemini-2.5-flash",
    api_key=google_api_key
)

google.pricing.set_in_per_1m(1.5)
google.pricing.set_out_per_1m(3)
google.pricing.set_currency("EUR")

response = google.chat(**chat_params)
print(response.content)
if response.usage is not None:
    print(response.usage.input_tokens, "tokens", f"({response.cost_input} {response.currency})")
    print(response.usage.output_tokens, "tokens", f"({response.cost_output} {response.currency})")
    print(response.usage.total_tokens, "tokens", f"({response.cost_total} {response.currency})")
```

## Logging

The library uses Python's standard `logging` module and does not configure handlers.
Loggers are module-based under `llm_api_adapter.*` (e.g., `llm_api_adapter.universal_adapter`).

* **Default behavior:** No handlers installed, effective level = `WARNING`.
* **No secrets are logged** — API keys and request bodies are excluded. Only event metadata and errors are logged. Adapter and client objects mask the key in `__repr__` (e.g. `api_key='sk-12345...cdef'`), so they are safe to log or print.

### Enable logs (console)

```python
import logging

logging.basicConfig(level=logging.INFO)  # or DEBUG
# Optionally limit logging to this library
logging.getLogger("llm_api_adapter").setLevel(logging.DEBUG)
```

### Write logs to a file

```python
import logging

handler = logging.FileHandler("llm_api_adapter.log")
handler.setFormatter(logging.Formatter(
    "%(asctime)s %(levelname)s %(name)s %(message)s"
))
root = logging.getLogger()
root.setLevel(logging.INFO)
root.addHandler(handler)
```

### Per-request correlation (optional)

```python
import logging
logger = logging.getLogger("llm_api_adapter")

req_id = "req-123"
logger = logging.LoggerAdapter(logger, {"request_id": req_id})
logger.info("starting call")
```

To include `request_id` in log output, add `%(request_id)s` to your log formatter.

### Reduce noise / silence logs

```python
import logging
logging.getLogger("llm_api_adapter").setLevel(logging.WARNING)   # silence info logs
logging.getLogger("urllib3").setLevel(logging.WARNING)           # if using requests
```

### Env-based log level toggle

```python
# app.py
import logging, os
level = os.getenv("LLM_ADAPTER_LOGLEVEL", "WARNING").upper()
logging.getLogger("llm_api_adapter").setLevel(level)
```

**Tip:** For HTTP-level debugging with `requests`, also set:

```python
import http.client as http_client, logging
http_client.HTTPConnection.debuglevel = 1
logging.getLogger("urllib3").setLevel(logging.DEBUG)
```

Use this only in development.

## Adoption

[Vertec](https://www.vertec.com/en-ch/kb/llm-client/) bundles llm-api-adapter with Vertec Cloud Suite and On-Premises, allowing customers to call OpenAI, Anthropic, and Google models from custom Python extensions through a unified interface.

## Related project

For production applications that need retries, multi-provider failover,
circuit breakers, and tool-session recovery, see
[`llm-api-resilience`](https://github.com/Inozem/llm_api_resilience).

Install it with:

```bash
python -m pip install llm-api-resilience
```

`llm-api-resilience` is built on top of this adapter and keeps its
provider-neutral interface unchanged.

## Development & Testing

Developer setup, deterministic unit and mocked-integration commands, paid E2E requirements, provider-key safety, documentation rules, and the release flow are documented in [CONTRIBUTING.md](CONTRIBUTING.md).

## License

This project is licensed under the terms of the MIT License.  
See the [LICENSE](LICENSE) file for details.
