Metadata-Version: 2.4
Name: keld
Version: 1.2.0
Summary: An OpenAI-compatible client that reports its own usage telemetry to Keld Atlas.
Author-email: Keld <support@keld.co>
Maintainer-email: Keld <support@keld.co>
License: MIT
Project-URL: Homepage, https://keld.co
Project-URL: Repository, https://github.com/ncx-ai/atlas-telemetry-python
Project-URL: Issues, https://github.com/ncx-ai/atlas-telemetry-python/issues
Project-URL: Changelog, https://github.com/ncx-ai/atlas-telemetry-python/blob/main/CHANGELOG.md
Keywords: observability,otlp,opentelemetry,openai,telemetry,llm,finops
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: System :: Monitoring
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: httpx>=0.24
Requires-Dist: openai>=1.30
Requires-Dist: opentelemetry-sdk>=1.44
Requires-Dist: opentelemetry-instrumentation-genai-openai==1.1b0
Provides-Extra: test
Requires-Dist: pytest>=7.4; extra == "test"
Requires-Dist: pytest-asyncio>=0.21; extra == "test"
Requires-Dist: respx>=0.20; extra == "test"
Dynamic: license-file

# keld

An OpenAI-compatible Python client that reports its own usage telemetry to
[Keld Atlas](https://keld.co). `keld.Keld` subclasses `openai.OpenAI`, so `chat.completions`,
`responses`, streaming, retries, every response type and every type hint are the openai SDK's own.

```python
from keld import Keld                                # replaces: from openai import OpenAI

client = Keld(                                       # replaces: client = OpenAI(...)
    base_url="https://api-gateway.keld.co/v1",       # where inference goes
    api_key=os.environ["KELD_TOKEN"],                # that endpoint's key
    atlas_token=os.environ["ATLAS_INGEST_TOKEN"],    # turns telemetry on
)

client.chat.completions.create(...)                  # unchanged from here on
```

`atlas_token` is the only thing that turns telemetry on — omit it and this is a plain inference
client that exports nothing. `base_url` points at anything speaking the OpenAI API: OpenAI,
`https://api.anthropic.com/v1/`, `https://bedrock-runtime.{region}.amazonaws.com/openai/v1`, or
your own gateway.

## Install

```bash
pip install keld
```

Python 3.10 or newer. There are no extras: the client and the instrumentation are both core
dependencies, so no install can import but silently under-report.

## Options

Everything is a keyword argument to `Keld` / `AsyncKeld`. Anything this SDK does not name is
forwarded to `openai.OpenAI` untouched, so the request on the wire is unchanged.

| Argument | Env fallback | What it is |
|---|---|---|
| `api_key` | `OPENAI_API_KEY` | The inference endpoint's key. Never sent to Atlas. |
| `base_url` | `OPENAI_BASE_URL` | Where inference goes. `/v1` is appended when it names only a host. |
| `timeout`, `max_retries`, `default_headers`, `http_client`, `organization`, `project`, … | as the openai SDK defines them | Pass-through. This SDK adds no headers and reads no other environment variable. |
| `atlas_token` | none, deliberately | A Keld Atlas **agent key** (`kagt_...`) from the Atlas Integrations page. Turns telemetry on. `ATLAS_INGEST_TOKEN` in the environment will not switch it on behind your back — pass it explicitly. **This is not an inference API key.** |
| `atlas_endpoint` | `ATLAS_ENDPOINT` | Where telemetry goes. Defaults to `https://atlas.keld.co`, so a token alone is a complete configuration. The SDK appends `/v1/logs` itself. |
| `atlas_metadata` | none, deliberately | `dict[str, str]` of the dimensions Atlas slices your spend by. |

`atlas_metadata` is attached to every call the process reports:

```python
client = Keld(
    base_url=..., api_key=..., atlas_token=...,
    atlas_metadata={
        "environment": "production",         # feeds Atlas's CapEx/OpEx classification
        "repo": "acme/checkout-service",     # the calling application
        "team": "checkout",                  # anything else Atlas slices by
    },
)
```

Keys reach Atlas bare, under exactly the names you give, so a new dimension needs a change at
the Atlas end and no release here. An entry with a non-string or empty key, a non-string value,
or a key the SDK writes itself (`service.name`, `tool`, `sdk.language`, `sdk.version`) is
dropped with a warning naming it; the rest of the mapping still reports, and nothing is raised.

## Sync or async

Two classes, because `openai` has two and they are not interchangeable. Pick the one that
matches the code you already have.

```python
from keld import Keld, AsyncKeld

# Sync — scripts, notebooks, batch jobs, Celery workers
client = Keld(base_url=..., api_key=..., atlas_token=...)
resp = client.chat.completions.create(model=..., messages=[...])

# Async — FastAPI, aiohttp, anything running an event loop
client = AsyncKeld(base_url=..., api_key=..., atlas_token=...)
resp = await client.chat.completions.create(model=..., messages=[...])
```

`AsyncKeld` subclasses `openai.AsyncOpenAI` and takes exactly the same arguments, telemetry
included. Neither is the default: if your handler is `async def`, the sync client blocks the
event loop and stalls every other request in the process while it waits on the model — so the
framework you already chose decides this, not us.

## What the SDK does on its own

- **`/v1` is appended to a `base_url` that names only a host**, so `https://api-gateway.keld.co`
  and `https://api-gateway.keld.co/v1` both work. One that already carries a path is left alone.
- **Streamed `chat.completions` calls get `stream_options={"include_usage": True}`** — the one
  place the SDK modifies an outgoing request. Without it an OpenAI-compatible endpoint omits
  token usage from a streamed response, and the call bills as zero tokens. Your own
  `stream_options` are merged with, never replaced; every chunk, including the usage-only chunk
  with empty `choices`, reaches your code unchanged; non-streaming requests are untouched.
- **`provider` is attributed from the endpoint's host.** OpenAI, Anthropic, Bedrock's per-region
  runtime hosts, OpenRouter, Together, Fireworks, Groq, DeepSeek, xAI, Mistral and any host under
  `.keld.co` are recognised by name; Ollama by its registered port `11434`; Azure by what the
  instrumentation reports. Anything else reports `hosted_vllm` with a one-time warning, and there
  is no override.
- **Telemetry never breaks inference.** Constructing the client never fails because of it, and a
  bug in our extraction or export is caught and logged. Provider errors reach you unchanged, and
  are also recorded as a zero-token event with an `error` attribute so error rates stay visible.
  Export runs on a daemon thread over a bounded queue that drops the oldest entries under
  backpressure, so it never blocks your request thread or on network I/O. Telemetry starts once
  per process: several clients still report each call once.
- **The queue drains when the interpreter exits.** Telemetry registers an `atexit` hook when it
  starts, so a script, a job, a CLI or a notebook cell that makes its calls and ends reports
  them. Without it the flush interval — a couple of seconds, on a daemon thread — would take the
  last batch with the process, silently. The hook waits at most 2 seconds, never raises, and is
  registered only once telemetry actually starts: importing `keld` in a process that passes no
  `atlas_token` leaves nothing behind.
- **An existing global `TracerProvider` is reused**, not replaced, so your own OTel exporters
  keep seeing openai spans. Without one, the SDK runs a private provider.

`keld.shutdown()` is what remains for the exits the hook cannot see: `os._exit()`, `os.execv()`,
a crash, or a signal your app does not handle — the SDK installs no signal handler of its own,
because attaching one would change how your process shuts down. It is also how to get the data
out of a process that stops making calls but keeps running, such as a worker between jobs or an
open notebook kernel. It takes a `timeout` (5 seconds by default, as a total budget), and pairs
safely with the hook: whichever drains first clears the queue and the other returns. That
function is the package's only public name besides the two clients.

## What gets reported

Every call is one OTLP log record carrying `provider`, `model`, `input_tokens`, `output_tokens`,
`cache_read_tokens`, `cache_creation_tokens` and `duration_ms`, plus a `session.id`,
`event.sequence` and `prompt.id` the SDK mints itself; `request_id` when the provider returns
one, and `error` when the call failed. Token counts always come from the provider's own
response — the SDK never estimates them.

`cost_usd` is sent only when the provider itself reports what the call cost — a Keld gateway
returns it as `usage.cost`, streamed or not. The SDK never prices a call: with no reported
cost there is no `cost_usd` attribute, and Atlas prices the call from its price table.

**Never sent:** prompt text, completion text, message bodies, or your API keys — the telemetry
path only reads the response of calls your own client makes.

Reported surfaces are `chat.completions.create`, `chat.completions.parse` and the Responses API
(`create` / `stream` / `retrieve`). Embeddings calls are deliberately not reported.

## Known limits

- **Native `anthropic.Anthropic()` and boto3 `bedrock-runtime` clients are not instrumented.**
  This SDK instruments the OpenAI-compatible surface and nothing else. Reach those providers
  through their OpenAI-compatible endpoints and they report like any other host; an app calling
  the native SDKs gets no telemetry at all.
- **Bedrock's OpenAI-compatible endpoint does not serve Claude.** It serves the OpenAI-format
  models (verified against `openai.gpt-oss-120b-1:0`). Claude on Bedrock is reachable only
  through APIs this SDK does not instrument.
- **A private or self-hosted host reports `hosted_vllm` and prices at $0.** There is no way to
  override the label; open an issue to have a host added to the table.
- **An endpoint that rejects `stream_options` will reject every streamed chat call.** The SDK
  always injects it and there is no flag to turn that off; a flag would trade a broken call for
  silent zero-token billing. Every endpoint we test against accepts it.

## Development

```bash
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[test]"
pytest -v
```

`[test]` is dev-only and the package's only extra. `tests/test_contract.py` pins the OTLP export
payload to `tests/fixtures/atlas-sdk-otlp-export.golden.json`, a copy of the normative fixture
published by the `keld-atlas` backend; if you touch the wire format, keep both copies in sync
rather than fixing the diff locally.

To cut a release: bump `__version__` in `src/keld/_version.py` (the single source of truth —
`pyproject.toml` reads it), move the `Unreleased` notes into a new `CHANGELOG.md` section, merge
to `main`, then `git tag v1.0.0 && git push origin v1.0.0`.

`.github/workflows/release.yml` verifies the tag matches `__version__`, builds an sdist and a
wheel, runs the suite against the built wheel on Python 3.10 through 3.13, and only then
publishes — via [PyPI Trusted Publishing](https://docs.pypi.org/trusted-publishers/), so no API
token lives in this repo. Running the workflow manually publishes to TestPyPI instead.

MIT licensed — see [LICENSE](LICENSE).
