Metadata-Version: 2.4
Name: evalrouter
Version: 0.4.0a7
Summary: Interactive evaluation workspace, command-line tools, and Python client for EvalRouter
Project-URL: Homepage, https://evalrouter.ai
Project-URL: Documentation, https://evalrouter.ai/developers/docs/cli
Project-URL: Model API, https://evalrouter.ai/developers/docs/model-api
License-Expression: LicenseRef-Kimpton-EvalRouter-SDK
License-File: LICENSE
Keywords: benchmarks,cli,evalrouter,llm,model-evaluation
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Typing :: Typed
Requires-Python: <3.14,>=3.12
Requires-Dist: httpx==0.28.1
Requires-Dist: rich==15.0.0
Requires-Dist: textual==8.2.8
Provides-Extra: benchmark
Requires-Dist: pydantic==2.13.5; extra == 'benchmark'
Provides-Extra: keyring
Requires-Dist: keyring==25.7.0; extra == 'keyring'
Provides-Extra: langsmith
Requires-Dist: langsmith==0.12.6; extra == 'langsmith'
Provides-Extra: strands
Requires-Dist: opentelemetry-api==1.45.0; extra == 'strands'
Requires-Dist: opentelemetry-sdk==1.45.0; extra == 'strands'
Requires-Dist: strands-agents==1.57.2; extra == 'strands'
Provides-Extra: tracing
Requires-Dist: opentelemetry-api==1.45.0; extra == 'tracing'
Requires-Dist: opentelemetry-sdk==1.45.0; extra == 'tracing'
Description-Content-Type: text/markdown

# EvalRouter CLI and Python client

Evaluate models in a full-screen terminal workspace or from scripts. Browse
benchmarks, choose models, optionally set a spending cap, follow runs, and
compare or export their results. The same package includes the typed Python
client, optional trace capture, and benchmark authoring tools.

Requires Python 3.12 or 3.13. Distributed under the proprietary Kimpton EvalRouter
SDK License, included as `LICENSE` in the package.

## Install and create your account

We recommend [uv](https://docs.astral.sh/uv/getting-started/installation/) to install the CLI in its own Python environment:

```sh
uv tool install --prerelease allow evalrouter
evalrouter
```

For this alpha, use `python -m pip install --pre evalrouter` in a Python 3.12 or 3.13 environment.
uv selects a compatible Python for the CLI and downloads one if needed. If
your shell cannot find `evalrouter`, run `uv tool update-shell` and restart your
terminal. Upgrade an existing installation with `evalrouter upgrade` (or
`evalrouter --upgrade`). Alpha, beta and rc installations include prereleases
automatically; stable installations stay on stable unless you pass `--pre`.
The command uses uv, pipx or pip in the current virtual environment. Source
installs, system Python and installer-pinned versions need their original installer.
Check the installed version with `evalrouter --version`. Use `evalrouter --help`
for explicit commands, or see the [CLI guide](https://evalrouter.ai/developers/docs/cli).

The interactive workspace checks for newer releases in the background and shows
an update notice when one is available. The check never installs an update or
delays opening the workspace. Set `EVALROUTER_NO_UPDATE_CHECK=1` to disable it.
Scripts, editable installs and explicitly pinned installations skip update checks.

### Complete command names with Tab

For zsh, add this line to `~/.zshrc` after your `compinit` setup, then start a new
terminal:

```sh
eval "$(evalrouter completion zsh)"
```

For bash, add this line to `~/.bashrc`:

```sh
eval "$(evalrouter completion bash)"
```

Type `evalrouter ru` and press Tab to complete `run`. With `evalrouter r`, Tab
shows matching commands such as `run`, `resume` and `results`. The shell setup
only completes top-level command names and makes no API request. It does not
change interactive prompts or run an evaluation.

This package was previously published as `kimpton-evalrouter-sdk`. The command
stays `evalrouter`. Use `from evalrouter import Client` for new integrations;
legacy `kimpton_evalrouter` imports remain compatible and refer to the same
public classes. To switch, run
`uv tool uninstall kimpton-evalrouter-sdk` (or
`pip uninstall -y kimpton-evalrouter-sdk evalrouter`) and install `evalrouter`.

The package supplies the `evalrouter` command; no repository checkout is needed.
[Create an EvalRouter account](https://evalrouter.ai/register), verify your email,
and follow the browser sign-in below.
Package installation does not grant account access or evaluation credit.

Sign in with your account:

```sh
evalrouter login
evalrouter whoami
```

Follow the browser sign-in; on a local computer the browser returns to the
waiting CLI. For a remote terminal, use `evalrouter login --no-browser` and
approve the code on a device with a browser. New accounts finish setup there.
The CLI selects your only workspace automatically, or lets you choose one.
Production is the default API; `evalrouter options` lists `--base-url` and the other
settings that apply to every command.
No API key or credential-storage setup is needed for a normal sign-in.
The optional `evalrouter[keyring]` extra enables a compatible OS keychain
backend; otherwise credentials use a private file readable only by your user.
In the workspace, **Account and workspace → Sign in** starts the same flow.
See the [sign-in guide](https://evalrouter.ai/developers/docs/cli-login) for
workspace switching, stored API keys and signing out.

CI and services can use `EVALROUTER_API_KEY` and `EVALROUTER_WORKSPACE_ID` supplied
privately by a secret manager. Never put credentials in arguments, URLs, request
files or logs.

## Interactive workspace

Run `evalrouter` without arguments in a terminal to open the full-screen
workspace. Choose **Account and workspace → Sign in** to connect your account.
Local sign-in returns from the browser to the waiting CLI; for a remote terminal,
use `evalrouter login --no-browser` and approve the code on a device with a browser.

**Run a benchmark** guides you through the benchmark, one to eight models, each
model's reasoning settings and coverage. Type to filter lists; Space or Enter
toggles a model and Tab continues. A supported native one-model evaluation
starts directly with the optional `--max-cost` cap or available wallet credit,
limited by the service maximum. Supported native separate multi-model groups
also start directly, with an optional aggregate cap. Their finite service bounds
do not reserve the wallet once per model: mandatory fees and live calls use
shared available credit. Matched comparisons retain their quote and required cap.

Choose a suggested sample, **Custom sample size**, or **Full benchmark** for
coverage. A custom sample accepts 1 to 1,000 tasks, limited to the benchmark's
task count when known. Full benchmark includes every task.

Home includes **Run history**, **Run groups**, and **Results**, plus model and
benchmark browsing, billing and account settings. **Results** offers **Single
run**, **Group run**, then **Compare runs**. Group status shows each model's state;
open a member run, return to its group, or keep watching live progress.
While watching, choose **Open live run in browser** or **Open live group in
browser** with Enter, B or a click. Live status continues in the terminal.

Escape goes back; Left also goes back when you are not editing text. Leaving a
live run stops watching only: submitted server work continues. Ctrl+Q exits. Use
PgUp/PgDn for long reviews and results, and F1 for keyboard help. Select output
with the mouse and press Ctrl+C to copy; with no selection, Ctrl+C goes back.
Password fields cannot be copied. `EVALROUTER_TUI=off` selects the older interactive
shell; explicit commands, plain output and JSON remain available.

## Run a model benchmark

Sign in, choose a ready benchmark and compatible model, then run:

```sh
evalrouter login
evalrouter catalog
evalrouter catalog --models
evalrouter run BENCHMARK_SLUG --model MODEL_ID --sample 3
```

Supported native single-model benchmarks start directly. A spending cap is
optional; use `--max-cost USD` to set one. Without a cap, the service bounds
admission using your available wallet credit and its service maximum. An
explicit cap is never silently raised. Native multi-model groups and matched
comparisons also start without a required quote or cap. `--mode matched` pins
the same benchmark tasks for every model; separate mode runs each model on its
compatible subset. Native actual-usage routes charge recorded usage, including
usage above a live-call hold for work already sent. A negative wallet stops new
dispatches until credit is added. Accepted historical quotes keep their terms.

For scripts, approve paid work explicitly and retain the operation identity:

```sh
evalrouter run BENCHMARK_SLUG --model MODEL_ID --sample 3 --yes --json
evalrouter run BENCHMARK_SLUG --model MODEL_A --model MODEL_B --mode matched --sample 3 --yes --json
evalrouter status
evalrouter wait RUN_ID
evalrouter results RUN_ID
evalrouter export RUN_ID --format json --output result.json
```

A sample is not a full-benchmark score. A local timeout or interrupted watcher
does not stop server work. Reuse the saved request identity when recovering an
uncertain submission; a new identity can start another evaluation.

Use `evalrouter --help` and each command's `--help` for supported options.
Account commands include `login`, `whoami`, `workspace` and `logout`;
`credits` reads balance and holds. Payments take place on the website.

## Author your own benchmark

`benchmark init`, `benchmark validate` and `benchmark bundle` run entirely on
your machine: no account, no network request, and nothing from your package is
imported or executed. The normative validator ships inside `evalrouter`, and
its one dependency (pydantic) is an optional extra. The base installation uses
httpx, Rich and Textual. With pip, use `pip install 'evalrouter[benchmark]'`:

```sh
uv tool install "evalrouter[benchmark]"
evalrouter benchmark init ./my-benchmark --namespace example --name my-benchmark
evalrouter benchmark validate ./my-benchmark
evalrouter benchmark bundle ./my-benchmark   # evalrouter-bundle-example-my-benchmark-<version>.zip
```

Without the extra these three commands stop before reading your package and
report `benchmark_extra_required` with the exact install string; `--help` and
every network command keep working. `init` writes an explicitly synthetic draft
and claims no license: replace the tasks, rights, references and limits before
submitting. Local validation is not platform admission — EvalRouter revalidates
every submission on its own servers.

Where private benchmark imports are enabled, `benchmark upload`, `import` and
`create` use workspace-private data with the service's built-in graders. These
are account-backed operations, separate from local package validation. See the
[private benchmark guide](https://evalrouter.ai/developers/docs/byob) for manifests,
supported formats, source management and permissions.

## Python client

Install this alpha with `python -m pip install --pre evalrouter` in your application's environment. The
import name is `evalrouter`; the client supports context-manager cleanup:

Canonical imports require a release containing this namespace. Older registry
releases through 0.3.24 use `kimpton_evalrouter`. Select alpha releases explicitly
to use the canonical `evalrouter` import and current tracing APIs.

```python
import os
from evalrouter import Client

with Client(
    base_url="https://api.evalrouter.ai",
    api_key=os.environ["EVALROUTER_API_KEY"],
    workspace_id=os.environ["EVALROUTER_WORKSPACE_ID"],
) as client:
    print(client.catalog.models())
```

Use `client.catalog`, `client.evals`, `client.quotes` and `client.runs` for
typed discovery, submission and result resources. Creating a quote starts no evaluation; submitting its accepted
ID can start paid work. Keep the quote ID and idempotency key before submission.
The [model API guide](https://evalrouter.ai/developers/docs/model-api) describes
request fields, limits and result contracts.

## Import LangSmith traces

The optional `langsmith` extra provides a local connector for completed agent
runs. Where your workspace has Traces, create an SDK destination connection in
**Traces → Connect agent**, then supply `LANGSMITH_API_KEY` and the connection's
EvalRouter environment variables through your secret manager. The source key
is never sent to EvalRouter.

```sh
python -m pip install --pre 'evalrouter[langsmith]'
python -m evalrouter.langsmith projects
python -m evalrouter.langsmith sync --project 'YOUR_PROJECT_ID' \
  --state './.private/evalrouter-langsmith-state.json' --metadata-only --watch
```

The process must stay running for automatic imports. Its private checkpoint
resumes progress across restarts. Set `--since` to an ISO timestamp to choose
the initial import window. This explicit metadata-only mode omits source names,
inputs, outputs, error text, tags, events and arbitrary metadata. Hierarchy,
timing, error status and reported numeric usage remain. The compatibility default
captures bounded field-name-redacted content if `--metadata-only` is omitted.

Checkpoint version 2 binds the content policy and destination. Unknown
acknowledgments replay a private adjacent `.delivery.json` payload unchanged;
keep it with the checkpoint. A policy change on an existing checkpoint is
refused. Version 1 upgrades only in its original content-capture mode and cannot
restore payload bytes lost by the older connector. Finish unresolved deliveries
before using a new checkpoint and connection for a different policy. Only
activity recorded by LangSmith can be imported; two stable closed reads are
observed coverage, not proof that every upstream child exists.

## Live agent capture

Install the alpha with `python -m pip install --pre 'evalrouter[tracing]'` in your agent's own environment.
Configure `EVALROUTER_BASE_URL`, `EVALROUTER_WORKSPACE_ID`,
`EVALROUTER_TRACE_CONNECTION_ID` and the restricted `EVALROUTER_TRACE_KEY` through
an authorized secret manager. Install alpha releases explicitly; an ordinary install selects the stable release.

```python
import json
import os
from anthropic import Anthropic
from evalrouter.tracing import TraceCapture

# Also install anthropic, and supply ANTHROPIC_API_KEY privately.
model = Anthropic(max_retries=0)
capture = TraceCapture(capture_content=True)
with capture.agent("Support assistant") as root:
    root.input({"message": "Explain what an evaluation benchmark is."})
    with capture.model(os.environ["ANTHROPIC_MODEL"]) as span:
        answer = model.messages.create(
            model=os.environ["ANTHROPIC_MODEL"], max_tokens=256,
            messages=[{"role": "user", "content": "Explain what an evaluation benchmark is."}],
        )
        text = "".join(block.text for block in answer.content if block.type == "text")
        span.output({"answer": text})
    root.output({"answer": text})
receipt = capture.seal(root.trace_id)
print(json.dumps(receipt))
if receipt["state"] != "accepted" or not receipt["complete"] or not capture.close():
    raise SystemExit(1)
```

This example makes a real paid model request. Content capture is explicitly
enabled; use `capture_content=False` to omit input, output and error payloads.
Caller attributes still require your metadata policy.
Common secret field names are redacted, but secrets in arbitrary text cannot be
detected. Content capture defaults to true; select metadata-only explicitly.

Root exit and `flush()` do not seal. Join owned children, then `seal(root.trace_id)`
to acknowledge closure. `receipt(root.trace_id)` reports local delivery state,
counts, accepted server UUID and server completion, independently of evaluation
completion. Unknown acknowledgments retain identical batches for later
flush/seal/close on the same instance. Zero newly inserted spans on replay is
success. `close()` returns false while active, failed or unknown, and true after
acknowledged closure and shutdown. A context manager calls close, but does not
raise on a false result; inspect receipts explicitly. There is no background HTTP
export. Buffers are process-local; retry before exit. Each instance retains at
most 64 local trace receipts, 512 active spans and 16 MiB of pending records.
When another root needs room, the oldest successfully sealed idle receipt is
evicted. Save the receipt/server UUID if you need it later; `receipt(id)` raises
`ValueError` after eviction. Pending, unknown, active or dropped traces are never
evicted. If all 64 are unsettled, a new root is not recorded (its handle has no
trace ID and `failed_exports` increases); resolve delivery or use another instance.
Late activity inherited from an evicted trace is rejected, never made a new root.
Close attempts every retained trace within one aggregate deadline and leaves the
provider alive if any closure fails. Capture loss prevents successful seal. The optional OpenTelemetry dependency stays
isolated from the base HTTP client.

Create a connection under **Traces → Connect agent** where Traces is enabled for
your workspace. Detailed delivery receipts are described in the trace connection setup in your workspace.


## Attach existing OpenTelemetry instrumentation

Install the alpha `evalrouter[tracing]` extra into your existing Python
3.12/3.13 agent environment. Pass the OTel SDK provider the framework actually
uses; do not replace its global provider. No agent inheritance or per-call
EvalRouter wrappers are required for supported spans:

```python
from evalrouter.tracing import TraceCapture

capture = TraceCapture(capture_content=False)
attachment = capture.attach(existing_provider, scope_names=("your.instrumentation",))
# Run the already-instrumented agent, then join all owned child work.
for trace_id in attachment.trace_ids:
    receipt = capture.seal(trace_id)
    print(receipt)  # IDs/counts/delivery state only; inspect accepted + complete.
# At application shutdown, check capture.close(); it never stops existing_provider.
```

Attachment exports metadata only, even when wrapper content capture is enabled.
It removes source names, resource attributes, events, prompts, tool arguments,
results and exception text. It retains original trace/span/parent IDs, closed
operation/kind/status enums and bounded numeric token counters. Source counters
are observations, not native usage charges. `evalrouter.otel.status=unset` means
unknown source success: wire `ok` only means no explicit source ERROR, never a
native grade or proof that a tool succeeded. Unknown/conflicting kinds are
counted in `attachment.unsupported_spans` and prevent successful sealing.
Provider flush queues records; it does not prove accepted ingestion or closure.

The optional `[strands]` extra pins Strands for applications that need it; the
base client has no Strands dependency. Attachment only sees selected spans that
the supplied provider actually records. It cannot recover unsampled or missing
work, or capture uninstrumented frameworks automatically. LangSmith project
imports and ACP evaluation execution are separate integrations.

## Native multi-model client requests

The methods below require a client build containing native direct-group support;
use the supplied candidate archive until that release is published. Native group
service availability is controlled separately. Matched comparison launches keep
their existing quote/cap flow.

```python
from evalrouter import Client

with Client() as client:
    group = client.run_groups.create({
        "models": [
            {"kind": "managed", "route_id": "YOUR_FIRST_ROUTE"},
            {"kind": "managed", "route_id": "YOUR_SECOND_ROUTE"},
        ],
        "selection": {"profile_ids": ["YOUR_EXACT_NATIVE_PROFILE"]},
        "coverage": {"mode": "sample", "sample_count": 1, "seed": 42},
        # Optional aggregate stop limit, in integer micro-USD:
        # "max_charge_microusd": "2000000",
    }, idempotency_key="persist-this-group-request-key")
    # If work stops, resume eligible native member chains with a NEW durable key:
    # continued = client.run_groups.resume(group["id"], {},
    #     idempotency_key="persist-this-continuation-request-key")
```

An omitted customer cap returns `max_charge_microusd=None`;
`authorization_limit_microusd` is a finite service bound. It is not the money
reserved or a budget selected by you. Preserve the exact body and key across
uncertain replies. Actual charges and available wallet credit remain authoritative.
