Metadata-Version: 2.4
Name: evalrouter
Version: 0.3.23
Summary: Interactive evaluation workspace, command-line tools, and Python client for EvalRouter
Project-URL: Homepage, https://evalrouter.ai
Project-URL: Documentation, https://evalrouter.ai/developers/docs/cli
Project-URL: Model API, https://evalrouter.ai/developers/docs/model-api
License-Expression: LicenseRef-Kimpton-EvalRouter-SDK
License-File: LICENSE
Keywords: benchmarks,cli,evalrouter,llm,model-evaluation
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Typing :: Typed
Requires-Python: <3.14,>=3.12
Requires-Dist: httpx==0.28.1
Requires-Dist: rich==15.0.0
Requires-Dist: textual==8.2.8
Provides-Extra: benchmark
Requires-Dist: pydantic==2.13.5; extra == 'benchmark'
Provides-Extra: keyring
Requires-Dist: keyring==25.7.0; extra == 'keyring'
Provides-Extra: langsmith
Requires-Dist: langsmith==0.12.6; extra == 'langsmith'
Provides-Extra: tracing
Requires-Dist: opentelemetry-api==1.45.0; extra == 'tracing'
Requires-Dist: opentelemetry-sdk==1.45.0; extra == 'tracing'
Description-Content-Type: text/markdown

# EvalRouter CLI and Python client

Evaluate models in a full-screen terminal workspace or from scripts. Browse
benchmarks, choose one or more models, review a spending cap, follow runs, and
compare or export their results. The same package includes the typed Python
client, optional trace capture, and benchmark authoring tools.

Requires Python 3.12 or 3.13. Distributed under the proprietary Kimpton EvalRouter
SDK License, included as `LICENSE` in the package.

## Install and create your account

We recommend [uv](https://docs.astral.sh/uv/getting-started/installation/) to install the CLI in its own Python environment:

```sh
uv tool install evalrouter
evalrouter
```

With pip, use `pip install evalrouter` in a Python 3.12 or 3.13 environment.
uv selects a compatible Python for the CLI and downloads one if needed. If
your shell cannot find `evalrouter`, run `uv tool update-shell` and restart your
terminal. Upgrade an existing installation with `uv tool upgrade evalrouter`.
Check the installed version with `evalrouter --version`. Use `evalrouter --help`
for explicit commands, or see the [CLI guide](https://evalrouter.ai/developers/docs/cli).

The interactive workspace checks for newer releases in the background and shows
an update notice when one is available. The check never installs an update or
delays opening the workspace. Set `EVALROUTER_NO_UPDATE_CHECK=1` to disable it.
Scripts, editable installs and explicitly pinned installations skip update checks.

### Complete command names with Tab

For zsh, add this line to `~/.zshrc` after your `compinit` setup, then start a new
terminal:

```sh
eval "$(evalrouter completion zsh)"
```

For bash, add this line to `~/.bashrc`:

```sh
eval "$(evalrouter completion bash)"
```

Type `evalrouter ru` and press Tab to complete `run`. With `evalrouter r`, Tab
shows matching commands such as `run`, `resume` and `results`. The shell setup
only completes top-level command names and makes no API request. It does not
change interactive prompts or run an evaluation.

This package was previously published as `kimpton-evalrouter-sdk`. The command
and the `kimpton_evalrouter` import name stay the same. To switch, run
`uv tool uninstall kimpton-evalrouter-sdk` (or
`pip uninstall -y kimpton-evalrouter-sdk evalrouter`) and install `evalrouter`.

The package supplies the `evalrouter` command; no repository checkout is needed.
[Create an EvalRouter account](https://evalrouter.ai/register), verify your email,
and follow the browser sign-in below.
Package installation does not grant account access or evaluation credit.

Sign in with your account:

```sh
evalrouter login
evalrouter whoami
```

Follow the browser sign-in; on a local computer the browser returns to the
waiting CLI. For a remote terminal, use `evalrouter login --no-browser` and
approve the code on a device with a browser. New accounts finish setup there.
The CLI selects your only workspace automatically, or lets you choose one.
Production is the default API; `evalrouter options` lists `--base-url` and the other
settings that apply to every command.
No API key or credential-storage setup is needed for a normal sign-in.
The optional `evalrouter[keyring]` extra enables a compatible OS keychain
backend; otherwise credentials use a private file readable only by your user.
In the workspace, **Account and workspace → Sign in** starts the same flow.
See the [sign-in guide](https://evalrouter.ai/developers/docs/cli-login) for
workspace switching, stored API keys and signing out.

CI and services can use `EVALROUTER_API_KEY` and `EVALROUTER_WORKSPACE_ID` supplied
privately by a secret manager. Never put credentials in arguments, URLs, request
files or logs.

## Interactive workspace

Run `evalrouter` without arguments in a terminal to open the full-screen workspace.
**Run a benchmark** uses one flow for one to eight models: choose a benchmark,
check compatible models, choose each model's reasoning immediately, set coverage
and a spending cap, then review the quote before approving. Space or Enter toggles
a model; Tab continues from any row. Type to filter lists.

For coverage, choose a suggested sample, **Custom sample size**, or **Full
benchmark**. A custom sample accepts 1 to 1,000 tasks, limited to the benchmark's
task count when known. Full benchmark includes every task.

The quote and fresh available credit stay beside the approval choices. A known
shortfall blocks starting; an unavailable balance is labelled honestly. The service
checks credit again at admission. Resizing cannot approve a quote, and an expired
quote requires another review. Escape goes back; Left also goes back when you are
not editing text. On a live evaluation these keys detach only the view: submitted
evaluations continue on the server. Ctrl+Q exits.

Home includes **Run history**, **Run groups**, **Results**, model and benchmark
browsing, billing, and account/workspace settings. **Results** offers **Single
run**, **Group run**, then **Compare runs**. Group status shows each model's
state and lets you open its individual run, return to the group, or watch live.
While watching, choose **Open live run in browser** or **Open live group in
browser** with Enter, B or a click. Live status continues in the terminal.
Results and errors remain until you choose an action. History includes a refresh
action. PgUp/PgDn scroll long reviews and results. Saved report paths appear on
the result page and again after exit.
Benchmark uploads open in the browser. `EVALROUTER_TUI=off` selects the older
interactive shell; explicit commands, plain output and JSON remain available.
Select output with the mouse and press Ctrl+C to copy; with no selection,
Ctrl+C goes back. Password fields cannot be copied. F1 opens keyboard help.

## Run an evaluation in one command

Choose a ready benchmark and a compatible model from `evalrouter catalog` and
`evalrouter catalog --models`, then replace the placeholders:

```sh
evalrouter run BENCHMARK_SLUG --model MODEL_ID
```

`run` matches the benchmark and model by slug, name or ID, suggesting the
closest names after a typo. It creates a free quote, shows the coverage,
published rates, spending cap and concurrency, and asks whether to start the run (Start run or Cancel).
You are charged for actual usage at those rates, never more than your cap.
Only a yes starts paid work. It then follows progress and prints the scores,
the change from your previous completed run of the same benchmark and model,
and the report link. In a terminal, missing choices are asked for; run
`evalrouter run` alone to pick everything. `--sample N` or `--full` sets
coverage, `--max-cost USD` the cap and `--dry-run` stops after the quote.

The model picker uses the selected benchmark's compatibility catalog. Where
the catalog supplies reasoning controls, it offers that route's supported
efforts or thinking-token budget and identifies fixed or provider-default settings.
In a model checklist, Space or Enter selects a model; Tab continues from any
row once enough models are selected. The quote rechecks the chosen settings.

Without a terminal (CI, agents, `--json` or `--no-input`) `run` never prompts
and refuses unless both `--yes` and `--max-cost` are given:

```sh
evalrouter run BENCHMARK_SLUG --model MODEL_ID --max-cost 5 --yes --json
```

The idempotency key of each accepted quote is saved locally before the run is
submitted (`~/.local/state/evalrouter`, or `EVALROUTER_STATE_DIR`; identifiers
only, never credentials). Repeating the same command after a dropped connection
resumes that run instead of starting another; `--new` starts a separate one.
Afterwards, `results RUN`, `export RUN --format html,json` and `wait RUN`
take the run's ID or a unique prefix of it. `results --diff` also compares
with the previous run of the same benchmark and model; that takes a few more
requests, so plain `results` does not. `run --from RUN` runs an earlier
evaluation again as a new run with the same setup; other `run` options override
it. Watching a run survives brief API
outages such as a 503 during a deployment. If the API stays unreachable for
90 seconds (`--disconnect-timeout SECONDS`), the CLI stops watching, says so and
prints the `evalrouter wait` command that reattaches; the run continues.

### Run several models together

Repeat `--model` to evaluate two to eight models on the same benchmark. Review
one group quote with each model's settings and cap before starting:

```sh
evalrouter run BENCHMARK_SLUG --model MODEL_A --model MODEL_B --sample 25 --max-cost 10
evalrouter group list
evalrouter group status GROUP_ID --watch
evalrouter group compare GROUP_ID --format html,json --output group-report
```

In the full-screen workspace, the group quote keeps the total spending cap and
approval on one page. Press Enter on the price row to reveal **Enter to confirm**,
then Enter again to start. **Enter another amount** changes the cap and requests
a fresh quote when the amount changes. **Cancel** starts nothing. Moving away
from the price row, pressing PgUp/PgDn or resizing clears the pending confirmation.

Each member keeps its own run ID, results and charges. `group status GROUP_ID -i`
opens an individual member; `group resume GROUP_ID --max-cost USD` reviews a
continuation for eligible unfinished work. Leaving the view stops watching only.
Saved group reports compare retained results without making new model calls.

## Resume a stopped run

```sh
evalrouter resume RUN_ID --dry-run
evalrouter resume RUN_ID
```

When a run stops early, for example because it reached its spending cap,
`resume` quotes a new run that covers the same tasks. Tasks the original run
already graded are reused at $0, with no model request and no fee; only the
remaining tasks are rerun and charged, within a new cap. The quote shows the
source run, how many tasks are reused and rerun, which are left out and why,
the cap, and when the reusable evidence and the quote expire.
Only a yes (or `--yes`) starts paid work; `--dry-run` never does. Without a
terminal, pass both `--yes` and `--max-cost USD`. A running run is only
watched. The original run and its results are never changed.

The accepted quote and its key are saved locally before submission, so
repeating the command after a dropped connection, or after adding credit when
the account was short, submits the same quote and cannot start a second run.
An expired quote is replaced by a fresh one, which must be confirmed again if
its cap differs. Repeating it with a different `--max-cost` or
`--include-cancelled` choice gets a fresh quote instead; if the earlier
submission already started a run, that run is shown and its scope stays as
accepted. A run can be resumed once: `resume` on a run that was already
resumed follows it to the newest run of that chain. A newest run that is still
running is watched; a finished one is quoted and asked about again, even with
`--yes`, and scripts are refused with its ID (`continuation_exists`, exit 2).
Tasks stopped by cancellation or the time limit are rerun only with
`--include-cancelled`.

Refused tasks and cancellations omitted without `--include-cancelled` remain
excluded from this and every later resume. The quote asks you to acknowledge
those permanent exclusions; scripts need `--accept-left-out`.

New quotes show unresolved effects separately as **Pending**. Safe tasks can
continue while those cells retain their original provenance. A later resume
rechecks the original outcome and can run it only when it is genuinely safe;
a waived charge with an unknown provider outcome does not qualify. Nothing
retries automatically. Older quotes without pending provenance keep their
original permanent exclusions, except that an unknown provider outcome can be
offered for an explicit new attempt (below).

A task whose provider call failed is decided from that call's record, never
from its error name. If its original answer was stored, the resume grades that
stored answer at no charge, with no new request (`saved_response_cells`). If the
call is still unresolved, the quote shows the task as pending until the work
that made the call has finished (services before September 30, 2026 wait for
reconciliation, up to 24 hours after the call started; the quote says when). If
the outcome is uncertain (the provider may have processed the request) or the
provider did not bill a transient failure, nothing retries it on its own:
`--retry-unknown` consents to pay for a new attempt of each such task within
the same cap. You are never charged for the failed original, and a late
original answer never adds a charge. If a new attempt fails the same way, a
later resume can offer another, and each one needs your consent again
(services before September 30, 2026 allow one new attempt per task).
`--retry-unknown-max N` refuses (never trims) a quote with more such tasks. With your own provider connection, that
provider may bill both requests, so `--accept-duplicate-provider-charges` is
also required. The JSON quote reports these as `continuation.unknown_retry`.

Resume supports one model, one benchmark and one attempt per task on the
reviewed native benchmark families (Inspect, SOB text and lm-eval) without a
judge model, within the source run's evidence retention and with compatible
execution semantics. Comparisons, several benchmarks or attempts, judge-graded
benchmarks, repository agents and your own benchmarks are refused rather than
rerun. Graded tasks from runs that finished before resume evidence was recorded,
including existing runs, cannot gain that evidence later. Reuse across service
updates requires compatible recorded runtime and evaluation behavior; the quote
refuses evidence that fails those checks. Tasks whose earlier charge or provider
outcome is unresolved are never called until a fresh quote proves eligibility. A server without resume
says so, and nothing is quoted or charged.

## Lists

`catalog`, `catalog --models`, `status`, `connections list` and `workspace list`
print a compact table sized to your terminal with one suggested next command;
`catalog` hides benchmarks in review or unavailable unless you pass `--all`, and
`catalog -i` browses them interactively; `status -i` keeps your run list live and
opens a run with Enter. Home's **Browse models** (or `catalog --models -i`)
opens a searchable model list. Enter shows the exact route, published prices
and reasoning support; Escape or Left returns to the list, then Home. Loading
can also be left with Escape or Left. Browsing never quotes or starts a run. Every list accepts `--json`,
`--jq '.data[].id'` (a small jq subset), `-w/--wide`, `--columns a,b`,
`--sort COL` (`--sort=-COL` descending), `--filter KEY=VALUE` and `--limit N`.
Piped output stays JSON. Colour follows `NO_COLOR`, `CI` and `TERM=dumb`.
Help, static benchmark and model pages, results and group status opened from
Home stay visible until Enter, Escape or Left, including in a plain terminal.

`run` shows the quote (published rates and, when the quote states one, the
minimum cap) before asking for the cap. The suggested cap is the standard $5,
or the quote's minimum when that is higher, and the CLI says so. It refuses a
cap that provably cannot start one task and asks before repeating a run that just finished
(scripts pass `--new`). `--all-metrics` lists every metric and comparison.

## Discover, quote, then run

```sh
evalrouter catalog --status active --json
evalrouter catalog --models --json
evalrouter catalog --slug BENCHMARK_SLUG --json
```

Choose a profile whose `quote_availability.status` is `ready_for_quote` and a
compatible managed model route from those responses. An active catalog entry
alone does not mean it is ready to run.

These three catalog reads are public, so the CLI sends them without your
credentials: they never read the keychain or refresh a sign-in. If the API
explicitly allows it (`Cache-Control` with `max-age`, and no `no-store`,
`no-cache` or `private`), the CLI keeps a local copy for at most five minutes
and says so on stderr when it uses one, except for online reads inside Home.
Home quietly warms its two initial benchmark lists in the background using the
same cache rules. `--refresh` reads the catalog now,
asking caches along the way to revalidate, and fails rather than show an older copy, and `--offline` uses only a fresh local
copy with no network. When a server marks a response `no-store`, the command still makes a request
and no offline copy is retained. Availability shown from a copy is informational: a quote always
checks it again. Workspace data, run status, results, quotes, credits and signed
URLs are never cached. The copy lives in `$XDG_CACHE_HOME/evalrouter` (or
`~/.cache/evalrouter`; `%LOCALAPPDATA%\evalrouter\cache` on Windows, or
`EVALROUTER_CACHE_DIR`), may be deleted at any time, and
`EVALROUTER_READ_CACHE=off` disables it.

Save this as `quote.json`, replacing both uppercase identifiers:

```json
{
  "model": {"kind": "managed", "route_id": "MODEL_ROUTE_ID"},
  "selection": {"profile_ids": ["BENCHMARK_PROFILE_ID"]},
  "coverage": {"mode": "sample", "sample_count": 3, "seed": 42},
  "max_charge_microusd": "1000000"
}
```

```sh
evalrouter quote --config quote.json --json
```

Review compatibility, coverage, rates, the cap, warnings and expiry. A quote
starts no paid work. EvalRouter does not publish cost forecasts: charges follow
actual usage at the quoted rates and never exceed `max_charge_microusd`, and
the credit held at start is released as accounting settles. Older servers may
still return the optional legacy field `estimated_charge_microusd` (a
conservative worst case, not a forecast of spend); the CLI does not display it. The example's $1 platform cap is not a price guarantee or
promise that a particular evaluation fits. Paid runs require available credit.

Persist the returned quote ID and an operation key before submitting:

```sh
evalrouter run --quote REVIEWED_QUOTE_ID --idempotency-key SAVED_OPERATION_KEY --json
evalrouter wait RUN_ID --wait-timeout 3600 --json
evalrouter results RUN_ID --json
evalrouter export RUN_ID --format json --output result.json
evalrouter export RUN_ID --format csv --output result.csv
evalrouter export RUN_ID --format html --output result.html
```

Replace `RUN_ID` with the returned ID. A lost submission response is recovered
with the same quote and operation key; a new key may start separate work.
Inspect terminal status, coverage, errors and billing alongside scores. A sample
is not a full-benchmark score. Exports refuse overwrite by default; use the
result's integer `--version` for repeatable reports.

Every command that saves a file prints its full path. An explicit `--output`
file is used exactly and never replaced (exports accept `--overwrite`); a
directory, or no file name, gets a default such as `evalrouter-<run>.json` or
`evalrouter-bundle-<namespace>-<name>-<version>.zip`, moving to the next free
`-2`, `-3`... name instead of replacing anything. Add `--open` or `--reveal`
to open the saved file or show it in your file manager; nothing opens
otherwise, and a missing viewer never fails the command.

### Comparing runs

```sh
evalrouter compare RUN_A RUN_B --format html --output my-comparison
```

`compare` takes 1 to 8 finished runs of the same benchmark version, dataset,
task population, task selection and output format in one workspace. Each run
must be a managed or connected model run; comparison, agent and environment
runs are refused, and so is a run whose export does not record one of these
identities. It reads each run's summary export (`export?format=summary`:
identities, frozen native summary, coverage, billing snapshot and export
policy, with no sample rows) and saves `comparison.json` and
`sources/<run>.summary.json` in a new private directory. `--evidence`
downloads each run's full JSON export instead (`sources/<run>.json`) and also
compares grader versions and graded task sets; the scores are the same.
`comparison.json` holds each run's model identity (route, revision, protocol,
recorded limits, generation and routing settings), coverage, errors, its
frozen native summary with the headline's own denominator, and its settled
EvalRouter charge. A run that did not complete, or graded fewer tasks than
planned, is marked partial in `comparison.json` and in its own row of the
terminal table.

For SOB Text runs that were continued (`evalrouter resume`), add
`--include-resumed` (`--chain` is the same flag): each ID then names one
chain, meaning the original run and its resumes, and the chain becomes one
series. The service decides every task from the one run that graded it,
re-derives the SOB native summary with the benchmark authors' code, and keeps
ungraded tasks as errors with their codes. Each score is shown with its runs
("original + 1 resume") and its graded-of-planned count. If a resume replayed
a reused answer and got a different score (`reuse_score_mismatch`), the
report warns and the chain is not complete. `evalrouter export RUN
--include-resumed` saves that chain result by itself. The command follows
each run to its newest resume at the time you run it, so a later resume makes
it return a longer chain; the saved chain documents keep the exact version.

Aggregate-only reports show the retained metric values and their recorded sample
counts, plus per-code error or stop counts from the saved chain. Sample identities
and answers remain withheld. Missing counts and confidence intervals remain
unknown. SOB @6 reports identify their generation variant and each model's graded
denominator; unscored outcomes are kept separate without assigning them zero.
An existing downloaded archive does not change when the CLI is upgraded: generate
a new report to use the updated presentation.

`sources/` keeps each export byte for byte, with its SHA-256 recorded. With
`--evidence`, every graded sample must carry its scored answer by default: the final completion
text for SOB, plus the per-sample metadata the export supplies (for SOB, the
response's token counts and finish reason). Exports do not contain complete
transcripts, reasoning, timings or per-attempt usage. Keep these files
private. Use `--without-raw` (which implies `--evidence`) to archive runs
whose answers have expired or are withheld by the export policy. A chain
summary report downloads member exports only with `--evidence`. Different transports, limits or generation
settings are listed as warnings. Nothing is regraded, and no deltas,
intervals or winner are computed. `--output NAME.json` puts the sources in
`NAME-sources/`.

## Supported commands

| Task | Commands |
| --- | --- |
| Discover benchmarks and models | `catalog`, `catalog --models`, `catalog --slug SLUG` |
| Quote and run | `quote`, `run`, `resume` |
| Follow evaluations | `status`, `wait`, `events`, `cancel` |
| View and save results | `results`, `export`, `compare`, `open` |
| Manage run groups | `group list/status/resume/compare` |
| Sign in and choose a workspace | `login`, `whoami`, `workspace list/switch`, `logout` |
| Connect a model endpoint | `connections list/create/check/update/disable` |
| Check credit | `credits` |
| Author a benchmark locally | `benchmark init/validate/bundle` |
| Submit or manage private benchmark data, where enabled | `benchmark submit/upload/import/create/list/archive/sources/archive-source` |
| Inspect or share submitted benchmark packages | `benchmark status/inspect/grant/revoke` |
| Regrade a private benchmark's retained outputs, where enabled | `regrade` |
| Prepare repository agents, where supported | `agent-builds options/preview/create/list/status/cancel` |
| Report a problem and track it | `feedback`, `feedback list`, `feedback status ID` |
| Configure the command line | `options`, `completion bash/zsh`, `--help`, `--version` |

Slashes above separate subcommands; use, for example, `evalrouter group list`.
`credits` shows the workspace balance, the amount held by running evaluations
and what is available for new ones. It works with an `evalrouter login` sign-in
or an API key with the `billing:read` scope. Credit is added on the website under
Settings > Billing; the CLI never takes a payment.
Run each command with `--help` for exact flags. A connected model's credential
is never a command argument: `connections create` asks with hidden input in a
terminal, and scripts use `--key-stdin` or `--key-env NAME`
(`connections update --rotate-key` replaces it). Connection checks can send
a small model request, and your model provider may bill separately.

For `wait` and `run --wait`, `--progress auto` shows elapsed time and status in
an interactive terminal, with a percentage when validated processed and planned
counts are available. Use `--progress plain` for structured progress or
`--progress off` to suppress it. Redirected stderr and `--json` retain structured
progress. A progress percentage counts processed samples; it is not a score.
`--progress json` writes `evalrouter.progress.v1` records on stderr, on every
change and every 30 seconds: phase, processed and planned counts, errors,
spend and elapsed time. `evalrouter events RUN --follow` prints the run's
event log as JSON lines until it ends; `--type`, `--tail N` and `--after N`
filter, start from the end and resume.
`evalrouter open RUN` opens the run's live report page, also when an agent
runs it with `--json`; `run --open` opens it as soon as the run starts. No
browser opens in CI, over SSH, without a display, with `--no-browser` or with
`EVALROUTER_NO_BROWSER=1`; the address is printed instead.
For a group, use `evalrouter open GROUP_ID --group --live` for live status or
`evalrouter open GROUP_ID --group` for its comparison. Both require the full
group ID and open the website without downloading a report.

Use `--config -` for JSON on stdin. `--json` writes one result on stdout and
progress on stderr. Exit codes: 0 success; 1 API/transport or failed run;
2 invalid input or rejected request (including a refused unattended `run`);
4 partial/cancelled waited run; 5 confirmation declined, nothing started;
6 stopped watching because the API stayed unreachable. Interrupting
a local wait does not cancel server work. Use `evalrouter cancel RUN_ID`
explicitly when cancellation is intended.

See the [CLI documentation](https://evalrouter.ai/developers/docs/cli) for account
setup, model selection, spending semantics, recovery and versioned exports.
The CLI uses the [model API](https://evalrouter.ai/developers/docs/model-api),
which is also documented for direct HTTP integrations.

## Author your own benchmark

`benchmark init`, `benchmark validate` and `benchmark bundle` run entirely on
your machine: no account, no network request, and nothing from your package is
imported or executed. The normative validator ships inside `evalrouter`, and
its one dependency (pydantic) is an optional extra. The base installation uses
httpx, Rich and Textual. With pip, use `pip install 'evalrouter[benchmark]'`:

```sh
uv tool install "evalrouter[benchmark]"
evalrouter benchmark init ./my-benchmark --namespace example --name my-benchmark
evalrouter benchmark validate ./my-benchmark
evalrouter benchmark bundle ./my-benchmark   # evalrouter-bundle-example-my-benchmark-<version>.zip
```

Without the extra these three commands stop before reading your package and
report `benchmark_extra_required` with the exact install string; `--help` and
every network command keep working. `init` writes an explicitly synthetic draft
and claims no license: replace the tasks, rights, references and limits before
submitting. Local validation is not platform admission — EvalRouter revalidates
every submission on its own servers.

Where private benchmark imports are enabled, `benchmark upload`, `import` and
`create` use workspace-private data with the service's built-in graders. These
are account-backed operations, separate from local package validation. See the
[private benchmark guide](https://evalrouter.ai/developers/docs/byob) for manifests,
supported formats, source management and permissions.

## Python client

Install with `pip install evalrouter` in your application's environment. The
import name is `kimpton_evalrouter`; the client supports context-manager cleanup:

```python
import os
from kimpton_evalrouter import Client

with Client(
    base_url="https://api.evalrouter.ai",
    api_key=os.environ["EVALROUTER_API_KEY"],
    workspace_id=os.environ["EVALROUTER_WORKSPACE_ID"],
) as client:
    print(client.catalog.models())
```

Use `client.quotes`, `client.runs`, `client.traces` and `client.datasets` for their
typed resources. Creating a quote starts no evaluation; submitting its accepted
ID can start paid work. Keep the quote ID and idempotency key before submission.
The [model API guide](https://evalrouter.ai/developers/docs/model-api) describes
request fields, limits and result contracts.

## Import LangSmith traces

The optional `langsmith` extra provides a local connector for completed agent
runs. Where your workspace has Traces, create a LangSmith connection in
**Traces → Connect agent**, then supply `LANGSMITH_API_KEY` and the connection's
EvalRouter environment variables through your secret manager. The source key
is never sent to EvalRouter.

```sh
pip install 'evalrouter[langsmith]'
python -m kimpton_evalrouter.langsmith projects
python -m kimpton_evalrouter.langsmith sync --project 'YOUR_PROJECT_ID' \
  --state './evalrouter-langsmith-state.json' --watch
```

The process must stay running for automatic imports. Its private checkpoint
resumes progress across restarts. Set `--since` to an ISO timestamp to choose
the initial import window. Inputs and outputs are bounded and field-name
redacted; only activity recorded by LangSmith can be imported.

## Live agent capture

Install `evalrouter[tracing]` and use `from kimpton_evalrouter.tracing import TraceCapture`.
Where your workspace has Traces, create a connection and configure its restricted
ingest key with `EVALROUTER_TRACE_KEY` and `EVALROUTER_TRACE_CONNECTION_ID`, alongside the API
origin and workspace ID. Wrap each run with `capture.agent(...)`, and calls with
`capture.model(...)` or `capture.tool(...)`. Use `span.input(value)` and
`span.output(value)` to capture content. Call `capture.flush()` before a
serverless invocation exits and `capture.close()` on shutdown. The optional
OpenTelemetry dependency is isolated from the base HTTP client.

In your workspace's **Traces** page, choose **Connect agent** for the
connection's setup examples and capture settings. Capture can include model inputs
and outputs; configure redaction before sending private content.

Use `client.traces` for connection, trace, evaluator and finding resources, and
`client.datasets` for curation, review, immutable versions and exports.
These operations require the corresponding service release and workspace permissions.
