Metadata-Version: 2.4
Name: evalrouter
Version: 0.3.11
Summary: Typed Kimpton evaluation client and command-line interface
License-Expression: LicenseRef-Kimpton-EvalRouter-SDK
License-File: LICENSE
Requires-Python: <3.14,>=3.12
Requires-Dist: httpx==0.28.1
Requires-Dist: rich==15.0.0
Provides-Extra: benchmark
Requires-Dist: pydantic==2.13.5; extra == 'benchmark'
Provides-Extra: keyring
Requires-Dist: keyring==25.7.0; extra == 'keyring'
Provides-Extra: langsmith
Requires-Dist: langsmith==0.12.6; extra == 'langsmith'
Provides-Extra: tracing
Requires-Dist: opentelemetry-api==1.45.0; extra == 'tracing'
Requires-Dist: opentelemetry-sdk==1.45.0; extra == 'tracing'
Description-Content-Type: text/markdown

# EvalRouter CLI

Evaluate models from your terminal: discover
supported benchmarks and models, review a spending cap, run an evaluation, and export results. Requires Python 3.12
or 3.13. Distributed under the proprietary [Kimpton EvalRouter SDK License](LICENSE).

## Install and create your account

We recommend [uv](https://docs.astral.sh/uv/getting-started/installation/) to install the CLI in its own Python environment:

```sh
uv tool install --python 3.12 evalrouter
evalrouter --help
```

With pip, use `pip install evalrouter`. uv supplies Python 3.12 if needed. If
your shell cannot find `evalrouter`, run `uv tool update-shell` and restart your
terminal. Upgrade an existing installation with `uv tool upgrade evalrouter`.

This package was previously published as `kimpton-evalrouter-sdk`. The command,
the `kimpton_evalrouter` import and every API are unchanged. To switch, run
`uv tool uninstall kimpton-evalrouter-sdk` (or
`pip uninstall -y kimpton-evalrouter-sdk evalrouter`) and install `evalrouter`.

The package supplies the `evalrouter` command; no repository checkout is needed.
[Create an EvalRouter account](https://evalrouter.ai/register), verify your email,
and follow the browser sign-in below.
Package installation does not grant account access or evaluation credit.

Sign in with your account:

```sh
evalrouter login
evalrouter whoami
```

Approve the browser code. New accounts finish setup on the website; the CLI
continues when that is complete. It selects your only workspace automatically.
Production is the default API; `evalrouter options` lists `--base-url` and the other
settings that apply to every command.
No API key or credential-storage setup is needed for a normal sign-in.

CI and services can use `EVALROUTER_API_KEY` and `EVALROUTER_WORKSPACE_ID` supplied
privately by a secret manager. Never put credentials in arguments, URLs, request
files or logs.

## Import LangSmith traces

The optional `langsmith` extra provides a local connector for completed agent
runs. Create a LangSmith connection in **Traces → Connect agent**, then supply
`LANGSMITH_API_KEY` and the connection's EvalRouter environment variables through
your secret manager. The source key is never sent to EvalRouter.

```sh
pip install 'evalrouter[langsmith]'
python -m kimpton_evalrouter.langsmith projects
python -m kimpton_evalrouter.langsmith sync --project 'YOUR_PROJECT_ID' \
  --state './evalrouter-langsmith-state.json' --watch
```

The process must stay running for automatic imports. Its private checkpoint
resumes progress across restarts. Set `--since` to an ISO timestamp to choose
the initial import window. Inputs and outputs are bounded and field-name
redacted; only activity recorded by LangSmith can be imported. This connector
requires a package build that includes the `langsmith` extra.

## Run an evaluation in one command

```sh
evalrouter run gpqa-diamond --model gpt-4o-mini
```

`run` matches the benchmark and model by slug, name or ID, suggesting the
closest names after a typo. It creates a free quote, shows the coverage,
conservative estimate, spending cap and concurrency, and asks whether to start the run (Start run or Cancel).
Only a yes starts paid work. It then follows progress and prints the scores,
the change from your previous completed run of the same benchmark and model,
and the report link. In a terminal, missing choices are asked for; run
`evalrouter run` alone to pick everything. `--sample N` or `--full` sets
coverage, `--max-cost USD` the cap and `--dry-run` stops after the quote.

Without a terminal (CI, agents, `--json` or `--no-input`) `run` never prompts
and refuses unless both `--yes` and `--max-cost` are given:

```sh
evalrouter run gpqa-diamond --model gpt-4o-mini --max-cost 5 --yes --json
```

The idempotency key of each accepted quote is saved locally before the run is
submitted (`~/.local/state/evalrouter`, or `EVALROUTER_STATE_DIR`; identifiers
only, never credentials). Repeating the same command after a dropped connection
resumes that run instead of starting another; `--new` starts a separate one.
Afterwards, `results RUN`, `export RUN --format html,json` and `wait RUN`
take the run's ID or a unique prefix of it. `results --diff` also compares
with the previous run of the same benchmark and model; that takes a few more
requests, so plain `results` does not. `run --from RUN` runs an earlier
evaluation again as a new run with the same setup; other `run` options override
it. Watching a run survives brief API
outages such as a 503 during a deployment. If the API stays unreachable for
90 seconds (`--disconnect-timeout SECONDS`), the CLI stops watching, says so and
prints the `evalrouter wait` command that reattaches; the run continues.

## Resume a stopped run

```sh
evalrouter resume RUN_ID --dry-run
evalrouter resume RUN_ID
```

When a run stops early, for example because it reached its spending cap,
`resume` quotes a new run that covers the same tasks. Tasks the original run
already graded are reused at $0, with no model request and no fee; only the
remaining tasks are rerun and charged, within a new cap. The quote shows the
source run, how many tasks are reused and rerun, which are left out and why,
the worst case, the cap, and when the reusable evidence and the quote expire.
Only a yes (or `--yes`) starts paid work; `--dry-run` never does. Without a
terminal, pass both `--yes` and `--max-cost USD`. A running run is only
watched. The original run and its results are never changed.

The accepted quote and its key are saved locally before submission, so
repeating the command after a dropped connection, or after adding credit when
the account was short, submits the same quote and cannot start a second run.
An expired quote is replaced by a fresh one, which must be confirmed again if
its cap differs. Repeating it with a different `--max-cost` or
`--include-cancelled` choice gets a fresh quote instead; if the earlier
submission already started a run, that run is shown and its scope stays as
accepted. A run can be resumed once: `resume` on a run that was already
resumed follows it to the newest run of that chain. A newest run that is still
running is watched; a finished one is quoted and asked about again, even with
`--yes`, and scripts are refused with its ID (`continuation_exists`, exit 2).
Tasks stopped by cancellation or the time limit are rerun only with
`--include-cancelled`.

Refused tasks and cancellations omitted without `--include-cancelled` remain
excluded from this and every later resume. The quote asks you to acknowledge
those permanent exclusions; scripts need `--accept-left-out`.

New quotes show unresolved effects separately as **Pending**. Safe tasks can
continue while those cells retain their original provenance. A later resume
rechecks the original outcome and can run it only when it is genuinely safe;
a waived charge with an unknown provider outcome does not qualify. Nothing
retries automatically. Older quotes without pending provenance keep their
original permanent exclusions, except that an unknown provider outcome can be
offered for one explicit new attempt (below).

A task whose provider call failed is decided from that call's record, never
from its error name. If its original answer was stored, the resume grades that
stored answer at no charge, with no new request (`saved_response_cells`). If the
call is still unresolved, the quote says when reconciliation settles it (up to
24 hours after it started). If the outcome is uncertain (the provider may have
processed the request; your charge for it was waived) or the provider did not
bill a transient failure, nothing retries it on its own: `--retry-unknown`
consents to pay for **one** new attempt of each such task, once ever, within
the same cap. The original outcome and charge stay unchanged, and a late
original answer never adds a charge. `--retry-unknown-max N` refuses (never
trims) a quote with more such tasks. With your own provider connection, that
provider may bill both requests, so `--accept-duplicate-provider-charges` is
also required. The JSON quote reports these as `continuation.unknown_retry`.

Resume supports one model, one benchmark and one attempt per task on the
reviewed native benchmark families (Inspect, SOB text and lm-eval) without a
judge model, within the source run's evidence retention and on the same
EvalRouter deployment. Comparisons, several benchmarks or attempts, judge-graded
benchmarks, repository agents and your own benchmarks are refused rather than
rerun. Graded tasks from runs that finished before resume evidence was recorded,
including existing runs, cannot gain that evidence later, and results are
never reused across EvalRouter deployments: after a deployment, those tasks are
refused rather than reused or rerun. Tasks whose earlier charge or provider
outcome is unresolved are never called until a fresh quote proves eligibility. A server without resume
says so, and nothing is quoted or charged.

## Lists

`catalog`, `catalog --models`, `status`, `connections list` and `workspace list`
print a compact table sized to your terminal with one suggested next command;
`catalog` hides benchmarks in review or unavailable unless you pass `--all`, and
`catalog -i` browses them interactively; `status -i` keeps your run list live and
opens a run with Enter. Every list accepts `--json`,
`--jq '.data[].id'` (a small jq subset), `-w/--wide`, `--columns a,b`,
`--sort COL` (`--sort=-COL` descending), `--filter KEY=VALUE` and `--limit N`.
Piped output stays JSON. Colour follows `NO_COLOR`, `CI` and `TERM=dumb`.

`run` shows the quote (expected cost from your previous run when known, worst
case, and the smallest workable cap) before asking for the cap, refuses a cap
that cannot fit one task, and asks before repeating a run that just finished
(scripts pass `--new`). `--all-metrics` lists every metric and comparison.

## Discover, quote, then run

```sh
evalrouter catalog --status active --json
evalrouter catalog --models --json
evalrouter catalog --slug BENCHMARK_SLUG --json
```

Choose a profile whose `quote_availability.status` is `ready_for_quote` and a
compatible managed model route from those responses. An active catalog entry
alone does not mean it is ready to run.

These three catalog reads are public, so the CLI sends them without your
credentials: they never read the keychain or refresh a sign-in. If the API
explicitly allows it (`Cache-Control` with `max-age`, and no `no-store`,
`no-cache` or `private`), the CLI keeps a local copy for at most five minutes
and says so on stderr when it uses one. `--refresh` reads the catalog now,
asking caches along the way to revalidate, and fails rather than show an older copy, and `--offline` uses only a fresh local
copy with no network. When a server marks a response `no-store`, the command still makes a request
and no offline copy is retained. Availability shown from a copy is informational: a quote always
checks it again. Workspace data, run status, results, quotes, credits and signed
URLs are never cached. The copy lives in `$XDG_CACHE_HOME/evalrouter` (or
`~/.cache/evalrouter`; `%LOCALAPPDATA%\evalrouter\cache` on Windows, or
`EVALROUTER_CACHE_DIR`), may be deleted at any time, and
`EVALROUTER_READ_CACHE=off` disables it.

Save this as `quote.json`, replacing both uppercase identifiers:

```json
{
  "model": {"kind": "managed", "route_id": "MODEL_ROUTE_ID"},
  "selection": {"profile_ids": ["BENCHMARK_PROFILE_ID"]},
  "coverage": {"mode": "sample", "sample_count": 3, "seed": 42},
  "max_charge_microusd": "1000000"
}
```

```sh
evalrouter quote --config quote.json --json
```

Review compatibility, coverage, cost components, warnings and expiry. A quote
starts no paid work. The example's $1 platform cap is not a price guarantee or
promise that a particular evaluation fits. Paid runs require available credit.

Persist the returned quote ID and an operation key before submitting:

```sh
evalrouter run --quote REVIEWED_QUOTE_ID --idempotency-key SAVED_OPERATION_KEY --json
evalrouter wait RUN_ID --wait-timeout 3600 --json
evalrouter results RUN_ID --json
evalrouter export RUN_ID --format json --output result.json
evalrouter export RUN_ID --format csv --output result.csv
evalrouter export RUN_ID --format html --output result.html
```

Replace `RUN_ID` with the returned ID. A lost submission response is recovered
with the same quote and operation key; a new key may start separate work.
Inspect terminal status, coverage, errors and billing alongside scores. A sample
is not a full-benchmark score. Exports refuse overwrite by default; use the
result's integer `--version` for repeatable reports.

Every command that saves a file prints its full path. An explicit `--output`
file is used exactly and never replaced (exports accept `--overwrite`); a
directory, or no file name, gets a default such as `evalrouter-<run>.json` or
`evalrouter-bundle-<namespace>-<name>-<version>.zip`, moving to the next free
`-2`, `-3`... name instead of replacing anything. Add `--open` or `--reveal`
to open the saved file or show it in your file manager; nothing opens
otherwise, and a missing viewer never fails the command.

### Comparing runs

```sh
evalrouter compare RUN_ID RUN_ID --format html --output my-comparison
```

`compare` takes 1 to 8 finished runs of the same benchmark version, dataset,
task population, task selection and output format in one workspace. Each run
must be a managed or connected model run; comparison, agent and environment
runs are refused, and so is a run whose export does not record one of these
identities. It reads each run's summary export (`export?format=summary`:
identities, frozen native summary, coverage, billing snapshot and export
policy, with no sample rows) and saves `comparison.json` and
`sources/<run>.summary.json` in a new private directory. `--evidence`
downloads each run's full JSON export instead (`sources/<run>.json`) and also
compares grader versions and graded task sets; the scores are the same.
`comparison.json` holds each run's model identity (route, revision, protocol,
recorded limits, generation and routing settings), coverage, errors, its
frozen native summary with the headline's own denominator, and its settled
EvalRouter charge. A run that did not complete, or graded fewer tasks than
planned, is marked partial in `comparison.json` and in its own row of the
terminal table.

For SOB Text runs that were continued (`evalrouter resume`), add
`--include-resumed` (`--chain` is the same flag): each ID then names one
chain, meaning the original run and its resumes, and the chain becomes one
series. The service decides every task from the one run that graded it,
re-derives the SOB native summary with the benchmark authors' code, and keeps
ungraded tasks as errors with their codes. Each score is shown with its runs
("original + 1 resume") and its graded-of-planned count. If a resume replayed
a reused answer and got a different score (`reuse_score_mismatch`), the
report warns and the chain is not complete. `evalrouter export RUN
--include-resumed` saves that chain result by itself. The command follows
each run to its newest resume at the time you run it, so a later resume makes
it return a longer chain; the saved chain documents keep the exact version.

`sources/` keeps each export byte for byte, with its SHA-256 recorded. With
`--evidence`, every graded sample must carry its scored answer by default: the final completion
text for SOB, plus the per-sample metadata the export supplies (for SOB, the
response's token counts and finish reason). Exports do not contain complete
transcripts, reasoning, timings or per-attempt usage. Keep these files
private. Use `--without-raw` (which implies `--evidence`) to archive runs
whose answers have expired or are withheld by the export policy. A chain
summary report downloads member exports only with `--evidence`. Different transports, limits or generation
settings are listed as warnings. Nothing is regraded, and no deltas,
intervals or winner are computed. `--output NAME.json` puts the sources in
`NAME-sources/`.

## Supported commands

`catalog` (`--models`, `--slug` or `--environment`), `connections list/create/check/update/disable`,
`quote`, `run`, `resume`, `status`, `wait`, `cancel`, `results`, `export`, `compare` and `open`.
Authoring your own benchmark package uses `benchmark init/validate/bundle`.
`credits` shows the workspace balance, the amount held by running evaluations
and what is available for new ones. It works with an `evalrouter login` sign-in
or an API key with the `billing:read` scope. Credit is added on the website under
Settings > Billing; the CLI never takes a payment.
Run each command with `--help` for exact flags. A connected model's credential
is never a command argument: `connections create` asks with hidden input in a
terminal, and scripts use `--key-stdin` or `--key-env NAME`
(`connections update --rotate-key` replaces it). Connection checks can send
a small model request, and your model provider may bill separately.

For `wait` and `run --wait`, `--progress auto` shows elapsed time and status in
an interactive terminal, with a percentage when validated processed and planned
counts are available. Use `--progress plain` for structured progress or
`--progress off` to suppress it. Redirected stderr and `--json` retain structured
progress. A progress percentage counts processed samples; it is not a score.
`--progress json` writes `evalrouter.progress.v1` records on stderr, on every
change and every 30 seconds: phase, processed and planned counts, errors,
spend and elapsed time. `evalrouter events RUN --follow` prints the run's
event log as JSON lines until it ends; `--type`, `--tail N` and `--after N`
filter, start from the end and resume.
`evalrouter open RUN` opens the run's live report page, also when an agent
runs it with `--json`; `run --open` opens it as soon as the run starts. No
browser opens in CI, over SSH, without a display, with `--no-browser` or with
`EVALROUTER_NO_BROWSER=1`; the address is printed instead.

Use `--config -` for JSON on stdin. `--json` writes one result on stdout and
progress on stderr. Exit codes: 0 success; 1 API/transport or failed run;
2 invalid input or rejected request (including a refused unattended `run`);
4 partial/cancelled waited run; 5 confirmation declined, nothing started;
6 stopped watching because the API stayed unreachable. Interrupting
a local wait does not cancel server work. Use `evalrouter cancel RUN_ID`
explicitly when cancellation is intended.

See the [CLI documentation](https://evalrouter.ai/developers/docs) for account
setup, model selection, spending semantics, recovery and versioned exports.
The CLI uses the [model API](https://evalrouter.ai/developers/docs/model-api),
which is also documented for direct HTTP integrations.

## Author your own benchmark

`benchmark init`, `benchmark validate` and `benchmark bundle` run entirely on
your machine: no account, no network request, and nothing from your package is
imported or executed. The normative validator ships inside `evalrouter`, and
its one dependency (pydantic) is an optional extra so the base installation
stays thin (httpx and rich only). With pip, use `pip install 'evalrouter[benchmark]'`:

```sh
uv tool install --python 3.12 "evalrouter[benchmark]"
evalrouter benchmark init ./my-benchmark --namespace example --name my-benchmark
evalrouter benchmark validate ./my-benchmark
evalrouter benchmark bundle ./my-benchmark   # evalrouter-bundle-example-my-benchmark-<version>.zip
```

Without the extra these three commands stop before reading your package and
report `benchmark_extra_required` with the exact install string; `--help` and
every network command keep working. `init` writes an explicitly synthetic draft
and claims no license: replace the tasks, rights, references and limits before
submitting. Local validation is not platform admission — EvalRouter revalidates
every submission on its own servers.

## Live agent capture

Install `evalrouter[tracing]` and use `from kimpton_evalrouter.tracing import TraceCapture`.
Create a connection in Traces and configure its restricted ingest key with
`EVALROUTER_TRACE_KEY` and `EVALROUTER_TRACE_CONNECTION_ID`, alongside the API
origin and workspace ID. Wrap each run with `capture.agent(...)`, and calls with
`capture.model(...)` or `capture.tool(...)`. Use `span.input(value)` and
`span.output(value)` to capture content. Call `capture.flush()` before a
serverless invocation exits and `capture.close()` on shutdown. The optional
OpenTelemetry dependency is isolated from the base HTTP client.

See [the capture guide](https://evalrouter.ai/developers/docs/live-traces) for complete
examples, redaction settings, coverage, retention, and delivery limits.

Use `client.traces` for connection, trace, evaluator and finding resources, and
`client.datasets` for curation, review, immutable versions and exports.
These operations require the corresponding service release and workspace permissions.
