Metadata-Version: 2.4
Name: evalrouter
Version: 0.3.0
Summary: Typed Kimpton evaluation client and command-line interface
License-Expression: LicenseRef-Kimpton-EvalRouter-SDK
License-File: LICENSE
Requires-Python: <3.14,>=3.12
Requires-Dist: httpx==0.28.1
Requires-Dist: rich==15.0.0
Provides-Extra: benchmark
Requires-Dist: pydantic==2.13.5; extra == 'benchmark'
Provides-Extra: keyring
Requires-Dist: keyring==25.7.0; extra == 'keyring'
Description-Content-Type: text/markdown

# EvalRouter CLI

Evaluate models or qualified repository agents from your terminal: discover
supported targets, review a spending cap, run an evaluation, and export results. Requires Python 3.12
or 3.13. Distributed under the proprietary [Kimpton EvalRouter SDK License](LICENSE).

## Install and create your account

We recommend [uv](https://docs.astral.sh/uv/getting-started/installation/) to install the CLI in its own Python environment:

```sh
uv tool install --python 3.12 evalrouter
evalrouter --help
```

With pip, use `pip install evalrouter`. uv supplies Python 3.12 if needed. If
your shell cannot find `evalrouter`, run `uv tool update-shell` and restart your
terminal. Upgrade an existing installation with `uv tool upgrade evalrouter`.

This package was previously published as `kimpton-evalrouter-sdk`. The command,
the `kimpton_evalrouter` import and every API are unchanged. To switch, run
`uv tool uninstall kimpton-evalrouter-sdk` (or
`pip uninstall -y kimpton-evalrouter-sdk evalrouter`) and install `evalrouter`.

The package supplies the `evalrouter` command; no repository checkout is needed.
[Create an EvalRouter account](https://evalrouter.ai/register), verify your email,
and create a workspace key in [API key settings](https://evalrouter.ai/settings/api-keys).
Package installation does not grant account access or evaluation credit.

Set `EVALROUTER_API_KEY` and `EVALROUTER_WORKSPACE_ID` through your secret manager
or private environment. Use `https://api.evalrouter.ai` for `EVALROUTER_BASE_URL`,
without `/v1`. Never put credentials in arguments, URLs, request files or logs.
The CLI uses an existing workspace key; browser sign-up is a separate step.

## Run an evaluation in one command

```sh
evalrouter run gpqa-diamond --model gpt-4o-mini
```

`run` matches the benchmark and model by slug, name or ID, suggesting the
closest names after a typo. It creates a free quote, shows the coverage,
conservative estimate, spending cap and concurrency, and asks `Run it? [Y/n]`.
Only a yes starts paid work. It then follows progress and prints the scores,
the change from your previous completed run of the same benchmark and model,
and the report link. In a terminal, missing choices are asked for; run
`evalrouter run` alone to pick everything. `--sample N` or `--full` sets
coverage, `--max-cost USD` the cap and `--dry-run` stops after the quote.

Without a terminal (CI, agents, `--json` or `--no-input`) `run` never prompts
and refuses unless both `--yes` and `--max-cost` are given:

```sh
evalrouter run gpqa-diamond --model gpt-4o-mini --max-cost 5 --yes --json
```

The idempotency key of each accepted quote is saved locally before the run is
submitted (`~/.local/state/evalrouter`, or `EVALROUTER_STATE_DIR`; identifiers
only, never credentials). Repeating the same command after a dropped connection
resumes that run instead of starting another; `--new` starts a separate one.
Afterwards, `results`, `export --format html,json`, `wait` and `open` default to
your last run and accept a unique run ID prefix. Watching a run survives brief
API outages such as a 503 during a deployment.

## Lists

`catalog`, `catalog --models`, `status`, `connections list` and `workspace list`
print a compact table sized to your terminal with one suggested next command;
`catalog` hides benchmarks in review or unavailable unless you pass `--all`, and
`catalog -i` browses them interactively. Every list accepts `--json`,
`--jq '.data[].id'` (a small jq subset), `-w/--wide`, `--columns a,b`,
`--sort COL` (`--sort=-COL` descending), `--filter KEY=VALUE` and `--limit N`.
Piped output stays JSON. Colour follows `NO_COLOR`, `CI` and `TERM=dumb`.

`run` shows the quote (expected cost from your previous run when known, worst
case, and the smallest workable cap) before asking for the cap, refuses a cap
that cannot fit one task, and asks before repeating a run that just finished
(scripts pass `--new`). `--all-metrics` lists every metric and comparison.

## Discover, quote, then run

```sh
evalrouter catalog --status active --json
evalrouter catalog --models --json
evalrouter catalog --slug BENCHMARK_SLUG --json
```

Choose a profile whose `quote_availability.status` is `ready_for_quote` and a
compatible managed model route from those responses. An active catalog entry
alone does not mean it is ready to run.
Save this as `quote.json`, replacing both uppercase identifiers:

```json
{
  "model": {"kind": "managed", "route_id": "MODEL_ROUTE_ID"},
  "selection": {"profile_ids": ["BENCHMARK_PROFILE_ID"]},
  "coverage": {"mode": "sample", "sample_count": 3, "seed": 42},
  "max_charge_microusd": "1000000"
}
```

```sh
evalrouter quote --config quote.json --json
```

Review compatibility, coverage, cost components, warnings and expiry. A quote
starts no paid work. The example's $1 platform cap is not a price guarantee or
promise that a particular evaluation fits. Paid runs require available credit.

Persist the returned quote ID and an operation key before submitting:

```sh
evalrouter run --quote REVIEWED_QUOTE_ID --idempotency-key SAVED_OPERATION_KEY --json
evalrouter wait RUN_ID --wait-timeout 3600 --json
evalrouter results RUN_ID --json
evalrouter export RUN_ID --format json --output result.json
evalrouter export RUN_ID --format csv --output result.csv
evalrouter export RUN_ID --format html --output result.html
```

Replace `RUN_ID` with the returned ID. A lost submission response is recovered
with the same quote and operation key; a new key may start separate work.
Inspect terminal status, coverage, errors and billing alongside scores. A sample
is not a full-benchmark score. Exports refuse overwrite by default; use the
result's integer `--version` for repeatable reports.

## Supported commands

`catalog` (`--models`, `--slug` or `--environment`), `connections list/create/check/update/disable`,
`quote`, `run`, `status`, `wait`, `cancel`, `results`, `export` and `open`.
Repository qualification uses `agent-builds options/preview/create/list/status/cancel`.
Authoring your own benchmark package uses `benchmark init/validate/bundle`.
Run each command with `--help` for exact flags. A connected model's credential
is never a command argument: `connections create` asks with hidden input in a
terminal, and scripts use `--key-stdin` or `--key-env NAME`
(`connections update --rotate-key` replaces it). Connection checks can send
a small model request, and your model provider may bill separately.

For `wait` and `run --wait`, `--progress auto` shows elapsed time and status in
an interactive terminal, with a percentage when validated processed and planned
counts are available. Use `--progress plain` for structured progress or
`--progress off` to suppress it. Redirected stderr and `--json` retain structured
progress. A progress percentage counts processed samples; it is not a score.

Use `--config -` for JSON on stdin. `--json` writes one result on stdout and
progress on stderr. Exit codes: 0 success; 1 API/transport or failed run;
2 invalid input or rejected request (including a refused unattended `run`);
4 partial/cancelled waited run; 5 confirmation declined, nothing started. Interrupting
a local wait does not cancel server work. Use `evalrouter cancel RUN_ID`
explicitly when cancellation is intended.

See the [CLI documentation](https://evalrouter.ai/developers/docs) for account
setup, model selection, spending semantics, recovery and versioned exports.
The CLI uses the [model API](https://evalrouter.ai/developers/docs/model-api),
which is also documented for direct HTTP integrations.

## Author your own benchmark

`benchmark init`, `benchmark validate` and `benchmark bundle` run entirely on
your machine: no account, no network request, and nothing from your package is
imported or executed. The normative validator ships inside `evalrouter`, and
its one dependency (pydantic) is an optional extra so the base installation
stays thin (httpx and rich only). With pip, use `pip install 'evalrouter[benchmark]'`:

```sh
uv tool install --python 3.12 "evalrouter[benchmark]"
evalrouter benchmark init ./my-benchmark --namespace example --name my-benchmark
evalrouter benchmark validate ./my-benchmark
evalrouter benchmark bundle ./my-benchmark --output my-benchmark.zip
```

Without the extra these three commands stop before reading your package and
report `benchmark_extra_required` with the exact install string; `--help` and
every network command keep working. `init` writes an explicitly synthetic draft
and claims no license: replace the tasks, rights, references and limits before
submitting. Local validation is not platform admission — EvalRouter revalidates
every submission on its own servers.

## Repository agents

Repository commands require version 0.2.0 or later and a service that has enabled
repository qualification for the workspace. Installing the package does not
enable that service. Check its current availability and supported targets first:

```sh
evalrouter agent-builds options --json
evalrouter catalog --environment ENVIRONMENT_FAMILY_PATH --json
```

The catalog lookup takes the family path before `@`; qualification inputs still
require the exact versioned reference returned by options. Stop if qualification
is disabled or the desired exact target/model is absent.
The supported runtime is locked Node/npm with an `evalrouter-agent.json` manifest;
custom images and Python agent runtimes are not supported. Save `build.json`:

```json
{
  "repository_url": "YOUR_PUBLIC_GITHUB_REPOSITORY_URL",
  "ref": "main",
  "manifest_path": "evalrouter-agent.json",
  "qualification": {
    "environment_ref": "EXACT_ENVIRONMENT_FROM_OPTIONS",
    "model": {"kind": "managed", "route_id": "MODEL_FROM_OPTIONS"}
  },
  "max_cost_microusd": "1000000"
}
```

For a private repository, connect GitHub in the web application and explicitly
select that repository for this workspace. Add
`source: {"github_connection_id": "CONNECTION_ID", "github_repository_id": 123}`
using the two actual IDs from that connection. These IDs select a grant; they
are not credentials. The CLI does not accept GitHub tokens or authorize access.

```sh
evalrouter agent-builds preview --config build.json --json
```

Preview resolves an immutable commit and returns the qualification scope, costs,
cap and expiry. It creates no job or credit hold and starts no sandbox/model
work. Review it, then save the quote ID and a durable operation key before:

```sh
evalrouter agent-builds create --quote REVIEWED_BUILD_QUOTE --idempotency-key SAVED_BUILD_KEY --json
evalrouter agent-builds status BUILD_JOB_ID --json
evalrouter agent-builds list --quote REVIEWED_BUILD_QUOTE --json
```

Creation reserves the quoted cap and may incur its stated costs. Recover an
unknown submission with the same quote and key or the quote-filtered list;
a new key may create new work. Cancel explicitly with
`evalrouter agent-builds cancel BUILD_JOB_ID`. A `cancelling` or `reconciling`
status remains unresolved. Wait for terminal cleanup and zero held funds.

Ready requires `status: ready`, `cleanup_confirmed: true`, zero held funds and
an `agent.ref`. It establishes only the returned scope, not benchmark quality.
Use that record's exact `qualification.environment_ref`, `qualification.split`
and `qualification.evaluation_coverage` in a separate evaluation quote:

```json
{
  "agent": "READY_AGENT_REF",
  "selection": {"environment": "RETURNED_ENVIRONMENT_REF", "split": "RETURNED_SPLIT"},
  "coverage": "REPLACE_WITH_RETURNED_EVALUATION_COVERAGE_OBJECT",
  "max_charge_microusd": "1000000"
}
```

Replace the coverage placeholder with the entire returned object; do not guess
tasks, counts or seeds. Use the existing `quote`, `run`, `wait`, `results` and
`export` commands above. Qualification and evaluation have separate caps and
charges. A qualification pass does not submit an evaluation automatically.
