Metadata-Version: 2.5
Name: benchspec
Version: 0.0.4
Summary: benchspec is a framework for evaluating AI agents with repeatable, isolated benchmarks. Write evals as Markdown, run each eval across named benchmark arms, and compare how agent behavior changes by harness, model, effort, and environment.
Project-URL: Homepage, https://github.com/theycallmeswift/benchspec
Project-URL: Documentation, https://github.com/theycallmeswift/benchspec/tree/dev/docs
Project-URL: Issues, https://github.com/theycallmeswift/benchspec/issues
Author-email: Swift <swift@majorleaguehacking.com>
License: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Framework :: Pytest
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.11
Requires-Dist: pytest>=8
Requires-Dist: python-dotenv>=1.0
Requires-Dist: pyyaml>=6
Provides-Extra: claude
Provides-Extra: codex
Provides-Extra: microsandbox
Requires-Dist: microsandbox<0.7,>=0.6.16; extra == 'microsandbox'
Provides-Extra: opencode
Provides-Extra: testing
Description-Content-Type: text/markdown

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="docs/assets/benchspec-wordmark-dark.svg">
  <img alt="benchspec" src="docs/assets/benchspec-wordmark-light.svg" width="188" height="48">
</picture>

**Benchmark what your agent does, not what it says.**

[![CI](https://github.com/theycallmeswift/benchspec/actions/workflows/ci.yml/badge.svg)](https://github.com/theycallmeswift/benchspec/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/benchspec)](https://pypi.org/project/benchspec/)
[![Python](https://img.shields.io/pypi/pyversions/benchspec)](https://pypi.org/project/benchspec/)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)

benchspec runs an agent (Claude Code, Codex, or OpenCode) against a task
in a fresh sandbox, checks what it actually did in the workspace, and reports the
result as a comparison: with your skill versus without, one model versus another,
one harness versus another.

<img src="docs/assets/benchmark-terminal.gif" width="800" alt="Animated terminal output: benchspec run prints a benchmark matrix with evals as rows, arms as columns, color-coded rates, and percentage-point deltas">

*End-of-run summary for the in-repo [`hello`](evals/e2e/hello/) suite: the
`baseline` and `trial` columns of a real three-sample run.* Rows are evals, columns
are arms (`baseline` ran the agent bare, `trial` installed the skill), and every
non-baseline cell shows its assertion pass rate plus the delta against the
baseline in percentage points. The scoped `` Skill `hello` invoked `` line is in
neither column. The same matrix lands in `benchmark.md`, with machine-readable
artifacts alongside.

Teams pick harnesses, models, and prompts by anecdote: run it once, eyeball the
transcript, trust the vibe. benchspec turns that guess into a measurement.
Write the goal once, run it across the configurations you care about, and read
off — in percentage points — how good each one actually is at accomplishing it.

> **When to use benchspec.** If the only question is whether one plugin helps
> Claude Code, `claude plugin eval` ships inside Claude Code and answers it on
> one credential, no container. The
> [comparison below](#benchspec-vs-the-alternatives) routes the other cases.

## What you need

| | |
|---|---|
| Platform | Any OS with a Docker daemon (default). Apple Silicon or Linux with `/dev/kvm` for the microsandbox opt-in. Python 3.11+. |
| Agent CLI | `claude`, `codex`, or `opencode` on `PATH`, with its credential (for Claude Code, `CLAUDE_CODE_OAUTH_TOKEN` or `ANTHROPIC_API_KEY`). |
| Binder credential | The binder is a fixed model call that classifies each assertion, required for every `analyze` and `run`: `GEMINI_API_KEY` by default, or `OPENROUTER_API_KEY` with `[tool.benchspec.binder] provider = "openrouter"`. |
| Judge credential | The judge runs on the host through an agent CLI, using either its env credential or the CLI's own login (`claude login`, `codex login`); the default is `claude-code` with `sonnet`. Prefer a different vendor from the arms (this repo's own suite judges Claude arms with Codex). |

Credentials can live in a repo-root `.env`. A graded run can touch up to three
vendors: the agent's, the binder's, and the judge's. Or exactly one: set
`provider = "openrouter"` on the binder, the judge, and the arms, and the whole
run needs only `OPENROUTER_API_KEY` (see
[`configuration.md`](docs/configuration.md#providers)). `lint` is free;
`analyze` and `run` spend API calls, and `run` also boots sandboxes. Preflight
lists every missing piece and exits before anything is spent.

## Getting Started

```bash
pip install benchspec
```

For the microsandbox isolation opt-in, add its extra:

```bash
pip install "benchspec[microsandbox]"
```

> **Pre-1.0.** The eval format and the artifact schemas are the surfaces most
> likely to change.

An eval is one Markdown file: a prompt, then a checklist of plain-prose claims
about the workspace after the agent is done. There is no checker syntax to learn;
the wording is the spec. `evals/hello/greets-by-name.eval.md`:

```markdown
---
---

## Prompt

You are working in a workspace rooted at your current working directory.
Greet Alice by name.

## Assertions

- [ ] ./Greetings/Alice.md contains the exact line 'Hello, Alice!'
- [ ] Skill `hello` invoked
  - if: {BENCHSPEC_ARM} != {BENCHSPEC_BASELINE}
- [ ] The greeting feels warm and personable, not curt or robotic
```

The benchmark is a block in `pyproject.toml`. Arms are the report columns; the
baseline is what the others are measured against:

```toml
[tool.benchspec]
default-set = "default"

[tool.benchspec.sets.default]
harness  = "claude-code"
model    = "sonnet"
baseline = "baseline"
arms = [
  { name = "baseline" },   # installs nothing
  { name = "trial" },      # setup.sh installs the skill
]
```

Then:

```bash
benchspec lint      # static checks on the assertions
benchspec analyze   # which assertions grade deterministically, which go to the judge
benchspec run       # every (eval × arm) in its own sandbox, graded, reported
```

The `` Skill `hello` invoked `` line needs a skill for the trial arm to install
and a `setup.sh` that installs it; the [quickstart](docs/quickstart.md) writes
both and takes an empty directory to that first graded report. The `- if:`
sub-bullet keeps that line out of the baseline's rate, so the delta measures the
skill ([scoping](docs/writing-evals.md#scoping-an-assertion-to-arms)).

## How a run works

```mermaid
flowchart LR
    E["greets-by-name.eval.md<br/>prompt + assertions"] --> A1["arm: baseline<br/>fresh sandbox,<br/>setup.sh installs nothing"]
    E --> A2["arm: trial<br/>fresh sandbox,<br/>setup.sh installs the hello skill"]
    A1 --> F1["facts: files, SHAs,<br/>final message, tool calls"]
    A2 --> F2["facts"]
    F1 --> G["binder: deterministic checkers<br/>everything else: LLM judge"]
    F2 --> G
    G --> R["benchmark.md + benchmark.json<br/>meta.json + index.jsonl"]
```

Each `(eval × arm)` pair is one parametrized pytest test. A cell:

1. **Boots a sandbox** — a Docker container by default, or a microsandbox
   microVM for the opt-in — from a cached snapshot with the agent CLI already
   installed. The first run builds the snapshot (about a minute); later runs
   reuse it. `benchspec sandbox:build` pays that cost up front,
   `benchspec sandbox:clean` reclaims the disk.
2. **Seeds the clean room** — the eval's optional `workspace/` files land in a
   fresh directory mounted at `/workspace`, the agent's working directory.
3. **Runs `setup.sh`**, where arms diverge: it sees `$BENCHSPEC_ARM`, so the
   baseline branch exits early and the trial branch copies the skill into place.
4. **Invokes the agent** on the eval's prompt.
5. **Collects the facts** — file tree, contents, SHA-256s, the final message,
   the tool calls.
6. **Grades** — the binder maps each assertion to a deterministic checker where
   it can do so without risk; the judge grades everything else from the
   collected evidence alone.

On either backend, nothing in the guest can write back to your checkout:
`setup.sh` reaches the skill under test through a read-only staged copy of your
repo at `/project` (what a `git clone` would contain — never `.env`, `.git`, or
earlier runs' artifacts). Credential exposure differs — microsandbox injects each
credential at the network boundary, Docker as a plain container environment
variable the agent can read. [`sandbox.md`](docs/sandbox.md) has the tradeoff.

`benchspec run` is pytest underneath, and everything after `--` goes to
pytest verbatim. The normal run is `benchspec run -- --count 3`: three samples
per cell (pytest-repeat), so every delta carries a noise band; a one-sample run
is flagged in the report. `-n 8` fans cells across eight sandboxes to keep it
fast, and `-k greets-by-name` (equivalently `pytest -k greets-by-name`) runs one
eval. The repo's own `make e2e` defaults to three samples and six workers through
the `COUNT` and `WORKERS` variables; `make e2e COUNT=1 WORKERS=1` is the quick
sequential pass.

## benchspec vs. the alternatives

Where each alternative is the better answer:

| | Reach for it when | What benchspec adds |
|---|---|---|
| [`claude plugin eval`](https://code.claude.com/docs/en/plugin-evals) | The question is whether one plugin helps Claude Code. It ships inside Claude Code, needs one credential and no container, and interviews you to write the suite. | A second harness, arms beyond with-and-without, a workspace seeded and hashed before the run, and a judge that need not share the arms' provider. |
| [Inspect AI](https://inspect.aisi.org.uk/) | You write Python, and you want off-the-shelf evals, remote execution at scale, or `pass@k` reducers. | The eval is a Markdown file of prose claims rather than a scorer you implement, and the report is an arms-versus-baseline delta without assembling one. |
| [Harbor](https://www.harborframework.com/) / Terminal-Bench | You want to rank agents on a standard published benchmark, with a long list of agents already integrated. | Your tasks, your baseline, your delta. |
| [Coder Eval](https://github.com/UiPath/coder_eval) | You want typed YAML criteria with weights and fractional credit, or its GitHub Action. | Prose assertions instead of a criterion schema, and pre-run hashes behind `left unchanged`. |
| [promptfoo](https://www.promptfoo.dev/) | What you are grading is a prompt and the reply it produced. | Grading of the workspace the agent left behind, not the response it wrote about it. |

benchspec's own cost is the top of this README: a Docker daemon, and up to
three vendors' credentials. Every run pays that back in artifacts —
`meta.json`, `index.jsonl`, `benchmark.json` — that another tool can aggregate
without knowing the directory layout.

[`docs/research/2026-09-13-alternatives-landscape.md`](docs/research/2026-09-13-alternatives-landscape.md)
has the long version: unit of evaluation, isolation, and statistics per tool.

## Documentation

| | |
|---|---|
| [`docs/quickstart.md`](docs/quickstart.md) | Empty directory to a graded two-arm run. |
| [`docs/concepts.md`](docs/concepts.md) | The vocabulary: eval, arm, set, baseline, binder, judge. |
| [`docs/writing-evals.md`](docs/writing-evals.md) | The eval format, workspaces, `setup.sh`, and how grading decides what binds. |
| [`docs/configuration.md`](docs/configuration.md) | Sets, arms, the judge, every CLI flag, exit codes. |
| [`docs/sandbox.md`](docs/sandbox.md) | Snapshots, mounts, credentials, host requirements. |
| [`docs/results.md`](docs/results.md) | Reading `benchmark.md` and the machine-readable artifacts. |
| [`docs/harnesses.md`](docs/harnesses.md) | `claude-code`, `codex`, `opencode`, and adding your own. |

## Development

Clone to a graded run of the in-repo suite in four steps. Steps 1 and 2 cost
nothing; step 3 is the first paid call.

```bash
git clone https://github.com/theycallmeswift/benchspec && cd benchspec
```

1. **Install.** [`uv`](https://docs.astral.sh/uv/) manages the venv;
   `make install` runs `uv sync`. Then the free checks:

   ```bash
   curl -LsSf https://astral.sh/uv/install.sh | sh   # if you don't have uv
   make install
   make test     # unit suite: no credentials, no sandbox
   ```

2. **Confirm the suite collects.** Needs no credentials and no Docker:

   ```bash
   make e2e EVAL_ARGS="--collect-only -q"   # 27 cells: 3 evals × 3 arms × 3 samples, then 36 through OpenRouter
   ```

3. **Set up the credentials.** `make e2e` runs [`evals/e2e/hello/`](evals/e2e/hello/)
   twice: three evals across three Claude Code arms, judged by Codex; then the same
   evals across Claude Code, Codex, and OpenCode arms with the binder, the judge,
   and every arm on OpenRouter. It needs `claude` and `codex` on `PATH`, a running
   Docker daemon, and four credentials in `.env`:

   ```bash
   cp .env.example .env
   ```

   | Variable | For |
   |---|---|
   | `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`) or `ANTHROPIC_API_KEY` | the arms |
   | `GEMINI_API_KEY` | the binder |
   | `OPENAI_API_KEY` | the Codex judge |
   | `OPENROUTER_API_KEY` | the second run: binder, judge, and arms through OpenRouter |

   Preflight lists every missing piece in one message and exits before anything
   is spent.

4. **Run it.** The cheapest real run is one eval, one sandbox:

   ```bash
   make e2e COUNT=1 WORKERS=1 EVAL_ARGS="-k greets-by-name"   # one eval, one sample, sequential
   make e2e                                                  # the whole suite, three samples, six sandboxes
   ```

   The first run builds the sandbox snapshot (about a minute); with the snapshot
   cached, the whole suite takes about a minute on six workers. The report lands
   in `tmp/evals/iteration_01/benchmark.md`.

Then `make lint` (ruff, ty, houserules; needs `GEMINI_API_KEY`) before a pull
request. On Claude Code on the web, the environment's setup script does steps 1
and 3 once for every session; see
[`sandbox.md`](docs/sandbox.md#claude-code-on-the-web).

Issues and pull requests are welcome at
[github.com/theycallmeswift/benchspec](https://github.com/theycallmeswift/benchspec).
Until `CONTRIBUTING.md` lands, [`docs/style/development.md`](docs/style/development.md)
is the code style, and `make test` plus `make lint` are the bar.

## License

[MIT](LICENSE).
