Benchmark report / {{RUN}} / executed {{DATE}}
{{TITLE}}
{{ONE_PARAGRAPH: what this run was for, and what it is not. A benchmark rules no threshold.}}
Arms
{{ARM_A}} → {{ARM_B}}
{{how each was installed}}
Corpora
{{CORPORA}}
{{tiers run, and tiers not run}}
Questions
{{N}}
{{judged / timing / planted unanswerable}}
Classification
{{CLASSIFICATION}}
no threshold ruled — SR-WORK-BENCHMARK decision 6
How to read this report / every number carries its direction
Which way is good, stated on every metric
| marker | means | metrics it sits on in this run |
|---|---|---|
| ↑ higher is better | a bigger number is a better engine | hit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable |
| ↓ lower is better | a smaller number is a better engine | query p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated |
| — neither | a change is a signal, not a score. Nothing here says which value is better — a person reads the rows | queries whose list moved · first differing rank · shard count · headroom · b and c |
“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.
A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.
The arms / evidence/ARMS.toml
Two engines, byte‑checked against one corpus
| arm A | arm B | |
|---|---|---|
| version | {{ARM_A}} | {{ARM_B}} |
| install | {{...}} | {{...}} |
| python | {{...}} | {{...}} |
| corpus sha256 | {{...}} | {{... identical, or the run is void}} |
| queries sha256 | {{...}} | {{...}} |
| enrichment | {{present|absent}} | {{present|absent}} |
{{Anything that makes this run non-reproducible -- an editable install of a dirty tree, an uncommitted change, a machine that was not quiet -- is stated HERE, not discovered later.}}
The null control / run first, as it always is
Arm A against itself
nullcontrol · {{corpus}} · {{path}}
{{n}} differed↓ lower is better
{{Identical ranked lists on every query, or the run stops here.}}
Why it is first
Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.
CAP‑3 / hit@k against the planted key
{{the one-line finding, or: no number exists for this run}}
| arm | version | hit@1↑ higher is better | hit@5↑ higher is better | hit@10↑ higher is better | hit@20↑ higher is better | hit@50↑ higher is better |
|---|---|---|---|---|---|---|
| A | {{ARM_A}} | {{}} | {{}} | {{}} | {{}} | {{}} |
| B | {{ARM_B}} | {{}} | {{}} | {{}} | {{}} | {{}} |
hit@5 is the headline; the other four are context and are never dropped for being undramatic.
CAP‑3 / headroom, stated in both directions
How much could have moved at all
One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.
| k | n | improvement headroom— neither | regression headroom— neither | b A wrong → B right↑ higher is better | c A right → B wrong↓ lower is better |
|---|---|---|---|---|---|
| 1 | {{}} | {{}} | {{}} | {{}} | {{}} |
| 5 | {{}} | {{}} | {{}} | {{}} | {{}} |
| 10 | {{}} | {{}} | {{}} | {{}} | {{}} |
Zero headroom is not agreement. Where the improvement column is 0 no delta is representable, so equal scores there report a ceiling, never a match. Say which rows are ceilings.
CAP‑4 / the answer layer, and the planted unanswerables
{{the one-line finding}}
| arm | answered— neither | declined↑ higher is better | fabricated↓ lower is better |
|---|---|---|---|
| A {{ARM_A}} | {{}} | {{}} | {{}} |
| B {{ARM_B}} | {{}} | {{}} | {{}} |
fabricated = answered a question planted to have no answer. Higher declined is better only on the planted unanswerables — on the judged set, declining is a miss. Say which set each column is over.
CAP‑1 · CAP‑2 / the ranked lists, and what moved between them
{{n}} of {{N}} lists differ — {{and where}}
{{entered / left counts, and the one sentence a reader should take away}}
CAP‑5 / the committed index
{{one line}}
| arm | index bytes↓ lower is better | bytes / document↓ lower is better | shards— neither |
|---|---|---|---|
| A | {{}} | {{}} | {{}} |
| B | {{}} | {{}} | {{}} |
Deterministic — unaffected by what else the machine was doing.
CAP‑6 / speed, arms interleaved A B A B
{{one line}} {{do not quote when the machine was not quiet}}
| arm | query p50↓ lower is better | query p95↓ lower is better | ingest↓ lower is better | build↓ lower is better |
|---|---|---|---|---|
| A | {{}} | {{}} | {{}} | {{}} |
| B | {{}} | {{}} | {{}} | {{}} |
State whether the machine was quiet. A loaded machine does not produce noise — it produces a clean, localised anomaly that reads like a finding. Interleaving A B A B protects the difference and never the absolute number.
Guard rails / what this run may never be used to say
What this run does not do
It rules no threshold
SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.
{{scope limit}}
{{tiers not run, paths not run, arms not compared}}
{{classification limit}}
{{what informed costs this run, or what blind buys it}}
{{captures with no number}}
{{named here as well as in their own section, so the gaps are countable from one slide}}
The run is work/regression/{{RUN}}/ — report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.