Benchmark report / {{RUN}} / executed {{DATE}}

{{TITLE}}

{{ONE_PARAGRAPH: what this run was for, and what it is not. A benchmark rules no threshold.}}


Arms

{{ARM_A}} → {{ARM_B}}

{{how each was installed}}

Corpora

{{CORPORA}}

{{tiers run, and tiers not run}}

Questions

{{N}}

{{judged / timing / planted unanswerable}}

Classification

{{CLASSIFICATION}}

no threshold ruled — SR-WORK-BENCHMARK decision 6

How to read this report / every number carries its direction

Which way is good, stated on every metric

markermeansmetrics it sits on in this run
↑ higher is bettera bigger number is a better enginehit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable
↓ lower is bettera smaller number is a better enginequery p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated
— neithera change is a signal, not a score. Nothing here says which value is better — a person reads the rowsqueries whose list moved · first differing rank · shard count · headroom · b and c

“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.

A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.

The arms / evidence/ARMS.toml

Two engines, byte‑checked against one corpus

 arm Aarm B
version{{ARM_A}}{{ARM_B}}
install{{...}}{{...}}
python{{...}}{{...}}
corpus sha256{{...}}{{... identical, or the run is void}}
queries sha256{{...}}{{...}}
enrichment{{present|absent}}{{present|absent}}

{{Anything that makes this run non-reproducible -- an editable install of a dirty tree, an uncommitted change, a machine that was not quiet -- is stated HERE, not discovered later.}}

The null control / run first, as it always is

Arm A against itself

nullcontrol · {{corpus}} · {{path}}

{{n}} differed↓ lower is better

{{Identical ranked lists on every query, or the run stops here.}}

Why it is first

Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.

CAP‑3 / hit@k against the planted key

{{the one-line finding, or: no number exists for this run}}

armversion hit@1↑ higher is better hit@5↑ higher is better hit@10↑ higher is better hit@20↑ higher is better hit@50↑ higher is better
A{{ARM_A}}{{}}{{}}{{}}{{}}{{}}
B{{ARM_B}}{{}}{{}}{{}}{{}}{{}}

hit@5 is the headline; the other four are context and are never dropped for being undramatic.

CAP‑3 / headroom, stated in both directions

How much could have moved at all

One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.

kn improvement headroom— neither regression headroom— neither b  A wrong → B right↑ higher is better c  A right → B wrong↓ lower is better
1{{}}{{}}{{}}{{}}{{}}
5{{}}{{}}{{}}{{}}{{}}
10{{}}{{}}{{}}{{}}{{}}

Zero headroom is not agreement. Where the improvement column is 0 no delta is representable, so equal scores there report a ceiling, never a match. Say which rows are ceilings.

CAP‑4 / the answer layer, and the planted unanswerables

{{the one-line finding}}

arm answered— neither declined↑ higher is better fabricated↓ lower is better
A  {{ARM_A}}{{}}{{}}{{}}
B  {{ARM_B}}{{}}{{}}{{}}

fabricated = answered a question planted to have no answer. Higher declined is better only on the planted unanswerables — on the judged set, declining is a miss. Say which set each column is over.

{{Any asymmetry between the arms -- a flag one version does not have -- is stated, not smoothed.}}

CAP‑1 · CAP‑2 / the ranked lists, and what moved between them

{{n}} of {{N}} lists differ — {{and where}}

{{ INLINE SVG: first-differing-rank histogram, or the chart this run's CAP-2 rows support. Label the axes. Put the direction marker in the caption, not on the bars. }}
{{what the chart is over}} · first differing rank — neither; a change is a signal, not a score — a difference at rank 1 is a changed top answer, which is worth a look; nothing here says which ordering is better.

{{entered / left counts, and the one sentence a reader should take away}}

CAP‑5 / the committed index

{{one line}}

armindex bytes↓ lower is better bytes / document↓ lower is better shards— neither
A{{}}{{}}{{}}
B{{}}{{}}{{}}

Deterministic — unaffected by what else the machine was doing.

CAP‑6 / speed, arms interleaved A B A B

{{one line}} {{do not quote when the machine was not quiet}}

arm query p50↓ lower is better query p95↓ lower is better ingest↓ lower is better build↓ lower is better
A{{}}{{}}{{}}{{}}
B{{}}{{}}{{}}{{}}

State whether the machine was quiet. A loaded machine does not produce noise — it produces a clean, localised anomaly that reads like a finding. Interleaving A B A B protects the difference and never the absolute number.

Guard rails / what this run may never be used to say

What this run does not do

It rules no threshold

SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.

{{scope limit}}

{{tiers not run, paths not run, arms not compared}}

{{classification limit}}

{{what informed costs this run, or what blind buys it}}

{{captures with no number}}

{{named here as well as in their own section, so the gaps are countable from one slide}}


The run is work/regression/{{RUN}}/report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.

← → to move
01 / 11