Benchmark report / 2026‑08‑28‑benchmark‑v1‑vs‑head / executed 2026‑08‑28

Every paired test
returned zero

Not a small delta — no query changed its outcome in either direction, on any tier, on any suite. The primary endpoint turned out to be saturated before it ran, and the supersession endpoint failed its predicted pass for a reason nobody had guessed.


Arms

1.0.0 → alpha.2

two engines resident at once, one venv each, CPython 3.11.15

Corpora

100 / 1k / 10k

one generated corpus, three tiers, byte‑identical per arm

Questions

240 × 3

marker queries per tier, + 40 chains + 20 unanswerables + decoys

Classification

informed

no delta stated — one session wrote the generator and read the scores

Rebuilt 2026‑09‑13 into the report template. Every number is read from the filed run; no test was re-run, and captures this run never took say so rather than being dropped.

How to read this report / every number carries its direction

Which way is good, stated on every metric

markermeansmetrics it sits on in this run
↑ higher is bettera bigger number is a better enginehit@5 · hit@10 · rank‑1 · declined‑when‑unanswerable · b
↓ lower is bettera smaller number is a better engineask p50/p95 · ingest · build · index bytes · wheel bytes · fabricated · chain inversions · c
— neithera change is a signal, not a scoreshards · discordant count · headroom · decoys reaching the top 5

“Better” means better on THIS instrument, and nothing more. A direction marker says which way a metric points; it never says the difference is real or large enough to act on. That is a person reading the rows.

Captures with no number for this run keep their section and say so — this run predates the capture set, and three of the seven were never taken.

The result, before the detail / the pre-registered endpoints

Two nulls, one failure, and a gate that did its job

idendpointverdictmeasured
B1retrieval quality, tier 1 000 — primaryinconclusive0 discordant of 240 · p = 1.0 — and hit@5 240/240 in BOTH arms
B2supersession — current ranks firstfailboth arms invert 21 of 40 at tier 1 000 — superseded_weight ships at 1.0, a no‑op
B3committed bytes and wheelpassindex 1.002 × · wheel 7.11 MB → 259 KB · no int8 vectors committed
B5 / B6query and ingest latencypassask p95 1.32 × (bar 1.5 ×) · cold ingest 1.04 × (bar 2.0 ×)
B7declining on unanswerablesinconclusive0 of 20 declined in both arms — a generated corpus can only test declining when nothing matches
B9null control — run firstpass300/300 identical rows · 0 discordant across seeds

B4 and B8 do not exist. The item says “thresholds B1–B9”; the frozen document defines seven ids. Nothing was skipped — two ids were never written.

The arms / evidence/ARMS.toml

Two engines resident at once

 arm Aarm B
versionfux‑engine 1.0.02.0.0‑alpha.2 @ 75ade57
python3.11.153.11.15
corpusone generated corpusits own copy of identical bytes
record shapefux.index.v1fux.index.v2
tf_fields["heading","body"]["body","heading","title","path","ctx"]
dense lanedocument records carry codeabsent — B3's one named check, read from a record

The frozen HEAD was written into the pre-registration before the first command ran. The differential law was checked before any --fast number was produced: ask --fast and ask --scan byte‑identical across all 240 queries, 0 mismatches.

The null control / B9, run first, in two halves

Arm A against itself

same corpus, twice

300/300↑ higher is better

identical rows

arm A vs arm A′, second seed

0 discordant↓ lower is better

p = 1.0

Run first, deliberately

No number in this report was produced before this gate passed.


A method correction filed later: a cross‑seed pairing compares different questions — it is a rate check, not a determinism check. B9 carries that weakness, named by the contested run that found it.

CAP‑3 / hit@k — and why it could not move

Both arms perfect, on every tier

tierarm rank‑1↑ higher is better hit@5↑ higher is better hit@10↑ higher is better MRR@10↑ higher is better chain inversions↓ lower is better
100A / B240/240240/240240/2401.00005/10 · 5/10
1 000A / B240/240240/240240/2401.000021/40 · 21/40
10 000A / B240/240240/240240/2401.000017/40 · 17/40

hit@20 and hit@50: no number exists for this run. The capture set that requires five values of k was ruled on 2026‑09‑13, sixteen days after this run filed. Its plan chose hit@5 and hit@10. Not re‑run to fill the gap — the report is frozen and a better number is a new run.

🔴 A marker planted in exactly one document has df = 1. It is already at rank 1 and nothing can move it — so the discordant count of 0 was fixed by the corpus design, not measured from the engines.

CAP‑3 / headroom — the column this run did not have

Zero headroom in the direction that mattered

endpointn improvement headroom— neither b↑ higher is better c↓ lower is better discordant— neither p
B1  hit@5, tier 1 00024000001.0
B2  current‑ranks‑first40190001.0
B7  declined20200001.0
B9  hit@5, A vs A′24000001.0

🔴 B1's improvement headroom is 0, and that is the whole verdict. pb and pc are structurally zero on this suite, so a null reports a ceiling, never a match. Ruling it PASS would report a property of the queries as a finding about the engines.

The reusable lesson, and it outlived the run: the pre-registration's power table asked how many queries and answered correctly. It never asked whether the queries could express the effect. A power calculation does not tell you the queries are hard.

CAP‑4 / the answer layer, and the 20 planted unanswerables

Nothing declined, in either version

arm planted unanswerable— neither declined↑ higher is better fabricated↓ lower is better
A  1.0.020020
B  2.0.0‑alpha.220020

On all 20, arm B returned band: partial, answerable: true — and for the absent‑entity half it named the missing term: missing: ["zq00000w"], coverage: 0.0009. Arm B knows the queried term is absent and answers anyway.

Arm A emits no confidence block at all, so this is a capability table and not a contrast. Stated, not smoothed — the asymmetry is permanent, because the flag cannot be added to a released version.

The controls: a decoy reached the top 5 for 1 of 50 shadowed queries at tier 1 000 and 1 of 208 at tier 10 000 — neither, identically in both arms.

CAP‑1 · CAP‑2 / the ranked lists, and what moved

Not captured by this run

No number exists for this run, in either capture. The ranked lists were not retained — the plan recorded pass/fail per query and discarded the ordering — so there is nothing to diff and no rank delta to file. The run is frozen and was not re‑executed to produce them.

What this costs a reader

Every statement about where the two engines agree rests on hit@5 being 240/240 — which is a ceiling. Whether the orderings below rank 5 were identical is unknown and unknowable from what was filed.

Why it is a section and not a deletion

A missing section reads as a capture that was never required. This one was required from 2026‑09‑13, and this run predates it — which is exactly what a reader needs to know before comparing it with a later run.

🔴 This run is also the first in the store to file PER‑QUERY ROWS for every query, arm and tier — which is why its b, c and every test above can be recomputed by anybody.

CAP‑5 / the committed index, and the wheel

The index held; the wheel collapsed

  arm A↓ lower is better arm B↓ lower is better ratio↓ lower is better bar
index bytes, 1 000 docs1 462 3421 465 0651.002 ×≤ 1.25 × pass
index bytes, 10 000 docs14 147 49214 117 8570.998 ×≤ 1.25 × pass
shards, 10 000 docs— neither256256
published wheel7 113 352258 9010.036 ×

bytes / document: not filed by this run. It is divisible from the two columns beside it; it is left uncomputed here because this report renders the filed rows and does not derive new ones.

CAP‑6 / speed, arms interleaved A B A B, tier 10 000

Both bars cleared, with the slower engine

  arm A↓ lower is better arm B↓ lower is better ratio↓ lower is better bar
ask p5079.3 ms103.7 ms1.31 ×
ask p9585.6 ms112.6 ms1.32 ×≤ 1.5 × pass
cold ingest median25.68 s26.78 s1.04 ×≤ 2.0 × pass
build median612 ms1 086 ms1.77 ×

1 200 timed queries per arm (240 × 5 repeats), 20 warm‑ups per arm discarded, three cold ingest repeats per arm.

Arm B's three ingest repeats were 25.4 / 26.8 / 38.8 s — the third is an outlier on a laptop that was doing other things. The median is reported and all three are filed.

Guard rails / what this run may never be used to say

What this run does not do

It licenses no improvement

B1 is inconclusive, not pass. The numeric condition resolved and the experiment reached almost nothing — ruling it pass would report the corpus as a finding about the engines.

It states no delta

informed: one session wrote the generator and read the scores. Never compared with a blind run, and not a generalisation estimate.

It captures three of seven

CAP‑1, CAP‑2 and two hit@k columns have no number here. The capture set postdates the run; the gaps are named, never filled after the fact.

B2's failure is not a defect report

Both arms invert identically because superseded_weight ships at 1.0. Post‑hoc at 0.5 the same arm goes 21/40 → 0/40, 21 fixed / 0 broken — the machinery works and is switched off. Labelled post‑hoc, and out of the verdict.


The run is work/regression/2026‑08‑28‑benchmark‑v1‑vs‑head/report.md, ANALYSIS.md, six VERDICT files, evidence/rows/. Frozen pre‑registration: work/benchmark/PRE-REGISTRATION-V1-VS-HEAD.md. This report carries no number that run does not. What must be here at all: records/0053_WORK-benchmark.md.

← → to move
01 / 12