Benchmark report / 2026‑08‑28‑benchmark‑contested / executed 2026‑08‑28
A suite that can
detect a ranking change
The marker suite was saturated before it ran. This is its replacement: four candidates sharing the query’s terms at equal term frequency, equal field and equal length, one correct answer, and a generator that asserts the headroom instead of assuming it.
Arms
unchanged
1.0.0 vs 75ade57 — only the instrument differs
Corpora
1 200 · 10 000
220 contested clusters; every finding re‑checked at 8× the corpus
Per‑query rows
1 900
120 proximity · 60 path · 40 heading
Classification
informed
no delta stated
⚠ Rebuilt 2026‑09‑13 into the report template. Every number is read from the filed run; no test was re‑run, and the four captures this run never took say so rather than being dropped.
How to read this report / every number carries its direction
Which way is good, stated on every metric
| marker | means | metrics it sits on in this run |
|---|---|---|
| ↑ higher is better | a bigger number is a better engine | target‑at‑rank‑1 % · marker hit@5 · b · candidates visible per cluster |
| ↓ lower is better | a smaller number is a better engine | c |
| — neither | a change is a signal, not a score | discordant count · headroom · p · chance level |
🔴 Read the chance level before the score. A four‑candidate cluster is 25 % by guessing. Both arms sit at 21.7 % — so a number that looks like a fifth of the questions is not better than chance, and “higher is better” starts from 25, not 0.
Captures with no number for this run keep their section and say so. Four of the seven have none — this suite measured ranking and nothing else, on purpose.
The result, before the detail / the pre-registered endpoints
A null the instrument could have broken
| id | endpoint | verdict | measured |
|---|---|---|---|
| C1 | proximity, A vs B‑core — primary | no detected change | 0 discordant of 120 · p = 1.0 · 94 queries of headroom · both arms 21.7 % |
| C2 | proximity, B‑core vs B‑tuned — ablation | tuned better | b = 94, c = 0 · p = 1.0e‑28 · 22 % → 100 % |
| C3 | path, A vs B‑core | B better | b = 60, c = 0 · p = 1.7e‑18 · 0 % → 100 % |
| C4 | heading — negative control | inconclusive | 0 discordant of 40 · 0 queries of headroom — saturated at 100 % in both arms |
| C5 | null control — halt gate | pass | 380/380 substantive rows identical, arm A twice. Run first. |
| C6 | headroom disclosure | reported | the line the previous pre‑registration had no way to print |
C3 is a capability delta, not a ranking win — arm A has no path tf field, so the contest is near‑tautological. C2 is an ablation and carries no version claim: it says the proximity reranker works and ships switched off (rerank_weight = 0.0).
The arms / evidence/ARMS.toml
Deliberately the same two engines
| arm A | arm B‑core | arm B‑tuned | |
|---|---|---|---|
| version | fux 1.0.0 | fux 2.0.0‑alpha.2 | B‑core + one key |
| what it is | the previous major | shipped defaults | an ablation — carries no version claim |
| corpora | t1200 · t10000 | identical bytes | identical bytes |
| regeneration | byte‑identical when regenerated, verified on both corpora | ||
🔴 The arms are deliberately identical to the v1‑vs‑head run’s, so only the instrument differs. That is what makes this run readable against that one — and it is the only thing about the two that is comparable.
The pre‑registration — sha256 e8417b33… — was written and on disk before the first corpus byte existed. Its ids are C1–C6: a new id space, because the B ids belong to a frozen document and may never be reused.
The null control / C5, a halt gate, run first
Arm A against itself
nullcontrol · t1200
380/380↑ higher is better
substantive rows identical, arm A twice
It is a gate, not a result
Everything above and below depends on it. A difference between two runs of the same arm is a broken harness, and it would be invisible in every table in this report.
The instrument has its own gate, and it is the point of the run
The generator’s --selftest asserts all four equalities — equal tf, equal field, equal length, path order ≠ role — and HALTS rather than reporting a number it cannot trust. The headroom is asserted, not assumed. Confirmed empirically: both bag‑of‑words arms landed at 21.7 % against a 25 % chance level, and all four candidates were visible in the top 10 in 120 of 120 clusters ↑ higher is better — so every contest was genuinely joined.
CAP‑3 / the ranking endpoint, in this suite's own terms
Both arms, indistinguishable from guessing
| kind | n | arm A↑ higher is better | arm B‑core↑ higher is better | chance— the floor | discordant— neither |
|---|---|---|---|---|---|
| proximity · t1200 | 120 | 21.7 % | 21.7 % | 25 % | 0 |
| proximity · t10000 | 120 | 29.2 % | 29.2 % | 25 % | 0 |
| path · t1200 | 60 | 0 % | 100 % | 25 % | 60 |
| heading · t1200 (control) | 40 | 100 % | 100 % | 25 % | 0 |
| marker hit@5 (control) | 120 | 120/120 | 120/120 | — | 0 |
Both arms rise together at 8× the corpus (21.7 % → 29.2 %) and the discordant count stays 0 in both directions. The instrument is not scale‑dependent and neither is the null.
hit@1, hit@10, hit@20 and hit@50: no number exists for this run. Its endpoint is target‑at‑rank‑1 within a 4‑candidate cluster, which is not hit@k over a corpus. The five‑column requirement postdates the run by sixteen days. Not re‑run to fill it. The retained marker hit@5 row above is a control, not the endpoint.
CAP‑3 / headroom — the column this run was built to print
94 of 120 clusters were free to change hands
| endpoint | n | improvement headroom— neither | b↑ higher is better | c↓ lower is better | p | what the null means |
|---|---|---|---|---|---|---|
| C1 proximity | 120 | 94 | 0 | 0 | 1.0 | about the engines — power 0.99 against a pb .25 / pc .05 effect |
| C3 path | 60 | 60 | 60 | 0 | 1.7e‑18 | a capability, not a ranking win |
| C4 heading (control) | 40 | 0 | 0 | 0 | 1.0 | 🔴 a ceiling — it returned the right answer for the wrong reason |
🔴 This is the first suite in the project whose null is load‑bearing. The previous run’s null could not be told apart from a broken instrument; this one can, because the corpus asserts that 94 clusters could have changed hands and none did.
C4 is Inconclusive, not Pass, and the run’s own headroom column is what caught it. A negative control with zero headroom did not discharge its job — it must be rebuilt before it is cited as a control again.
CAP‑4 / the answer layer
Not asked by this run
No number exists for this run. This is a ranking suite: it measures whether an engine puts the target first among four equally‑evidenced candidates. It contains no unanswerable questions and never calls answer, so there is nothing to report as answered, declined or fabricated. Not re‑run to produce one.
Where the question was asked instead
Its sibling on the same two arms — 2026‑08‑28‑benchmark‑v1‑vs‑head — planted 20 unanswerables and both arms answered all 20. That run’s CAP‑4 section carries it.
Why the gap is stated
“Correct” here is declared by construction, not true: the base documents state no facts. This suite measures whether an engine prefers co‑occurrence and whether it can see a filename — never whether a document answers anything.
CAP‑1 · CAP‑2 / the ranked lists, and what moved
Not captured by this run
No number exists for this run, in either capture. The filed rows are per‑query outcomes — 1 900 of them, one per query per arm per tier, which is what makes every test above recomputable — but the ordered lists themselves were not retained, so there is no diff and no per‑document rank delta. The run is frozen and was not re‑executed.
Within the contest, the equivalent question was answered: all four candidates were visible in the top 10 in 120 of 120 clusters. That is a statement about the contest being joined, not about the corpus‑wide ordering.
CAP‑5 · CAP‑6 / index size and speed
Neither was measured, and the run says so
CAP‑5 — committed index size: no number exists for this run. The plan measured quality only; no index byte count, bytes/document or shard count was taken on either corpus.
CAP‑6 — speed: no number exists for this run. The report states it outright — “no wall‑clock claim… latency was not re‑measured and the previous run’s B5/B6 stand unchallenged.” Not re‑run to fill it, and the earlier run’s numbers are that run’s, on a different corpus.
Why this is a deliberate gap
The two arms and the two corpora were fixed so that only the instrument differed from the sibling run. Re‑measuring bytes and latency would have added nothing the sibling did not already carry — and the sibling’s numbers were taken on a different corpus, so they are not this run’s.
What the capture set changed
From 2026‑09‑13 a run files all seven or is not filed as a benchmark. That is the rule this gap argued for: two sibling runs that cannot be read as one because each took a different half.
Guard rails / what this run may never be used to say
What this run does not do
It states no delta
One session wrote the generator and read the scores, so the run is informed. It may be filed, listed and cited. It is not a generalisation estimate.
A null is not equality
C1 says these engines do not separate these clusters. With headroom that is a real statement about the engines — it is still not “they retrieve equally well”.
No default may move from here
C2 is an ablation on a corpus that rewards exactly what the reranker does; c = 0 is a property of the generator. On hand‑graded text the same reranker is worth 28 → 32. Doing nothing is legitimate.
It captures three of seven
CAP‑1, CAP‑2, CAP‑4 and CAP‑5/6 have no number here, and hit@k is not this suite’s endpoint. The gaps are named, never filled after the fact.
The run is work/regression/2026‑08‑28‑benchmark‑contested/ — report.md, ANALYSIS.md, five VERDICT files, DISCLOSURE‑C6.md, evidence/rows/. Frozen pre‑registration: work/benchmark/PRE-REGISTRATION-CONTESTED.md. This report carries no number that run does not. What must be here at all: records/0053_WORK-benchmark.md.