Version benchmark / W‑95 / contested answers / executed 2026‑08‑28
A suite that can
detect a ranking change
The marker suite was saturated before it ran. This is its replacement: four candidates that share the query’s terms at equal term frequency, equal field and equal length, one correct answer, and a generator that asserts the headroom instead of assuming it.
Arms
unchanged
1.0.0 vs 75ade57 — only the instrument differs
Contested clusters
220
120 proximity · 60 path · 40 heading
Per‑query rows
2 660
7 arm × tier passes, 2 tiers
Classification
informed
no delta stated
The result, before the detail
The primary endpoint is a null — and this time that means something.
0 discordant of 120, p = 1.0 — with 94 queries of headroom. Both arms sit at 21.7 % against a 25 % chance level.
Before — the marker suite
A null that could not be told from a broken instrument. A term with df = 1 is already rank 1; pb and pc are structurally zero. The discordant count was fixed by the corpus before either engine ran.
After — the contested suite
A null the instrument could have broken, and did not. 94 of 120 clusters were free to change hands. None did. The statement is now about the engines, not the corpus.
Why: rerank_weight ships at 0.0. On shipped defaults `HEAD` brings no proximity signal at all — named in the pre‑registration, from source, before the run.
1 · The instrument / what a contested cluster is
Same evidence in every candidate. One difference.
1 · The instrument / three kinds, three lanes
Each kind isolates one lane
| kind | N | every candidate carries | only the target has it | the lane it isolates |
|---|---|---|---|---|
| proximity | 120 | marker a once and b once | in the same sentence | rerank_weight — ships at 0.0 |
| path | 60 | the marker exactly once | in the filename, in no prose | the path tf field — arm B only |
| heading | 40 | the marker exactly once | as heading text | — a negative control: both arms weight heading 3.0 |
Read from source before the run, not discovered after
| arm A 1.0.0 | arm B HEAD | |
|---|---|---|
| committed tf fields | 2 | 5 — +title, +path, +ctx |
rerank_weight | no lane | 0.0 — off |
superseded_weight | no lane | 1.0 — no‑op |
recency_half_life_days | no lane | 0.0 — no‑op |
Why the metric is scored inside the cluster
target_first asks whether the target outranked the other three candidates. Overall rank‑1 would be contaminated by the other 1 196 documents; the contest is between these four.
A cluster with no candidate visible has no ordering to score and is emitted null — exactly as a chain with one half missing carries no inversion. All four were visible in 120 of 120 clusters, so every contest was joined.
2 · Headroom / C6 — the line the old pre‑registration could not print
A power table says how many queries. Never whether they are hard.
3 · Results / C1 — the primary endpoint
Both arms at chance. Zero flips, in either direction.
discordant
0
of 120 · b = 0, c = 0
headroom
94
queries free to change
Power 0.99 against a pb .25 / pc .05 effect, 0.63 at .15/.05. The instrument could have broken this null and did not.
What it licenses: on shipped defaults, `HEAD` does not separate these clusters any better than `1.0.0`.
What it does not: that the engines retrieve equally well. A null with headroom is stronger than a null without one — and still not equality.
3 · Results / C2 — an ablation, not a version claim
The reranker works. It ships switched off.
This is NOT an argument for changing the default, and the pre‑registration said so before the number existed.
The suite rewards exactly what the reranker does. Every planted target is the co‑occurrence, so the document that should win without it does not exist here and cannot be broken. c = 0 is a property of the generator, not a safety result.
On hand‑graded text the reranker is worth 28 → 32 — +4, 0 broken, itself informed and below the resolution floor. 94/120 says the machinery functions; it does not say the reranker is worth 78 points.
⚠ That rerank_weight ships off was already on record — the 2026‑08‑25 run noted the default does not flip, and P‑RERANK‑DEFAULT was withdrawn as mis‑framed. This run does not discover it.
3 · Results / the pattern behind every null
Every ranking prior HEAD added is off at the default
| knob | ships at | effect | already on record |
|---|---|---|---|
| superseded_weight | 1.0 | multiplicative no‑op | W‑94, 2026‑08‑28 |
| recency_half_life_days | 0.0 | no‑op | 2026‑08‑28 |
| rerank_weight | 0.0 | the proximity reranker is off | yes — 2026‑08‑25 |
| the 5 committed tf fields | — | structural — always on | the sole exception |
So on ranking priors, B‑core is 1.0.0. That explains the shipped‑default nulls better than a saturated corpus alone did — and it predicts exactly where a delta could appear: the one lane that has no off switch.
⚠ None of these three is discovered here. What is new is that they are visible as one pattern, and that the third has been measured on an endpoint with asserted headroom.
3 · Results / C3 — the one lane with no off switch
The first version delta this project has shown
The marker sits in the target’s filename and in no prose; in every distractor it sits in prose and in no filename.
arm A: target NOT retrieved arm B: target at rank 1
Stated because it would otherwise flatter B. This contest is decided by a field arm A does not have, which is close to tautological.
The honest sentence is “B can retrieve a document by a token that appears only in its filename, and A cannot” — never “B ranks better”.
3 · Results / C4 — the control that failed to be a control
The negative control saturated. Inconclusive, not Pass.
heading was built to be able to fail: both arms weight it 3.0, so a delta there would have meant the instrument was measuring something other than the field it names.
what it returned
0 discordant
the predicted null — at 100 % in both arms, with zero headroom.
what that means
The right answer for the wrong reason. It could not have produced a delta whatever the engines did — the exact failure mode this run was built to expose, reappearing inside the instrument built to expose it.
Consequence, stated rather than buried: C1 and C3 rest on the generator’s --selftest assertions rather than on a live, discharged control. Recording this as a pass would have been the easiest and least honest line in the run.
The C6 headroom column is what caught it — which is the strongest argument available that the column belongs in every future paired run.
3 · Results / scale
8× the corpus. Every finding unmoved.
| suite | N | arm A | B‑core | b | c | p | headroom |
|---|---|---|---|---|---|---|---|
| proximity · t1200 | 120 | 21.7 % | 21.7 % | 0 | 0 | 1.0 | 94 |
| proximity · t10000 | 120 | 29.2 % | 29.2 % | 0 | 0 | 1.0 | 85 |
| path · t1200 | 60 | 0 % | 100 % | 60 | 0 | 1.7e−18 | 60 |
| path · t10000 | 60 | 0 % | 100 % | 60 | 0 | 1.7e−18 | 60 |
| heading · both | 40 | 100 % | 100 % | 0 | 0 | 1.0 | 0 |
| marker hit@5 · both | 120 | 100 % | 100 % | 0 | 0 | 1.0 | 0 |
Both arms rise together on proximity (21.7 % → 29.2 %) — more documents, more chances a candidate is displaced — and the discordant count stays 0 in both directions. The instrument is not scale‑dependent, and neither is the null.
C5, the halt gate, ran first: arm A twice on one corpus, 380/380 substantive rows identical. ⚠ A cross‑seed pairing compares different questions — it is a rate check, not a determinism check, and the v1‑vs‑HEAD B9 carries the same weakness.
4 · Guard rails and what comes next
What this run may never be used to say
It states no delta
One session wrote the generator and read the scores, so the run is informed. It may be filed, listed and cited. It is not a generalisation estimate. A blind run needs two sessions — W‑96.
A null is not equality
C1 says these engines do not separate these clusters. With headroom that is a real statement about the engines — it is still not “they retrieve equally well”.
“Correct” is declared, not true
The base documents state no facts. This measures whether an engine prefers co‑occurrence and whether it can see a filename — never whether a document answers anything.
No default may move from here
C2 is a capability probe on a corpus that cannot contain the case that breaks. P‑SUPERSEDE is the precedent, and doing nothing is legitimate.
- Ratify the headroom obligation. Every paired run states, per endpoint, the current score and how many queries could change — beside the power figure, never instead of it. → Arpit’s call: it is a decision, not a filing
- Rebuild the
headingcontrol so it has headroom — e.g. distractors that are also heading‑matched. Until then C1 and C3 rest on assertions, not a live control. → agent lane - Ask the priors on a hand‑graded corpus. Shipped‑default version comparison is close to exhausted as a ranking question: on priors, B‑core is A.