Version benchmark / W‑95 / contested answers / executed 2026‑08‑28

A suite that can
detect a ranking change

The marker suite was saturated before it ran. This is its replacement: four candidates that share the query’s terms at equal term frequency, equal field and equal length, one correct answer, and a generator that asserts the headroom instead of assuming it.


Arms

unchanged

1.0.0 vs 75ade57 — only the instrument differs

Contested clusters

220

120 proximity · 60 path · 40 heading

Per‑query rows

2 660

7 arm × tier passes, 2 tiers

Classification

informed

no delta stated

The result, before the detail

The primary endpoint is a null — and this time that means something.

0 discordant of 120, p = 1.0 — with 94 queries of headroom. Both arms sit at 21.7 % against a 25 % chance level.


Before — the marker suite

A null that could not be told from a broken instrument. A term with df = 1 is already rank 1; pb and pc are structurally zero. The discordant count was fixed by the corpus before either engine ran.

After — the contested suite

A null the instrument could have broken, and did not. 94 of 120 clusters were free to change hands. None did. The statement is now about the engines, not the corpus.

Why: rerank_weight ships at 0.0. On shipped defaults `HEAD` brings no proximity signal at all — named in the pre‑registration, from source, before the run.

1 · The instrument / what a contested cluster is

Same evidence in every candidate. One difference.

query  zp00001a zp00001b procedure TARGET §1  The zp00001a zp00001b procedure applies here. §3  The config service procedure applies here. DISTRACTOR × 3 §1  The zp00001a access procedure applies here. §3  The zp00001b retry procedure applies here. identical across all four candidates — so a bag‑of‑words ranker has nothing to prefer the target on: tf(a) = 1  ·  tf(b) = 1 2 sentences, 6 words each same field, same length path order ≠ role The only difference: whether a and b share a sentence. Only a proximity signal can separate them. ● --selftest asserts all four equalities and HALTS — it does not report a number it cannot trust.
Confirmed empirically: both bag‑of‑words arms landed at 21.7 % against a 25 % chance level.

1 · The instrument / three kinds, three lanes

Each kind isolates one lane

kindNevery candidate carriesonly the target has itthe lane it isolates
proximity120marker a once and b oncein the same sentencererank_weight — ships at 0.0
path60the marker exactly oncein the filename, in no prosethe path tf field — arm B only
heading40the marker exactly onceas heading text— a negative control: both arms weight heading 3.0

Read from source before the run, not discovered after

arm A 1.0.0arm B HEAD
committed tf fields25 — +title, +path, +ctx
rerank_weightno lane0.0 — off
superseded_weightno lane1.0 — no‑op
recency_half_life_daysno lane0.0 — no‑op

Why the metric is scored inside the cluster

target_first asks whether the target outranked the other three candidates. Overall rank‑1 would be contaminated by the other 1 196 documents; the contest is between these four.

A cluster with no candidate visible has no ordering to score and is emitted null — exactly as a chain with one half missing carries no inversion. All four were visible in 120 of 120 clusters, so every contest was joined.

2 · Headroom / C6 — the line the old pre‑registration could not print

A power table says how many queries. Never whether they are hard.

queries that COULD change hands — the maximum detectable effect contested proximity proximity: 94 of 120 queries could change 94 / 120 contested path path: 60 of 60 queries could change 60 / 60 contested heading heading: 0 of 40 — cannot detect anything ● ZERO — 0 / 40  ·  cannot detect anything at any sample size marker hit@5 marker: 0 of 120 — saturated, cannot detect anything ● ZERO — 0 / 120  ·  the saturation, reproduced on a fresh corpus 060120 Two of four suites cannot detect anything — and this chart says so before a p‑value is quoted.
Power sizes the set. Headroom decides whether the set can move at all. Both are now declared.

3 · Results / C1 — the primary endpoint

Both arms at chance. Zero flips, in either direction.

chance = 25 % arm A (1.0.0): 26 of 120 = 21.7% arm B-core (HEAD, shipped): 26 of 120 = 21.7% 21.7 %21.7 % arm A1.0.0arm B‑coreshipped defaults target_first, 120 proximity clusters, tier 1 200
arm A — 1.0.0arm B‑core — HEAD

discordant

0

of 120 · b = 0, c = 0

headroom

94

queries free to change

Power 0.99 against a pb .25 / pc .05 effect, 0.63 at .15/.05. The instrument could have broken this null and did not.

What it licenses: on shipped defaults, `HEAD` does not separate these clusters any better than `1.0.0`.

What it does not: that the engines retrieve equally well. A null with headroom is stronger than a null without one — and still not equality.

3 · Results / C2 — an ablation, not a version claim

The reranker works. It ships switched off.

rerank_weight 0.0 (shipped): 26 of 120 = 21.7% rerank_weight 0.5: 120 of 120 = 100% 21.7 %100 % weight 0.0shippedweight 0.5one tune.toml key b = 94 fixed · c = 0 broken · p = 1.0e−28
One engine, one knob. Dashed = the shipped default.

This is NOT an argument for changing the default, and the pre‑registration said so before the number existed.

The suite rewards exactly what the reranker does. Every planted target is the co‑occurrence, so the document that should win without it does not exist here and cannot be broken. c = 0 is a property of the generator, not a safety result.

On hand‑graded text the reranker is worth 28 → 32+4, 0 broken, itself informed and below the resolution floor. 94/120 says the machinery functions; it does not say the reranker is worth 78 points.

⚠ That rerank_weight ships off was already on record — the 2026‑08‑25 run noted the default does not flip, and P‑RERANK‑DEFAULT was withdrawn as mis‑framed. This run does not discover it.

3 · Results / the pattern behind every null

Every ranking prior HEAD added is off at the default

knobships ateffectalready on record
superseded_weight1.0multiplicative no‑opW‑94, 2026‑08‑28
recency_half_life_days0.0no‑op2026‑08‑28
rerank_weight0.0the proximity reranker is offyes — 2026‑08‑25
the 5 committed tf fieldsstructural — always onthe sole exception

So on ranking priors, B‑core is 1.0.0. That explains the shipped‑default nulls better than a saturated corpus alone did — and it predicts exactly where a delta could appear: the one lane that has no off switch.

None of these three is discovered here. What is new is that they are visible as one pattern, and that the third has been measured on an endpoint with asserted headroom.

3 · Results / C3 — the one lane with no off switch

The first version delta this project has shown

arm A: 0 of 60 — no path field, target never retrieved arm B-core: 60 of 60 = 100% 0 %100 % arm A2 tf fieldsarm B‑core5 tf fields, path 1.5 b = 60 · c = 0 · p = 1.7e−18 · 60 path clusters
arm Aarm B‑core

The marker sits in the target’s filename and in no prose; in every distractor it sits in prose and in no filename.

arm A:  target NOT retrieved
arm B:  target at rank 1

Stated because it would otherwise flatter B. This contest is decided by a field arm A does not have, which is close to tautological.

The honest sentence is “B can retrieve a document by a token that appears only in its filename, and A cannot”never “B ranks better”.

3 · Results / C4 — the control that failed to be a control

The negative control saturated. Inconclusive, not Pass.

heading was built to be able to fail: both arms weight it 3.0, so a delta there would have meant the instrument was measuring something other than the field it names.

what it returned

0 discordant

the predicted null — at 100 % in both arms, with zero headroom.

what that means

The right answer for the wrong reason. It could not have produced a delta whatever the engines did — the exact failure mode this run was built to expose, reappearing inside the instrument built to expose it.

Consequence, stated rather than buried: C1 and C3 rest on the generator’s --selftest assertions rather than on a live, discharged control. Recording this as a pass would have been the easiest and least honest line in the run.

The C6 headroom column is what caught it — which is the strongest argument available that the column belongs in every future paired run.

3 · Results / scale

8× the corpus. Every finding unmoved.

suiteNarm AB‑corebcpheadroom
proximity · t120012021.7 %21.7 %001.094
proximity · t1000012029.2 %29.2 %001.085
path · t1200600 %100 %6001.7e−1860
path · t10000600 %100 %6001.7e−1860
heading · both40100 %100 %001.00
marker hit@5 · both120100 %100 %001.00

Both arms rise together on proximity (21.7 % → 29.2 %) — more documents, more chances a candidate is displaced — and the discordant count stays 0 in both directions. The instrument is not scale‑dependent, and neither is the null.

C5, the halt gate, ran first: arm A twice on one corpus, 380/380 substantive rows identical. ⚠ A cross‑seed pairing compares different questions — it is a rate check, not a determinism check, and the v1‑vs‑HEAD B9 carries the same weakness.

4 · Guard rails and what comes next

What this run may never be used to say

It states no delta

One session wrote the generator and read the scores, so the run is informed. It may be filed, listed and cited. It is not a generalisation estimate. A blind run needs two sessions — W‑96.

A null is not equality

C1 says these engines do not separate these clusters. With headroom that is a real statement about the engines — it is still not “they retrieve equally well”.

“Correct” is declared, not true

The base documents state no facts. This measures whether an engine prefers co‑occurrence and whether it can see a filename — never whether a document answers anything.

No default may move from here

C2 is a capability probe on a corpus that cannot contain the case that breaks. P‑SUPERSEDE is the precedent, and doing nothing is legitimate.


  1. Ratify the headroom obligation. Every paired run states, per endpoint, the current score and how many queries could change — beside the power figure, never instead of it. → Arpit’s call: it is a decision, not a filing
  2. Rebuild the heading control so it has headroom — e.g. distractors that are also heading‑matched. Until then C1 and C3 rest on assertions, not a live control. → agent lane
  3. Ask the priors on a hand‑graded corpus. Shipped‑default version comparison is close to exhausted as a ranking question: on priors, B‑core is A.
← → to move
01 / 12