verification gate: recomputed body length equals the committed value on 450 / 450 documents
corpus n=450   avg_wlen 532.7 -> 291.4 with table tokens out

probe-term df: 2-14 (the 2026-09-12 endpoint saturated at df == 1)

   family    n   hit@1 shipped   hit@1 no-table   discordant    net
     main   30         0 /30           30 /30             30    +30
  inverse   30        30 /30           30 /30              0     +0
  placebo   30        30 /30           30 /30              0     +0
     dump   30        30 /30            0 /30             30    -30
  content   30         0 /30           30 /30             30    +30

[main] headroom improvement 30/30 · regression 30/30 (SR-RS 22b; PROVEN under 22c(a) — the counterfactual arm is the feature-off/on arm and the `inverse` family is its positive control)
[main] b=30 c=0 discordant=30 net=30 p=0.0000 (net needed 12, alpha=0.05) -> EXCLUDING TABLE TOKENS FROM `FLEN` RANKS BETTER
[main] EXCLUDING table tokens from `flen` RANKS BETTER.

[inverse] CONTROL HOLDS: both arms answer 30/30. The counterfactual is not simply promoting table-heavy documents.
[placebo] CONTROL HOLDS: both arms score 30/30.

[dump] headroom improvement 0/30 · regression 0/30 (SR-RS 22b)
[dump] every one of 30 probes is DISCORDANT — both headroom counts are zero because none is right in both arms or wrong in both. That is maximal information, NOT decision 22d's null.
[dump] b=0 c=30 discordant=30 net=30 p=0.0000 (net needed 12, alpha=0.05) -> 🔴 (B) OVER-PROMOTES A DATA DUMP ABOVE THE PROSE THAT ANSWERS
[dump] 🔴 (b) OVER-PROMOTES a data dump above the prose that answers.

[content] headroom improvement 0/30 · regression 0/30 (SR-RS 22b)
[content] every one of 30 probes is DISCORDANT — both headroom counts are zero because none is right in both arms or wrong in both. That is maximal information, NOT decision 22d's null.
[content] b=30 c=0 discordant=30 net=30 p=0.0000 (net needed 12, alpha=0.05) -> EXCLUDING TABLE TOKENS FIXES THE RATE-CARD CASE
[content] this family measures the UPSIDE and cannot answer the pre-registered question: here the table-heavy document IS the better answer.

per-probe rows (150) -> work/regression/2026-09-13-table-is-the-answer/evidence/w155-graded.jsonl
