{{ "rejected" if p.rejected else "not rejected" }}
{% if p.pairwise | length > 1 %}
Arm
vs
Pairs
b
c
Difference
95% CI
p
Adjusted p
{% for w in p.pairwise %}
{{ w.treatment }}
{{ w.control }}
{{ w.n_pairs }}
{{ w.b }}
{{ w.c }}
{{ w.difference | pp }}
{{ w.ci | diff_ci }}
{{ w.pvalue | p }}
{{ w.adjusted_pvalue | p }}
{% endfor %}
{% elif p.pairwise %}
{% set w = p.pairwise[0] %}
Discordant blocks: {{ w.b }} where only {{ w.treatment }} succeeded, {{ w.c }} where only {{ w.control }} did, out of {{ w.n_pairs }}.
{% endif %}
{% if charts.forest %}{{ charts.forest }}Black: the primary analysis. Grey: arms treated as independent samples (sensitivity analysis).{% endif %}
{% if r.sequential %}
{% set q = r.sequential %}
Group-sequential looks
{{ q.planned_looks }} planned looks at {% for t in q.planned_fractions %}{{ "%.0f" | format(t * 100) }}%{{ ", " if not loop.last }}{% endfor %} of {{ q.planned_blocks }} blocks, {{ spending_labels.get(q.spending, q.spending) }} error spending.{% if q.stopped_at %} Stopped at look {{ q.stopped_at }}.{% endif %}
Look
Kind
Information
Blocks
Z
Boundary
Crossed
Recorded
{% for k in q.looks %}
{{ k.look }}
{{ k.kind }}
{{ "%.0f" | format(k.fraction * 100) }}%
{{ k.blocks }}
{{ "–" if k.z is none else "%.2f" | format(k.z) }}
{{ "%.3f" | format(k.boundary) }}
{{ "yes" if k.crossed else "no" }}
{{ k.recorded_decision or "–" }}
{% endfor %}
{% endif %}
{% if r.crossover %}
{% set c = r.crossover %}
Crossover rounds
{{ c.cycles }} cycles ({{ c.ab }} with {{ c.treatment }} first, {{ c.ba }} with {{ c.control }} first); {{ c.cycles_used }} complete. Period effect (second round minus first): {{ c.period_effect | pp }}.
{% if charts.rounds %}{{ charts.rounds }}{% endif %}
Rounds
Round
Cycle
Period
Arm
Successes
Completed
Rate
{% for w in c.rounds %}
{{ w.round }}
{{ w.cycle }}
{{ w.period }}
{{ w.arm }}
{{ w.successes }}
{{ w.completed }}
{{ w.rate | rate }}
{% endfor %}
{% endif %}
{% if r.runner %}
Runner
What the runner measured, per arm. Descriptive only: no test is run on it.
Arm
Trials
Requests
Errors
Median latency (ms)
Max p95 latency (ms)
Abnormal exits
{% for u in r.runner %}
{{ u.arm }}
{{ u.trials }}
{{ "–" if u.requests is none else u.requests }}
{{ "–" if u.errors is none else u.errors }}
{{ "–" if u.latency_ms_median is none else "%.1f" | format(u.latency_ms_median) }}
{{ "–" if u.latency_ms_p95 is none else "%.1f" | format(u.latency_ms_p95) }}
{{ "–" if u.abnormal_exits is none else u.abnormal_exits }}
{% endfor %}
{% endif %}
{% if r.ladder %}
{% set lad = r.ladder %}
{% set a = lad.association %}
Checkpoint ladder
Mantel test of a linear association between training step and success: Z = {{ "–" if a.statistic is none else "%.2f" | format(a.statistic) }}, {{ a.pvalue | p_eq }} ({{ "primary" if a.primary else "pre-registered secondary" }} analysis).
{% if charts.ladder %}{{ charts.ladder }}{% if lad.plateau_arm %}Shaded: checkpoints shown to be within {{ "%.1f" | format(lad.margin * 100) }} pp of the final one.{% endif %}{% endif %}
{% if lad.margin is not none and lad.plateau %}
Plateau: each checkpoint against {{ lad.arms[-1] }}, non-inferiority margin {{ "%.1f" | format(lad.margin * 100) }} pp ({{ lad.plateau_method }}), tested from the latest checkpoint backwards. Plateau from: {{ lad.plateau_arm or "–" }}.
Checkpoint
{{ lad.arms[-1] }} minus it
Interval
Tested
Within margin
{% for row in lad.plateau %}
{{ row.arm }}
{{ row.difference | pp }}
{{ row.ci | diff_ci }}
{{ "yes" if row.tested else "no" }}
{{ "yes" if row.noninferior else "no" }}
{% endfor %}
{% endif %}
{% endif %}
{% if r.sensitivity %}
Sensitivity analysis: arms as independent samples
Comparison
Counts
Difference
Newcombe 95% CI
Boschloo p
Fisher p
{% for c in r.sensitivity %}
{{ c.treatment }} vs {{ c.control }}
{{ c.k1 }}/{{ c.n1 }} vs {{ c.k2 }}/{{ c.n2 }}
{{ c.difference | pp }}
{{ c.ci | diff_ci }}
{{ c.boschloo_p | p }}
{{ c.fisher_p | p }}
{% endfor %}
{% endif %}
{% set staged = r.stages | selectattr("n") | list %}
{% if staged %}
Progress stages
Furthest stage reached. Success means reaching {{ s.success_stage }}.
{% if charts.stages %}{{ charts.stages }}{% endif %}
{% if charts.funnel %}{{ charts.funnel }}Share of all trials that reached each stage. The dotted line marks the success stage.{% endif %}
Stage
{% for st in staged %}
{{ st.arm }}: reached / entered
{% endfor %}
{% for name in s.stages %}
{% set i = loop.index0 %}
{{ name }}
{% for st in staged %}{% set step = st.funnel[i] %}
{% if charts.timing %}{{ charts.timing }}Share of all trials that had succeeded by each time. Every trial is observed until it ends, so there is no censoring.{% endif %}
Arm
Successes
Median (s)
95% CI (bootstrap)
{% for t in r.timing %}
{{ t.arm }}
{{ t.n_successes }}
{{ "%.1f" | format(t.median_s) if t.median_s is not none else "–" }}
{{ ("%.1f–%.1f" | format(t.ci.low, t.ci.high)) if t.ci else "–" }}
{% endfor %}
Outcomes per condition
{% if charts.conditions %}{{ charts.conditions }}Success rate per condition and arm; grey cells have no completed trial.{% endif %}
Successes / completed trials per condition and arm
Condition
{% for arm in s.arms %}
{{ arm }}
{% endfor %}
{% for row in r.conditions %}
{{ row.condition }}
{% for c in row.cells %}
{{ c.successes }}/{{ c.completed }}
{% endfor %}
{% endfor %}
Sessions and validity checks
{% if charts.sessions %}{{ charts.sessions }}{% endif %}
Started (UTC)
Operator
Rig
{% for arm in s.arms %}
{{ arm }}
{% endfor %}
{% for row in r.sessions %}
{{ row.started_at.strftime("%Y-%m-%d %H:%M") }}
{{ row.operator }}
{{ row.rig }}
{% for c in row.cells %}
{{ c.successes }}/{{ c.completed }}
{% endfor %}
{% endfor %}
{% if r.drift %}
Check
Arm
Groups
Method
p
Flag
{% for d in r.drift %}
across {{ d.grouping }}s
{{ d.arm }}
{{ d.groups }}
{{ d.method }}
{{ d.pvalue | p }}
{{ "flagged" if d.flagged else "" }}
{% endfor %}
{% else %}
Drift checks need at least two sessions.
{% endif %}
{% set inv = r.invalid %}
Invalid trials (invalid/attempts): {% for a, n in inv.per_arm.items() %}{{ a }} {{ n }}/{{ inv.attempts[a] }}{{ ", " if not loop.last }}{% endfor %}{% if inv.method %}; {{ inv.method }} p = {{ inv.pvalue | p }}{% endif %}{% if inv.flagged %} (flagged){% endif %}.