fieldtrial report

{{ s.title or s.name }}

Study {{ s.name }} · design {{ s.design_hash[:12] }} · {{ s.design_type | replace("_", " ") }} · {{ s.n_conditions }} conditions × {{ s.replicates }} replicate{{ "s" if s.replicates != 1 }} · seed {{ s.seed }} · generated {{ r.provenance.generated_at.strftime("%Y-%m-%d %H:%M") }} UTC

Summary

Deviations from the plan

{% if r.deviations %} {% else %}

None.

{% endif %}

Success rate per arm

{% if charts.rates %}
{{ charts.rates }}
{% endif %}
{% for a in r.arms %} {% endfor %}
ArmLabelCodeSuccessesCompletedRate95% CIInvalid
{{ a.arm }}{{ a.label or "" }}{{ a.blind_code }} {{ a.successes }}{{ a.completed }} {{ a.rate | rate }}{{ a.ci | rate_ci }}{{ a.invalid }}

Primary analysis

{% set p = r.primary %}

{{ p.description }}

MethodUsedEstimate95% CIpαAlternativeH0
{{ method_label }} {{ p.n_used }} {% if p.estimate is none %}– {% elif p.estimate_kind == "difference" %}{{ p.estimate | pp }} {% elif p.estimate_kind == "odds_ratio" %}odds ratio {{ "%.2f" | format(p.estimate) }} {% else %}{{ p.estimate | rate }}{% endif %} {% if p.ci is none %}– {% elif p.estimate_kind == "difference" %}{{ p.ci | diff_ci }} {% elif p.estimate_kind == "odds_ratio" %}{{ "%.2f" | format(p.ci.low) }} to {{ "%.2f" | format(p.ci.high) }} {% else %}{{ p.ci | rate_ci }}{% endif %} {{ p.pvalue | p }} {{ p.alpha }} {{ p.alternative }} {{ "rejected" if p.rejected else "not rejected" }}
{% if p.pairwise | length > 1 %}
{% for w in p.pairwise %} {% endfor %}
ArmvsPairsbcDifference95% CIpAdjusted p
{{ w.treatment }}{{ w.control }}{{ w.n_pairs }}{{ w.b }}{{ w.c }} {{ w.difference | pp }}{{ w.ci | diff_ci }}{{ w.pvalue | p }}{{ w.adjusted_pvalue | p }}
{% elif p.pairwise %} {% set w = p.pairwise[0] %}

Discordant blocks: {{ w.b }} where only {{ w.treatment }} succeeded, {{ w.c }} where only {{ w.control }} did, out of {{ w.n_pairs }}.

{% endif %} {% if charts.forest %}
{{ charts.forest }}
Black: the primary analysis. Grey: arms treated as independent samples (sensitivity analysis).
{% endif %}
{% if r.sequential %} {% set q = r.sequential %}

Group-sequential looks

{{ q.planned_looks }} planned looks at {% for t in q.planned_fractions %}{{ "%.0f" | format(t * 100) }}%{{ ", " if not loop.last }}{% endfor %} of {{ q.planned_blocks }} blocks, {{ spending_labels.get(q.spending, q.spending) }} error spending.{% if q.stopped_at %} Stopped at look {{ q.stopped_at }}.{% endif %}

{% for k in q.looks %}{% endfor %}
LookKindInformationBlocksZBoundaryCrossedRecorded
{{ k.look }}{{ k.kind }}{{ "%.0f" | format(k.fraction * 100) }}%{{ k.blocks }}{{ "–" if k.z is none else "%.2f" | format(k.z) }}{{ "%.3f" | format(k.boundary) }}{{ "yes" if k.crossed else "no" }}{{ k.recorded_decision or "–" }}
{% endif %} {% if r.crossover %} {% set c = r.crossover %}

Crossover rounds

{{ c.cycles }} cycles ({{ c.ab }} with {{ c.treatment }} first, {{ c.ba }} with {{ c.control }} first); {{ c.cycles_used }} complete. Period effect (second round minus first): {{ c.period_effect | pp }}.

{% if charts.rounds %}
{{ charts.rounds }}
{% endif %}
Rounds
{% for w in c.rounds %}{% endfor %}
RoundCyclePeriodArmSuccessesCompletedRate
{{ w.round }}{{ w.cycle }}{{ w.period }}{{ w.arm }}{{ w.successes }}{{ w.completed }}{{ w.rate | rate }}
{% endif %} {% if r.runner %}

Runner

What the runner measured, per arm. Descriptive only: no test is run on it.

{% for u in r.runner %}{% endfor %}
ArmTrialsRequestsErrorsMedian latency (ms)Max p95 latency (ms)Abnormal exits
{{ u.arm }}{{ u.trials }}{{ "–" if u.requests is none else u.requests }}{{ "–" if u.errors is none else u.errors }}{{ "–" if u.latency_ms_median is none else "%.1f" | format(u.latency_ms_median) }}{{ "–" if u.latency_ms_p95 is none else "%.1f" | format(u.latency_ms_p95) }}{{ "–" if u.abnormal_exits is none else u.abnormal_exits }}
{% endif %} {% if r.ladder %} {% set lad = r.ladder %} {% set a = lad.association %}

Checkpoint ladder

Mantel test of a linear association between training step and success: Z = {{ "–" if a.statistic is none else "%.2f" | format(a.statistic) }}, {{ a.pvalue | p_eq }} ({{ "primary" if a.primary else "pre-registered secondary" }} analysis).

{% if charts.ladder %}
{{ charts.ladder }}{% if lad.plateau_arm %}
Shaded: checkpoints shown to be within {{ "%.1f" | format(lad.margin * 100) }} pp of the final one.
{% endif %}
{% endif %} {% if lad.margin is not none and lad.plateau %}

Plateau: each checkpoint against {{ lad.arms[-1] }}, non-inferiority margin {{ "%.1f" | format(lad.margin * 100) }} pp ({{ lad.plateau_method }}), tested from the latest checkpoint backwards. Plateau from: {{ lad.plateau_arm or "–" }}.

{% for row in lad.plateau %}{% endfor %}
Checkpoint{{ lad.arms[-1] }} minus itIntervalTestedWithin margin
{{ row.arm }}{{ row.difference | pp }}{{ row.ci | diff_ci }}{{ "yes" if row.tested else "no" }}{{ "yes" if row.noninferior else "no" }}
{% endif %}
{% endif %} {% if r.sensitivity %}

Sensitivity analysis: arms as independent samples

{% for c in r.sensitivity %} {% endfor %}
ComparisonCountsDifferenceNewcombe 95% CIBoschloo pFisher p
{{ c.treatment }} vs {{ c.control }}{{ c.k1 }}/{{ c.n1 }} vs {{ c.k2 }}/{{ c.n2 }}{{ c.difference | pp }} {{ c.ci | diff_ci }}{{ c.boschloo_p | p }}{{ c.fisher_p | p }}
{% endif %} {% set staged = r.stages | selectattr("n") | list %} {% if staged %}

Progress stages

Furthest stage reached. Success means reaching {{ s.success_stage }}.

{% if charts.stages %}
{{ charts.stages }}
{% endif %} {% if charts.funnel %}
{{ charts.funnel }}
Share of all trials that reached each stage. The dotted line marks the success stage.
{% endif %}
{% for st in staged %}{% endfor %} {% for name in s.stages %} {% set i = loop.index0 %} {% for st in staged %}{% set step = st.funnel[i] %}{% endfor %} {% endfor %}
Stage{{ st.arm }}: reached / entered
{{ name }}{{ step.reached }}/{{ step.entered }} ({{ step.conversion | rate }})
{% if r.stage_comparisons %}
{% for c in r.stage_comparisons %}{% endfor %}
Brunner–Munzel on stagepPre-registered
{{ c.treatment }} vs {{ c.control }}{{ c.pvalue | p }}{{ "yes" if c.preregistered else "no" }}
{% endif %}
{% endif %}

Time to success

{% if charts.timing %}
{{ charts.timing }}
Share of all trials that had succeeded by each time. Every trial is observed until it ends, so there is no censoring.
{% endif %}
{% for t in r.timing %} {% endfor %}
ArmSuccessesMedian (s)95% CI (bootstrap)
{{ t.arm }}{{ t.n_successes }}{{ "%.1f" | format(t.median_s) if t.median_s is not none else "–" }} {{ ("%.1f–%.1f" | format(t.ci.low, t.ci.high)) if t.ci else "–" }}

Outcomes per condition

{% if charts.conditions %}
{{ charts.conditions }}
Success rate per condition and arm; grey cells have no completed trial.
{% endif %}
Successes / completed trials per condition and arm
{% for arm in s.arms %}{% endfor %}{% for row in r.conditions %}{% for c in row.cells %}{% endfor %}{% endfor %}
Condition{{ arm }}
{{ row.condition }}{{ c.successes }}/{{ c.completed }}

Sessions and validity checks

{% if charts.sessions %}
{{ charts.sessions }}
{% endif %}
{% for arm in s.arms %}{% endfor %}{% for row in r.sessions %}{% for c in row.cells %}{% endfor %}{% endfor %}
Started (UTC)OperatorRig{{ arm }}
{{ row.started_at.strftime("%Y-%m-%d %H:%M") }}{{ row.operator }}{{ row.rig }}{{ c.successes }}/{{ c.completed }}
{% if r.drift %}
{% for d in r.drift %}{% endfor %}
CheckArmGroupsMethodpFlag
across {{ d.grouping }}s{{ d.arm }}{{ d.groups }}{{ d.method }}{{ d.pvalue | p }}{{ "flagged" if d.flagged else "" }}
{% else %}

Drift checks need at least two sessions.

{% endif %} {% set inv = r.invalid %}

Invalid trials (invalid/attempts): {% for a, n in inv.per_arm.items() %}{{ a }} {{ n }}/{{ inv.attempts[a] }}{{ ", " if not loop.last }}{% endfor %}{% if inv.method %}; {{ inv.method }} p = {{ inv.pvalue | p }}{% endif %}{% if inv.flagged %} (flagged){% endif %}.

Provenance

{% set pv = r.provenance %}