acceptance report
Generated {{ report.generated_at }} by juried {{ report.version }}.
Judge: {{ report.judge.provider }} / {{ report.judge.model }} ({% if report.judge.temperature is none %}temperature not set{% else %}temperature {{ report.judge.temperature }}{% endif %}).
Defaults: {{ report.defaults.runs }} runs per scenario, gate at {{ report.defaults.threshold }} on the lower bound of the Wilson 95% interval.
{{ report.summary.criteria }}criteria
{{ report.summary.scenarios }}scenarios
{{ report.summary.gates_passed }}gates upheld
{{ report.summary.gates_failed }}gates failed
{{ report.summary.transport_errors }}transport errors
{% if report.summary.criteria_without_scenarios %}
Criteria with no scenarios: {{ report.summary.criteria_without_scenarios | join(', ') }}.
{% endif %}
{% for criterion in report.criteria %}
{{ criterion.title }} ({{ criterion.id }})
{% if criterion.description %}
{{ criterion.description }}
{% endif %}
{% if not criterion.scenarios %}
No scenarios were run for this criterion.
{% else %}
| Scenario |
Kind |
Runs upheld |
Rate |
Lower bound |
Threshold |
Upper bound |
Latency |
Gate |
{% for scenario in criterion.scenarios %}
| {{ scenario.name }} |
{{ scenario.kind.replace('_', ' ') }} |
{{ scenario.passes }} / {{ scenario.runs }}{% if scenario.transport_errors %} {{ scenario.transport_errors }} transport{% endif %} |
{{ scenario.pass_rate | percent }} |
{{ scenario.interval.lower | percent }} |
{{ scenario.threshold | percent }} |
{{ scenario.interval.upper | percent }} |
{% if scenario.latency.measured %}{{ scenario.latency.mean_ms | round | int }} ms max {{ scenario.latency.max_ms | round | int }} ms{% else %}cached{% endif %} |
{{ 'upheld' if scenario.gate_passed else 'failed' }} |
{% endfor %}
A scenario is upheld when the lower bound of its 95% interval meets the threshold.
{% for scenario in criterion.scenarios %}
{{ scenario.name }} ({{ scenario.id }}) {{ 'upheld' if scenario.gate_passed else 'failed' }}
{% for turn in scenario.history %}
{{ turn.role }}{{ turn.content }}
{% endfor %}
user{{ scenario.message }}
expected{{ scenario.expected }}
{% if scenario.failures %}
{{ scenario.failures | length }} failing run{{ 's' if scenario.failures | length != 1 else '' }}
{% for failure in scenario.failures %}
attempt {{ failure.attempt }}{{ 'transport error' if failure.outcome == 'transport_error' else 'failed' }}
{% if failure.response is not none %}
response{{ failure.response }}
{% endif %}
judge{{ failure.reason }}{% if failure.model %} ({{ failure.model }}, {{ failure.judged_at }}){% endif %}
{% endfor %}
{% endif %}
{% endfor %}
{% endif %}
{% endfor %}