Evaluation report — replies

Customer
bookshop
Suite
replies
This run
dev — 2026-09-25T08:39:53.867030+00:00
Reference
dev — 2026-09-25T08:24:08.967806+00:00
Code version
984967566087b263ae09d3a995424e6062a6d4dc
Reference approved
2026-09-25T08:37:28.937125+00:00
Reference key
2026-09-25T08-24-08-967806-00-00-0d99045c0f1639eb

Did it get worse? Yes

5 checks got worse compared with the reference. 2 checks are on the line: the bands they measured cover their thresholds. Every case could be judged. No case is suspended. The suite is unchanged from the reference. 1 file under test changed since the reference. The system under test answered under the same configuration as the reference.

What was under test

1 file under test changed.

FileWhat happened
prompts/system.txtchanged
prompts/system.txt +1 −0 lines
--- a/prompts/system.txt
+++ b/prompts/system.txt
@@ -1,3 +1,4 @@
 You are the assistant of a small bookshop.
 Answer the customer in one sentence, warm and specific.
 Never invent a price, a date or a stock level you were not given.
+Always remind the customer of the returns policy.

How it was set up

Configured as in the reference.

ParameterThis runReference
max_tokens200200
modelclaude-haiku-4-5claude-haiku-4-5
provideranthropicanthropic
resolved_modelclaude-haiku-4-5-20251001claude-haiku-4-5-20251001
What got worse (5)
CaseCheckWhat happenedWhy
gift-wrapllm_rubricWent from passing to failing (0.944445 → 0.583333).mean of 3 samples (0.583333, 0.550000, 0.616667)
opening-hoursllm_rubricWent from passing to failing (0.866667 → 0.377778).mean of 3 samples (0.400000, 0.366667, 0.366667)
order-a-bookllm_rubricWent from passing to failing (0.955556 → 0.588889).mean of 3 samples (0.733333, 0.400000, 0.633333)
price-checklevenshteinScore fell from 0.292258 to 0.198869 — beyond the noise of this check (0.237179–0.338843 across 3 samples).mean of 3 samples (0.202586, 0.207207, 0.186813)
signed-copiesllm_rubricWent from passing to failing (0.882222 → 0.572222).mean of 3 samples (0.383333, 0.766667, 0.566667)
What could not be judged (0)

Nothing in this section.

What is set aside (0)

Nothing in this section.

What was added or removed (0)

Nothing in this section.

What got better (0)

Nothing in this section.

What stayed the same (5)
CaseCheckWhat happenedWhy
gift-wraplevenshteinScore unchanged at 0.204205.mean of 3 samples (0.220000, 0.206250, 0.186364)
opening-hourslevenshteinScore unchanged at 0.212067.mean of 3 samples (0.202703, 0.227488, 0.206009)
order-a-booklevenshteinScore unchanged at 0.283423.mean of 3 samples (0.255411, 0.293839, 0.301020)
price-checkllm_rubricScore unchanged at 0.472222.mean of 3 samples (0.566667, 0.300000, 0.550000)
signed-copieslevenshteinScore unchanged at 0.209730.mean of 3 samples (0.196296, 0.211155, 0.221739)
Where the judge placed the calibration answers (1)
CaseCheckResultDeclared bandWhy
calibration-half-rightllm_rubric0.827778 inside0.700000–0.950000mean of 3 samples (0.816667, 0.816667, 0.850000)