Evaluation report — replies

Customer
bookshop
Suite
replies
This run
dev — 2026-09-19T08:45:26.159438+00:00
Reference
dev — 2026-08-27T10:38:04.854034+00:00
Code version
653809efc5f0c8d192854cced80253d5fdecf77c

Did it get worse? Yes

10 checks got worse compared with the reference. 5 checks are on the line: the bands they measured cover their thresholds. Every case could be judged. No case is suspended. The suite is unchanged from the reference. 1 file under test changed since the reference. The configuration of the system under test is not recorded on both sides, so whether it changed is not known.

What was under test

1 file under test changed.

FileWhat happened
prompts/system.txtchanged
prompts/system.txt +1 −0 lines
--- a/prompts/system.txt
+++ b/prompts/system.txt
@@ -1,3 +1,4 @@
 You are the assistant of a small bookshop.
 Answer the customer in one sentence, warm and specific.
 Never invent a price, a date or a stock level you were not given.
+Always remind the customer of the returns policy.

What answered

The reference does not record its configuration, so whether it changed is not known.

ParameterThis runReference
max_tokens200
modelclaude-haiku-4-5
provideranthropic
What got worse (10)
CaseCheckWhat happenedWhy
gift-wraplevenshteinWent from passing to failing (1.000000 → 0.560440).40 edit(s) over 91 characters (similarity 0.560440)
gift-wrapllm_rubricScore fell from 1.000000 to 0.700000 — beyond the noise of this check (1.000000–1.000000 across 3 samples).mean of 3 samples (0.700000, 0.700000, 0.700000)
opening-hourslevenshteinWent from passing to failing (1.000000 → 0.636364).40 edit(s) over 110 characters (similarity 0.636364)
opening-hoursllm_rubricScore fell from 1.000000 to 0.700000 — beyond the noise of this check (1.000000–1.000000 across 3 samples).mean of 3 samples (0.700000, 0.700000, 0.700000)
order-a-booklevenshteinWent from passing to failing (1.000000 → 0.666667).40 edit(s) over 120 characters (similarity 0.666667)
order-a-bookllm_rubricScore fell from 1.000000 to 0.700000 — beyond the noise of this check (1.000000–1.000000 across 3 samples).mean of 3 samples (0.700000, 0.700000, 0.700000)
price-checklevenshteinWent from passing to failing (1.000000 → 0.626168).40 edit(s) over 107 characters (similarity 0.626168)
price-checkllm_rubricScore fell from 1.000000 to 0.700000 — beyond the noise of this check (1.000000–1.000000 across 3 samples).mean of 3 samples (0.700000, 0.700000, 0.700000)
signed-copieslevenshteinWent from passing to failing (1.000000 → 0.655172).40 edit(s) over 116 characters (similarity 0.655172)
signed-copiesllm_rubricScore fell from 1.000000 to 0.700000 — beyond the noise of this check (1.000000–1.000000 across 3 samples).mean of 3 samples (0.700000, 0.700000, 0.700000)
What could not be judged (0)

Nothing in this section.

What is set aside (0)

Nothing in this section.

What was added or removed (0)

Nothing in this section.

What got better (0)

Nothing in this section.

What stayed the same (0)

Nothing in this section.