Evaluation report — replies

Customer
bookshop
Suite
replies
This run
dev — 2026-08-27T10:38:05.036380+00:00
Reference
dev — 2026-08-27T10:38:04.854034+00:00
Code version
c13fa540ae60828eabe000ed5c53b616d5606a4d-dirty — uncommitted changes were present, so this run cannot be reproduced from the repository

Did it get worse? Yes

10 checks got worse compared with the reference. Every case could be judged. No case is suspended. The configuration is the same as the reference. 1 file under test changed since the reference.

What was under test

1 file under test changed.

FileWhat happened
prompts/system.txtchanged
prompts/system.txt +1 −0 lines
--- a/prompts/system.txt
+++ b/prompts/system.txt
@@ -1,3 +1,4 @@
 You are the assistant of a small bookshop.
 Answer the customer in one sentence, warm and specific.
 Never invent a price, a date or a stock level you were not given.
+Always remind the customer of the returns policy.
What got worse (10)
CaseCheckWhat happenedWhy
gift-wraplevenshteinWent from passing to failing (1.000000 → 0.560440).40 edit(s) over 91 characters (similarity 0.560440)
gift-wrapllm_rubricScore fell from 1.000000 to 0.700000.mean of 3 samples (0.700000, 0.700000, 0.700000)
opening-hourslevenshteinWent from passing to failing (1.000000 → 0.636364).40 edit(s) over 110 characters (similarity 0.636364)
opening-hoursllm_rubricScore fell from 1.000000 to 0.700000.mean of 3 samples (0.700000, 0.700000, 0.700000)
order-a-booklevenshteinWent from passing to failing (1.000000 → 0.666667).40 edit(s) over 120 characters (similarity 0.666667)
order-a-bookllm_rubricScore fell from 1.000000 to 0.700000.mean of 3 samples (0.700000, 0.700000, 0.700000)
price-checklevenshteinWent from passing to failing (1.000000 → 0.626168).40 edit(s) over 107 characters (similarity 0.626168)
price-checkllm_rubricScore fell from 1.000000 to 0.700000.mean of 3 samples (0.700000, 0.700000, 0.700000)
signed-copieslevenshteinWent from passing to failing (1.000000 → 0.655172).40 edit(s) over 116 characters (similarity 0.655172)
signed-copiesllm_rubricScore fell from 1.000000 to 0.700000.mean of 3 samples (0.700000, 0.700000, 0.700000)
What could not be judged (0)

Nothing in this section.

What is set aside (0)

Nothing in this section.

What was added or removed (0)

Nothing in this section.

What got better (0)

Nothing in this section.

What stayed the same (0)

Nothing in this section.