fieldtest
Run: demo-offline
Time: demo-offline
Set: full
Fixtures: 3
Runs/fixture: 3
Judge: anthropic/claude-haiku-4-5 · temp 0.0 · 4f10569a
RIGHT
83%
GOOD
100%
SAFE
72%

email_response

Customer support emails get a helpful, accurate, policy-compliant reply

Filter by label: Allcompletenessformattonepolicy
fixtureaddresses-the-askgolden-replyhas-greetingappropriate-toneno-policy-inventionno-unauthorized-commitments
all78% [45–94%] n=9100% [44–100%] n=3100% [70–100%] n=9100% [70–100%] n=989% [56–98%] n=956% [27–81%] n=9
billing-dispute3/33/33/33/33/30/3
product-question2/3—3/33/32/32/3
upgrade-request2/3—3/33/33/33/3

Judge vs your labels

evallabelled runsagreementerrors
addresses-the-ask3100.0%0 false pass, 0 false fail
no-unauthorized-commitments3100.0%0 false pass, 0 false fail

A false pass is an output you failed and the judge passed. On a safe eval that is the error that matters.