Guided · beginner · 20 min · Free · no account
The candidate that scores higher
Calder's new agent release scores 9 of 11 against the old one's 8. Read the paired results case by case, find the case where the better score hides a worse action, and say what should happen to this release.
Brief
- Customer
- Calder Equipment (synthetic customer from the Production AI Agents course; the release packet is a synthetic record written to teach evaluation, not the output of any model)
- Situation
- Two releases of the support agent ran the same eleven recorded cases. The candidate scored 9 of 11; the baseline 8. The product lead is pleased. The evaluation owner is on leave and left a paired packet and a one-line note:
read it case by case. - Your role
- You are the FDE asked to sign off. You have the packet,
jq, and the console. Nothing here is a real model's output; it is the shape of a real release comparison. - Goal
- List which cases improved and which got worse, find the case whose score hides an effect nobody approved, and say whether this release ships.
Objectives
- 1.Count the cases that went from failure to success (you commit to an answer)
- 2.Find the case whose candidate run produced an effect nobody approved (you commit to an answer)
- 3.Find the case with no recorded outcome (you commit to an answer)
- 4.Say what happens to this release (you commit to an answer)
Mission console
The console opens full screen: the customer's system on the left (topology, traces, metrics, queue, logs), your terminal, the sandbox and the mission on the right. Nothing you type leaves your browser; the world is deterministic, and reset puts everything back.
about 20 min · 8 steps at par
