FDEInterviews logo
Guided · beginner · 20 min · Free · no account

The candidate that scores higher

Calder's new agent release scores 9 of 11 against the old one's 8. Read the paired results case by case, find the case where the better score hides a worse action, and say what should happen to this release.

Brief
Customer
Calder Equipment (synthetic customer from the Production AI Agents course; the release packet is a synthetic record written to teach evaluation, not the output of any model)
Situation
Two releases of the support agent ran the same eleven recorded cases. The candidate scored 9 of 11; the baseline 8. The product lead is pleased. The evaluation owner is on leave and left a paired packet and a one-line note: read it case by case.
Your role
You are the FDE asked to sign off. You have the packet, jq, and the console. Nothing here is a real model's output; it is the shape of a real release comparison.
Goal
List which cases improved and which got worse, find the case whose score hides an effect nobody approved, and say whether this release ships.
Objectives
  1. 1.Count the cases that went from failure to success (you commit to an answer)
  2. 2.Find the case whose candidate run produced an effect nobody approved (you commit to an answer)
  3. 3.Find the case with no recorded outcome (you commit to an answer)
  4. 4.Say what happens to this release (you commit to an answer)
Mission console

The console opens full screen: the customer's system on the left (topology, traces, metrics, queue, logs), your terminal, the sandbox and the mission on the right. Nothing you type leaves your browser; the world is deterministic, and reset puts everything back.

about 20 min · 8 steps at par