FDEInterviews logo
Challenge · intermediate · 30 min · Premium

Calibrate the judge

Calder's model judge scores the agent's proposals and it disagrees with the human reviewer on the cases that matter. Compare the two on the labelled set, name which way the judge is wrong, fix the rubric it was given, and rerun.

Brief
Customer
Calder Equipment (synthetic customer from the Production AI Agents course; every label and score on this box is synthetic and made up for the exercise)
Situation
The evaluation runner uses a model as judge to score proposals as acceptable or not. A reviewer hand-labelled twenty cases last week. The judge agrees with the reviewer on most of them and disagrees on a specific kind, and the release gate trusts the judge.
Your role
You are the FDE who owns the evaluation. You have the labelled set, the judge's scores, the judge rubric (which you may edit), and a tool that reruns the judge over the set under the current rubric.
Goal
Count the disagreements each way, name the kind of case the judge accepts that the reviewer rejects, fix the rubric so that kind is rejected, and rerun to show the false accepts gone.
Objectives
  1. 1.Count the false accepts (you commit to an answer)
  2. 2.Name the kind of case the judge wrongly accepts (you commit to an answer)
  3. 3.Fix the rubric so a replacement without an inspection is rejected
  4. 4.Rerun the judge and show the false accepts gone
Mission workspace

This mission is part of premium, with every challenge and incident in the lab. The brief above is the whole problem; the box, the hints and the debrief are what you unlock.

Get full Premium access · ₹2,000 / $25

Every answer, concept and course, plus Premium PDF guides, companion files, the full practice-test bank and work-sample downloads. Referral Premium excludes guide PDFs and their companion files.

6 months · One payment · No auto-renewal

Study alongside free video lessons.