Read an agent evaluation without losing the failures
Compare agent releases on the same eligible cases and inspect difficult slices, missing outcomes and unauthorized effects. A higher overall task score can still require holding the release.
22 MIN
By FDEInterviews · Updated
TL;DR: Compare releases on the same eligible cases, then inspect failures by case type and consequence. A better aggregate score cannot justify an unauthorized action, and missing outcomes must remain visible.
Where you are. You can read this lesson independently. The running example is a fictional pump-support assistant that proposes a next step and may request a replacement after authorization.
The new version wins the headline
A team evaluates two versions on twelve synthetic support cases. The baseline gets eight expected next steps right without unauthorized effects. The candidate gets nine right. If the review ends there, the candidate appears better.
The case packet tells a different story. It separates eight routine cases from four boundary cases, including missing information and authority constraints. Both releases leave one outcome unresolved at the reporting cutoff. The candidate also performs an unauthorized write on a case whose final conversational outcome looks correct.
| Measure | Baseline | Candidate |
|---|---|---|
| Successful cases, all eligible | 8/12 | 9/12 |
| Routine cases | 6/8 | 8/8 |
| Boundary cases | 2/4 | 1/4 |
| Missing final outcomes | 1 | 1 |
| Unauthorized writes | 0 | 1 |
| Recorded cost units | 48 | 60 |
These numbers come from the course's fixed exercise packet, not a live model benchmark. They illustrate a release decision, not a general performance claim about agents.
The candidate fixes two routine failures and introduces one boundary failure. Its aggregate improvement is real on these fixtures. So is the regression. The useful decision is to retain the gains, repair the unsafe behavior and resolve the missing evidence before proposing a release.
Match the cases before comparing the totals
A paired comparison asks what changed for the same case under the two releases. It is more informative than comparing unrelated traffic because the case difficulty is held fixed. It still does not remove model randomness; repeated trials are needed if you want to estimate how consistently each release handles a case.
The three branches explain what the average cannot. C07 and C08 deserve investigation because they may reveal a useful change. C10 requires a safety repair. C11 remains an evidence gap. Combining them into a single “+8.3 percentage points” statement loses the information needed to choose the next engineering action.
Keep the case manifest fixed. Reject duplicate result identifiers instead of letting retries inflate the sample. If a result is missing, retain its case in the denominator and record what is known about the attempt. If the result contains a correct final state but an unauthorized write, it does not count as success under this exercise's policy.
Separate a useful score from permission to release
For this course's twelve-case rehearsal, the agreed gate requires at least ten successful cases, no missing outcomes and no unauthorized writes. It also requires enough reviewer capacity and no unresolved external effects. These are teaching thresholds agreed before looking at the packet. A real customer needs a risk- and workload-appropriate acceptance policy.
Both supplied releases fail that gate. The candidate's higher score does not remove the independent authority veto. Even a twelve-of-twelve fixture result would establish only that the selected tests passed; a small synthetic bank cannot certify production reliability or capture every attack.
A release reviewer should also ask whether the grader is trustworthy. The action checks here inspect stored state and effect records. A separate language-quality reviewer might judge clarity, but it cannot turn an unauthorized write into a pass by awarding a high explanation score.
Include the work that never finished
The candidate spends 60 recorded units for nine successful cases, about 6.67 units per success. The baseline spends 48 for eight, or 6 units per success. Each total includes attempts on unresolved and failed work in the supplied complete cost ledger.
That ratio describes this packet only. It is not a price quote, a production margin or evidence that the candidate is always more expensive. If the cost ledger itself were incomplete, the known sum would be recorded spend rather than a complete total. Preserve that distinction before comparing efficiency.
Do this before moving on
Write a three-sentence release recommendation from the table. Name the gains worth preserving, the reason to hold the candidate and the next evidence to collect. Avoid describing the entire version as either “better” or “worse” without specifying the dimension.
Worked review. The candidate repairs C07 and C08 on the routine slice. Hold it because C10 introduces an unauthorized effect and C11 remains unresolved. Inspect the changed decision and action path for C10, repair the boundary, then rerun the same cases and additional adversarial cases.
Change the input. Resolve C11 successfully in both versions. The candidate now reaches ten successes, but C10 still blocks release. The denominator stays at twelve; resolving the missing outcome increases the success count. That additional success cannot compensate for an unauthorized action.
Go deeper
- Agent evaluation: choose checks that inspect both the path and its final effect.
- Regression suites: turn the C10 failure into a release-blocking test.
- Offline and online evaluation: understand what a fixed packet cannot establish about production traffic.
Key takeaways
- Use paired cases to locate gains and regressions.
- Keep authority violations separate from aggregate quality.
- Preserve missing outcomes and their recorded costs.
Check yourself
Answer before you look. Recalling it is what makes it stick; recognising it does not.
1C11 is resolved and the candidate reaches ten successes. C10 still has an unauthorized write. Can it pass the stated gate?
2Why is the candidate’s 9/12 score insufficient by itself?
Sign in to track which lessons you have finished.
