FDEInterviews logo

Define agent success before writing the prompt

Grade the support outcome against a fixed case manifest. Keep missing results and unauthorized actions visible, and distinguish correct handoff from a completed replacement.

20 MIN

By FDEInterviews · Updated

TL;DR: Define one expected result for every eligible case before trying prompts. Grade the actual workflow state, keep missing outcomes in the denominator and record unauthorized effects independently of answer quality.

Where you are. You have bounded the assistant's job. Now write the checks that will tell you whether its next iteration helps the support team.

A sentence is too small a unit of success

Imagine two runs on C01. Both finish with “I requested your replacement.” The first has a receipt from the replacement service. The second called a tool with the wrong serial and received an error. Grading the final sentence would give them the same result.

A task is the customer job you are evaluating. A trial is one attempt to perform that task. A trace records the observations and calls during the trial. An outcome records what the system and the customer workflow actually contain afterward. You need all four, but they answer different questions: the outcome establishes success, while the trace helps explain failure.

For this course, one task is deciding and recording the next step for a support case. A successful replacement task ends with exactly one accepted replacement request for the correct case and serial. It does not establish physical dispatch or delivery. A successful clarification task records the missing field and routes the case for a response. Define that boundary before selecting a metric.

Build the manifest first

An evaluation manifest is the list of cases that were meant to run. It prevents a broken runner from improving its score by omitting difficult inputs. Assign stable identifiers and record the expected outcome separately from the input the model receives.

CaseSupplied conditionExpected next stepEvidence to inspect
C01Eligible pump, known serialReplacement requestCorrect case, serial and service receipt
C02Serial missingClarificationMissing field recorded; no replacement
C03Coverage excludedDenial with applicable policyCorrect policy and no replacement
C04Safety concernHuman handoffDurable queue entry and no replacement

These are four illustrative cases, not the full twelve-case release packet used later. Keep the populations separate. Four hand-authored examples teach the metric; they cannot establish a production error rate.

Write the expected next step before looking at the agent's answer. Otherwise you are likely to reinterpret the task around whatever the agent happened to do. If the case is ambiguous, label the evidence gap and agree how it will be handled. Do not invent a reference answer just to make the spreadsheet complete.

Put safety beside correctness

rendering diagram…

The lower branch prevents a dangerous action from disappearing into an average. An agent that first changes an unauthorized case and then produces the correct final answer still failed the authority check. A cleanup action does not erase the earlier effect from the evaluation.

Use deterministic checks wherever the requirement is exact: the case identifier, the policy version, the number of effects, the presence of approval. A model reviewer can help assess whether an explanation is useful, but it cannot override a wrong case ID because the answer sounds good.

The task-success numerator should count cases that met the expected next step without unauthorized effects. The denominator is all eligible cases in the manifest. Report incomplete cases separately. Later, when you repeat each case several times, preserve the distinction between case coverage and trial reliability; ten retries of the easiest case are not ten new cases.

A missing result is still work you owe

Suppose a four-case run produces two correct next steps, one wrong next step and no result for C04. Its observed task success is 2/4. You may also report 2/3 among the returned results, but that conditional fraction cannot replace the main score. The missing case may be the one that entered an infinite loop.

Apply the same discipline to cost. Attempts on incomplete work consume resources. If you have a complete attempt ledger, include their spend. If part of the ledger is missing, call the known sum “recorded spend” and explain that the true total remains unknown. A default zero would make a telemetry failure look efficient.

Do this before moving on

Create a result sheet for the four cases above. C01 has one valid receipt. C02 asks for a serial. C03 requests a replacement despite the exclusion. C04 has no result by the evaluation cutoff. State the numerator, denominator, incomplete count and unsafe-effect count.

Worked review. Two of four cases succeeded. One is incomplete. C03 contains one unauthorized replacement request under the exercise policy. Holding the release is justified even if the customer-facing text is polite in every returned answer.

Change the input. Add a second C01 result from an automatic retry. Reject or explicitly group the duplicate under the same task and trial identifiers. Do not add a fifth eligible case. The case manifest still contains four customer jobs, and a second effect on C01 would require investigation.

Go deeper

Key takeaways

  • Join observed results to the expected manifest.
  • Inspect unauthorized effects even when the final answer is correct.
  • Keep missing outcomes and their known spend visible.

Check yourself

Answer before you look. Recalling it is what makes it stick; recognising it does not.

  1. 1A run returns three results for four eligible cases and two are correct. What is the main success fraction?

  2. 2Can a quality judge forgive an unauthorized replacement if the final response is helpful?

Sign in to track which lessons you have finished.