Word error rate is the number every vendor quotes and the one that predicts the least. The offline suite is built from confirmed real calls with noise and codecs applied, scored on entity accuracy, tail latency and false interruptions. Production adds what a recording cannot simulate: real callers and the reopen rate.
Design the eval for a voice agent taking inbound support calls. What do you measure offline, and what can only production tell you?
Word error rate is the number every vendor quotes and the one that predicts the least. The offline suite is built from confirmed real calls with noise and codecs applied, scored on entity accuracy, tail latency and false interruptions. Production adds what a recording cannot simulate: real callers and the reopen rate.
Updated Sep 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
The strong candidate names the metrics that move when a voice agent fails in production (entity errors, tail latency, interruptions) and knows the sizing: 400 calls to bound a 5% interruption rate to about two points. They also separate what a recorded corpus can measure from what only live traffic can. The weak candidate proposes WER and a satisfaction survey.
No comments yet — be the first to share your approach.
