59Design the eval for a voice agent taking inbound support calls. What do you measure offline, and what can only production tell you?▼hardNewSierraDecagonCresta2 replies◆ premiumWord error rate is the number every vendor quotes and the one that predicts the least. The offline suite is built from confirmed real calls with noise and codecs applied, scored on entity accuracy, tail latency and false interruptions. Production adds what a recording cannot simulate: real callers and the reopen rate.Open full answer →
64Your support agent resolved 69% of conversations last month and 60% this week. Nobody deployed anything. What do you check first?▼hardNewSierraDecagon2 replies◆ premiumNobody deployed anything is true of your repository and false of the system. The model provider, a tool behind an API, the knowledge base and the traffic mix all change without a commit. Segmentation comes before theory: a broken intent and a mix shift produce the same headline and need different fixes.Open full answer →
65Design a simulated-customer regression suite for a support agent in production. What does it catch, what does it miss, how big must it be?▼hardNewSierraDecagonAnthropic2 replies◆ premiumA nightly suite of fifty simulated conversations per intent catches a collapsed tool with certainty and a five-point drift almost never. Knowing which is which, and sizing each layer for what it can actually see, is the whole design. The grader checks outcomes, not transcripts, or the agent learns to sound resolved.Open full answer →