76Your refund agent succeeds on 70% of test tasks. The customer asks if it is reliable enough to ship. How do you answer?▼hardNewSierraAnthropicOpenAI2 replies◆ premium70 percent is an average over tasks and runs, and it hides the number the customer cares about: does the same request get the same right outcome every time? Two agents with identical averages can differ from shippable to unusable.Open full answer →
77Users say your assistant forgets what they told it, and sometimes remembers things wrong. How do you evaluate its long-term memory?▼hardNewOpenAIAnthropicSierra2 replies◆ premiumSingle-session evals pass while users complain, because memory fails across sessions and at three separate stages. The answer is a planted-fact harness that names which stage lost the fact, plus the two tests teams skip: forgetting on request and leaking across users.Open full answer →
78Your customer-support agent is going to production next month. Define its SLOs, its error budget and what happens when the budget runs out.▼hardNewSierraDecagonIntercom2 replies◆ premiumUptime and latency are the easy SLIs. The hard ones are whether the case was actually resolved, which you only learn days later, and what to do with an action that must never happen, which no error budget should cover.Open full answer →