A 6% word error rate is a good number that measures the wrong thing. Per-digit accuracy of 98% loses 18% of ten-digit account numbers, and no aggregate metric contains that. Talking over callers is a turn-taking problem, not a model problem, and it has its own metrics too.
Your voice agent talks over callers and mishears account numbers, yet word error rate is 6%. Diagnose both, and say what to measure instead.
A 6% word error rate is a good number that measures the wrong thing. Per-digit accuracy of 98% loses 18% of ten-digit account numbers, and no aggregate metric contains that. Talking over callers is a turn-taking problem, not a model problem, and it has its own metrics too.
Updated Sep 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
The strong candidate separates the two symptoms into two mechanisms (endpointing and entity recognition) and then attacks the metric, because both failures are invisible to overall WER. The compounding arithmetic on digit strings is the thing to say out loud. Candidates who propose 'a better speech model' for both have not located either failure.
No comments yet — be the first to share your approach.
