Not leakage, not skew, not drift: the model learned exactly what you asked, and you asked for the wrong thing. The proxy-label mismatch that passes every test, the audit that catches it, and why this is the failure no metric can see.
The model aced every offline eval and the customer says it's useless. The labels look fine. What now?
Not leakage, not skew, not drift: the model learned exactly what you asked, and you asked for the wrong thing. The proxy-label mismatch that passes every test, the audit that catches it, and why this is the failure no metric can see.
Updated Aug 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
This is the failure mode that survives leakage checks, skew replay, and drift monitoring, because the pipeline is correct and the label is the bug. The screen is whether the candidate audits the label definition against the business outcome rather than re-checking the model, and the reserved probe is label noise versus label bias: random noise caps accuracy, but a systematic proxy mismatch makes a high-scoring model actively wrong, which is far more dangerous and far harder to see.
No comments yet — be the first to share your approach.
