The eval set says recall fell 2.3 points, which is inside the noise you would expect from a re-index. Support says the product got worse. Both are right, and the reason is that one document class lost 49 points of recall while everything else improved slightly.
A chunking change shipped last week. Aggregate recall barely moved, and support says answers got worse. Find the regression.
The eval set says recall fell 2.3 points, which is inside the noise you would expect from a re-index. Support says the product got worse. Both are right, and the reason is that one document class lost 49 points of recall while everything else improved slightly.
Updated Sep 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
This is a statistics question hiding in a retrieval costume, and the candidates who do well say so early. The aggregate is a weighted average, so a segment worth 8% of traffic can lose half its recall and move the headline by two points. The number worth carrying into the interview is the power calculation: the drop needs about eleven thousand queries per arm to detect in aggregate and sixteen inside the right stratum. Anyone who reaches for a bigger eval set rather than a partitioned one is solving it the expensive way.
No comments yet — be the first to share your approach.
