Three correct queries scored 0 of 3 on exact match and 3 of 3 on execution. The golden set is question, reference SQL and result set, stratified by the ways queries go wrong, with abstention as its own column. A hundred cases gives plus or minus eight points; four hundred gives four.
Build the eval for a text-to-SQL feature over a customer's warehouse: what is in the golden set, what is the metric, how many cases?
Three correct queries scored 0 of 3 on exact match and 3 of 3 on execution. The golden set is question, reference SQL and result set, stratified by the ways queries go wrong, with abstention as its own column. A hundred cases gives plus or minus eight points; four hundred gives four.
Updated Sep 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
Scored on whether the candidate can say what the metric compares and why, and whether they know the size arithmetic. Strong answers build the set from questions people actually asked, including the ones that should be refused, and report per stratum. Weak answers propose exact match, or a single accuracy over a set of easy single-table questions the pilot happened to include.
No comments yet — be the first to share your approach.
