FDEInterviews logo
LLM & GenAI Fundamentals / 57
hardNewAnthropicOpenAI

Here are 30 examples where our prompt gets the wrong answer. Improve it, show me the eval before and after, and do not overfit to the 30.

This is the live practical several labs run, and it is scored on protocol more than on the prompt you end with. Read all thirty before editing, cluster them, hold some out, change one thing at a time, and report with intervals, because 24 of 30 fixed is somewhere between 63% and 91%.

Updated Sep 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.

This is the live practical several labs run, and it is scored on protocol more than on the prompt you end with. Read all thirty before editing, cluster them, hold some out, change one thing at a time, and report with intervals, because 24 of 30 fixed is somewhere between 63% and 91%.

20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 517 remaining answers · ₹2,000 / $25
UP NEXT ON YOUR JOURNEY
FEDITOR'S NOTE

The candidates who do well treat the thirty as symptoms and the eval as the deliverable. The ones who do badly open the prompt in the first minute, fix example one, and by minute forty have a prompt that contains most of the thirty examples verbatim and a held-out score that went down. The interviewer is watching for the split, the regression set, and whether the final report admits what did not get fixed.

DISCUSSION · 0

No comments yet — be the first to share your approach.