FDEInterviews logo
RAG & Agent System Design / 65
hardNewSierraDecagonAnthropic

Design a simulated-customer regression suite for a support agent in production. What does it catch, what does it miss, how big must it be?

A nightly suite of fifty simulated conversations per intent catches a collapsed tool with certainty and a five-point drift almost never. Knowing which is which, and sizing each layer for what it can actually see, is the whole design. The grader checks outcomes, not transcripts, or the agent learns to sound resolved.

Updated Sep 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.

A nightly suite of fifty simulated conversations per intent catches a collapsed tool with certainty and a five-point drift almost never. Knowing which is which, and sizing each layer for what it can actually see, is the whole design. The grader checks outcomes, not transcripts, or the agent learns to sound resolved.

20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 517 remaining answers · ₹2,000 / $25
UP NEXT ON YOUR JOURNEY
FEDITOR'S NOTE

This is the follow-on to the live-metric triage question and it is asked at the same companies. The strong candidate separates two jobs the suite is asked to do (catch a break overnight, measure slow drift) and gives them different instruments with the arithmetic that says why. The weak candidate proposes 'run a thousand simulated conversations' with no grader design and no statement of what the number means.

DISCUSSION · 0

No comments yet — be the first to share your approach.