When can you trust LLM-as-judge, and how do you calibrate it against human labels?
Position bias, verbosity bias, self-preference, and the agreement-rate workflow that turns a sloppy judge into eval infrastructure you can defend. AI labs probe this hard; here's the calibrated answer.
Updated Aug 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
Position bias, verbosity bias, self-preference, and the agreement-rate workflow that turns a sloppy judge into eval infrastructure you can defend. AI labs probe this hard; here's the calibrated answer.
Lead with where the obvious approach breaks, because that is the judgment they are screening for — most candidates jump straight to the happy path and lose the room.
Then walk the failure back through the pipeline in order, naming the one metric the customer's exec sponsor actually cares about before you propose the fix.