Train a reward model for a coding agent from 100K noisy human scores biased toward short solutions.
Human 1-5 scores are biased toward short, simple code. Debias the labels, fuse test-pass rate without letting the agent game it, and pick PPO with a fused reward over a DPO ranking loss. Plus the hard cap on the action space so the agent can't delete the repo.
Updated Aug 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
Human 1-5 scores are biased toward short, simple code. Debias the labels, fuse test-pass rate without letting the agent game it, and pick PPO with a fused reward over a DPO ranking loss. Plus the hard cap on the action space so the agent can't delete the repo.
Lead with where the obvious approach breaks, because that is the judgment they are screening for — most candidates jump straight to the happy path and lose the room.
Then walk the failure back through the pipeline in order, naming the one metric the customer's exec sponsor actually cares about before you propose the fix.