FDEInterviews logo
📊 Evaluation & ML Foundations
Core

LLM-as-a-Judge

LLM-as-a-judge uses a strong model to grade outputs against an explicit rubric, so evaluation scales past the few hundred examples a human can read by hand. It only counts as evaluation once you have calibrated the judge against 50-100 human labels and reported how well it agrees, because an uncalibrated judge is just a confident opinion.

a free account unlocks the core curriculum tier · no card
RELATED CONCEPTS
LESSONS THAT TEACH THIS
PRACTICE THIS IN REAL QUESTIONS