FDEInterviews logo
📊 Evaluation & ML Foundations
Core

Evaluating RAG Systems

The central rule of RAG evaluation is to score retrieval and generation separately, because they fail for different reasons and you cannot fix what you cannot isolate. Retrieval is graded against a golden set with recall@k, precision@k, MRR, and NDCG; generation is graded for faithfulness and answer relevance, usually with an LLM judge. Recall@k is the ceiling on everything downstream.

a free account unlocks the core curriculum tier · no card
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS