FDEInterviews logo
🛡️ AI Security, Privacy & Governance
Advanced

Mechanistic Interpretability

Mechanistic interpretability tries to reverse-engineer the actual computations inside a model rather than treating it as a black box: finding the features it represents and the circuits that combine them. The current toolkit centers on sparse autoencoders that decompose dense activations into interpretable features, causal tests like activation patching that prove a feature matters, and steering that turns a behavior up or down at inference. Be honest in interviews: nobody can fully explain a frontier model, you cannot prove a behavior is absent, and feature labels are human guesses.

Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS