FDEInterviews logo
🧠 Foundations of LLMs & GenAI
Core

Policy Optimization: PPO and GRPO

PPO and GRPO are the reinforcement-learning optimizers that turn a reward signal into model weight updates during RLHF and reasoning training. PPO uses a clipped objective and a KL anchor to keep updates stable and close to the reference model; GRPO drops the learned value network and instead normalizes rewards across a group of sampled answers. FDE loops probe these because the KL anchor, reward hacking, and GRPO's cost savings are where reasoning pipelines actually succeed or break.

a free account unlocks the core curriculum tier · no card
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS