1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain RLHF: the reward model and PPO stages.
30-second answerSay your answer out loud first, then reveal.

Reward model
- Usually the same architecture as the LLM, with a scalar head.
- Trained with a pairwise (Bradley–Terry) loss: score(chosen) should exceed score(rejected).
PPO stage
- Generate responses → score them with the RM → update the policy to increase expected reward.
- KL penalty to the reference (SFT) model prevents reward hacking (exploiting RM blind spots, e.g. getting verbose because longer answers scored higher) and keeps the language fluent.
- Needs several models in memory (policy, reference, reward model, value/critic model), which makes it complex and expensive.
Why preferences instead of demonstrations: for humans it's easier to judge which answer is better than to write the perfect answer. And preferences capture subtle qualities (helpfulness, tone, safety).
Known issues: reward hacking, sycophancy (models learn that agreeing pleases raters), annotator bias, cost of human labels. RLAIF / Constitutional AI uses AI feedback guided by principles to scale labelling.
Related
This is what real progress feels like.