Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q24IntermediateConcept

Explain RLHF: the reward model and PPO stages.

30-second answerSay your answer out loud first, then reveal.
The RLHF pipeline: SFT on demonstrations gives the policy π_SFT, which generates two or more responses per prompt that humans rank to train a reward model; PPO then maximises the reward minus a KL penalty to π_SFT, giving the aligned policy.

Reward model

  • Usually the same architecture as the LLM, with a scalar head.
  • Trained with a pairwise (Bradley–Terry) loss: score(chosen) should exceed score(rejected).

PPO stage

  • Generate responses → score them with the RM → update the policy to increase expected reward.
  • KL penalty to the reference (SFT) model prevents reward hacking (exploiting RM blind spots, e.g. getting verbose because longer answers scored higher) and keeps the language fluent.
  • Needs several models in memory (policy, reference, reward model, value/critic model), which makes it complex and expensive.

Why preferences instead of demonstrations: for humans it's easier to judge which answer is better than to write the perfect answer. And preferences capture subtle qualities (helpfulness, tone, safety).

Known issues: reward hacking, sycophancy (models learn that agreeing pleases raters), annotator bias, cost of human labels. RLAIF / Constitutional AI uses AI feedback guided by principles to scale labelling.

This is what real progress feels like.