1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is DPO, and how does it compare to RLHF with PPO?
30-second answerSay your answer out loud first, then reveal.
DPO loss (intuition)
loss = −log σ( β · [ (log π(chosen) − log π_ref(chosen))
− (log π(rejected) − log π_ref(rejected)) ] )Make the policy prefer "chosen" over "rejected" more than the reference model does, scaled by β (which plays the role of the KL penalty).
RLHF (PPO) vs DPO
| RLHF (PPO) | DPO | |
|---|---|---|
| Reward model | Separate, trained first | Implicit |
| Training | Online RL: sample, score, update | Offline supervised-style on fixed pairs |
| Models in memory | Policy, ref, RM, critic | Policy + ref |
| Stability / tuning | Finicky | Simpler |
| Exploration | Generates new samples during training | Limited to the dataset (offline) |
Variants and relatives: IPO, KTO (works with thumbs-up/down instead of pairs), ORPO (combines SFT and preference loss), SimPO, and online/iterative DPO (generate new pairs with the current model). For reasoning, GRPO-style RL with verifiable rewards is widely used (Q43).
Practical tip: DPO quality depends heavily on preference data quality and on staying on-distribution. Pairs generated by the model being trained work better than pairs from some other model.
Related
Every expert started right here.