Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q25IntermediateConcept

What is DPO, and how does it compare to RLHF with PPO?

30-second answerSay your answer out loud first, then reveal.

DPO loss (intuition)

text
loss = −log σ( β · [ (log π(chosen) − log π_ref(chosen))
                     − (log π(rejected) − log π_ref(rejected)) ] )

Make the policy prefer "chosen" over "rejected" more than the reference model does, scaled by β (which plays the role of the KL penalty).

RLHF (PPO) vs DPO

RLHF (PPO)DPO
Reward modelSeparate, trained firstImplicit
TrainingOnline RL: sample, score, updateOffline supervised-style on fixed pairs
Models in memoryPolicy, ref, RM, criticPolicy + ref
Stability / tuningFinickySimpler
ExplorationGenerates new samples during trainingLimited to the dataset (offline)

Variants and relatives: IPO, KTO (works with thumbs-up/down instead of pairs), ORPO (combines SFT and preference loss), SimPO, and online/iterative DPO (generate new pairs with the current model). For reasoning, GRPO-style RL with verifiable rewards is widely used (Q43).

Practical tip: DPO quality depends heavily on preference data quality and on staying on-distribution. Pairs generated by the model being trained work better than pairs from some other model.

Every expert started right here.