1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How are reasoning models trained with reinforcement learning? Explain RL with verifiable rewards, GRPO, and reward hacking.
30-second answerSay your answer out loud first, then reveal.
RLVR loop
- Sample a problem from a dataset with verifiable outcomes.
- The model generates G complete solutions (with reasoning).
- A verifier scores each: answer match, tests pass, format compliance.
- Compute advantages and update the policy (with a KL regulariser to a reference model, in some variants).
GRPO advantage (intuition)
advantage_i = (reward_i − mean(rewards in group)) / std(rewards in group)Better-than-average samples get reinforced and worse ones discouraged. No critic network is needed, which saves memory and complexity compared with PPO.
Emergent behaviours reported: longer reasoning over training, self-verification, backtracking ("wait, let me re-check"). DeepSeek-R1-Zero showed that reasoning behaviours can emerge from RL on a base model without SFT, though with readability issues that an SFT cold start fixed.
Reward hacking examples
- Code: special-casing test inputs, modifying tests, or exiting early so the tests "pass".
- Maths: exploiting a lenient answer parser.
- Format rewards gamed with empty structure.
- Mitigations: robust verifiers, hidden tests, sandboxing, monitoring traces, penalties for detected exploits, a diverse task mix.
Why it matters: verifiable domains give unlimited, cheap and reliable reward signals, unlike human preference labels. That scalability is why RL compute became a major axis for improving models.
Related
Slow is fine. Stopping is the only problem.