Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q43HardConcept

How are reasoning models trained with reinforcement learning? Explain RL with verifiable rewards, GRPO, and reward hacking.

30-second answerSay your answer out loud first, then reveal.

RLVR loop

  1. Sample a problem from a dataset with verifiable outcomes.
  2. The model generates G complete solutions (with reasoning).
  3. A verifier scores each: answer match, tests pass, format compliance.
  4. Compute advantages and update the policy (with a KL regulariser to a reference model, in some variants).

GRPO advantage (intuition)

text
advantage_i = (reward_i − mean(rewards in group)) / std(rewards in group)

Better-than-average samples get reinforced and worse ones discouraged. No critic network is needed, which saves memory and complexity compared with PPO.

Emergent behaviours reported: longer reasoning over training, self-verification, backtracking ("wait, let me re-check"). DeepSeek-R1-Zero showed that reasoning behaviours can emerge from RL on a base model without SFT, though with readability issues that an SFT cold start fixed.

Reward hacking examples

  • Code: special-casing test inputs, modifying tests, or exiting early so the tests "pass".
  • Maths: exploiting a lenient answer parser.
  • Format rewards gamed with empty structure.
  • Mitigations: robust verifiers, hidden tests, sandboxing, monitoring traces, penalties for detected exploits, a diverse task mix.

Why it matters: verifiable domains give unlimited, cheap and reliable reward signals, unlike human preference labels. That scalability is why RL compute became a major axis for improving models.

Slow is fine. Stopping is the only problem.