1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are reasoning models, and what is test-time compute?
30-second answerSay your answer out loud first, then reveal.
How they differ from standard chat models
- They produce a "thinking" phase (sometimes hidden or summarised) before the final answer.
- Training uses RL with rewards for correct final answers on maths, code and logic, where correctness can be checked automatically (Q43).
- They show emergent behaviours like backtracking ("wait, that's wrong...") and verification.
Ways to spend test-time compute
- Longer reasoning traces (thinking budget).
- Parallel sampling + selection: generate N answers, then majority-vote or pick with a verifier.
- Search: tree search over reasoning steps guided by a reward model.
Trade-offs
| Good for | Not worth it for |
|---|---|
| Maths, competitive coding, complex debugging | Simple lookups, classification, chit-chat |
| Multi-step planning, agent tasks | Latency-critical UX |
| Hard analysis with verifiable answers | Very high-volume, cheap tasks |
Thinking tokens are billed and add latency, so many APIs expose a "reasoning effort" or thinking-budget control. Route by task difficulty.
Common mistakes
- Assuming the visible reasoning text is a faithful explanation of how the model got its answer. It can be incomplete or post-hoc.
Related
Every expert started right here.