Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q15EasyConcept

What are reasoning models, and what is test-time compute?

30-second answerSay your answer out loud first, then reveal.

How they differ from standard chat models

  • They produce a "thinking" phase (sometimes hidden or summarised) before the final answer.
  • Training uses RL with rewards for correct final answers on maths, code and logic, where correctness can be checked automatically (Q43).
  • They show emergent behaviours like backtracking ("wait, that's wrong...") and verification.

Ways to spend test-time compute

  1. Longer reasoning traces (thinking budget).
  2. Parallel sampling + selection: generate N answers, then majority-vote or pick with a verifier.
  3. Search: tree search over reasoning steps guided by a reward model.

Trade-offs

Good forNot worth it for
Maths, competitive coding, complex debuggingSimple lookups, classification, chit-chat
Multi-step planning, agent tasksLatency-critical UX
Hard analysis with verifiable answersVery high-volume, cheap tasks

Thinking tokens are billed and add latency, so many APIs expose a "reasoning effort" or thinking-budget control. Route by task difficulty.

Common mistakes

  • Assuming the visible reasoning text is a faithful explanation of how the model got its answer. It can be incomplete or post-hoc.

Every expert started right here.