1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is a Large Language Model, and how does "predicting the next token" lead to useful behaviour?
30-second answerSay your answer out loud first, then reveal.

Generation is autoregressive: the model produces one token, appends it to the input, and runs again until it emits a stop token or hits a length limit.
Why next-token prediction is powerful: to predict the next word of "The derivative of x² is ___", a solution to a bug report, or the next line of a legal argument, the model must internalise the patterns behind them. The objective is simple, but data at scale makes it rich.
Three training stages (high level)
- Pretraining: next-token prediction on trillions of tokens. Produces a base model that continues text but doesn't reliably follow instructions.
- Supervised fine-tuning (SFT): trains on (instruction, good response) examples so it behaves like an assistant.
- Preference tuning / RL (RLHF, DPO, RL with verifiable rewards): aligns outputs with human preferences and improves reasoning.
Common mistakes
- Saying the model "looks up answers in a database." It doesn't. Knowledge is stored implicitly in the weights, which is why it can be wrong.
- Saying it "just predicts the next word, so it can't reason." That undersells what's needed to predict well, but it is fair to say it has no built-in fact-checking.
Follow-ups to expect
- What exactly is a token? (Q2.)
- How does it decide which token to pick? (Q8.)
Related
Little by little, you're building something great.