Next-token prediction
Next-token prediction is the one task a language model is trained for: given the text so far, give every token in the vocabulary a probability of coming next.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Answering questions, summarising, writing code: a chat model does all of it by repeating this one step. The model first gives each possible next token a score, and the scores are turned into probabilities that add up to 1.
The video's diagram says a model generates output by calculating the probabilities of the next likely token, with The cat sat on the ... and three candidates: mat, roof and sofa. It then compares this with how people reply: after Hi, you might answer Hi, Hey, Hi Mayank, Hello Mayank or Hi Sir, and you pick the one that fits best from experience.
The scores below are made up for the video's cat example, so the arithmetic fits on the page. A real model computes one score for each of the 201,088 tokens in its vocabulary, with the whole text before it in view.
Syntax: probability of a word = exp(its score) / sum of exp(every score). This formula is called softmax.
Scores for each candidate word
# made-up scores for the word after "The cat sat on the"
scores = {"mat": 4.0, "floor": 3.1, "sofa": 2.6, "roof": 1.2, "moon": -1.5}A higher score means the model thinks that word is more likely next. Scores can be any number, negative too, so on their own they are hard to read.
Turning scores into probabilities with softmax
import math
def softmax(scores):
exps = {word: math.exp(score) for word, score in scores.items()}
total = sum(exps.values())
return {word: e / total for word, e in exps.items()}math.exp makes every score positive and stretches the gaps between them. Dividing by the total makes the results add up to 1.
Probabilities for the word after The cat sat on the
probs = softmax(scores)
for word, p in probs.items():
print(f"{word:6} {p:.3f}")
print("sum:", round(sum(probs.values()), 3))mat 0.582 floor 0.237 sofa 0.144 roof 0.035 moon 0.002 sum: 1.0
What the probabilities say
- mat is the favourite at 0.582, but not certain.
- floor and sofa together hold about 0.38, a large share for words that are not the favourite.
- moon gets 0.002: small, not zero. Every word keeps some chance.
- sum: 1.0 confirms the five probabilities add up to 1.
Picking the likeliest word
best = max(probs, key=probs.get)
print("the likeliest next word:", best)the likeliest next word: mat
Always taking the top word is called greedy decoding. It gives the same answer every time. Chat APIs usually pick at random instead, weighted by these probabilities, which is why the same question can get different answers; Temperature covers that choice.
Scores vs probabilities
| Scores (logits) | Probabilities | |
|---|---|---|
| Range | Any number | 0 to 1 |
| Add up to | Anything | 1 |
| Readable as a chance | No | Yes |
| Made by | The model | softmax over the scores |
Where next-token probabilities show up
- Sampling settings. Temperature and top_p reshape or cut these probabilities before a word is picked.
- Prompts. Every token in the prompt changes the scores of the next one; that is the only way a prompt has any effect.
- Wrong answers. The model picks likely words, and a likely word is not always a true one (Generating text).
logprobs=True is refused with a 400 error. Some other providers and models return them; check the model's page before building on them.Related
- Previous: Tokens
- Next: Generating text
- Reference: Softmax function
- Change the score of
"moon"to5.0and run it again. Which word is likeliest now? - Add
"bed": 2.9toscoresand see how the other probabilities shrink. - Halve every score before calling
softmax. Do the probabilities spread out or bunch up?
This is what real progress feels like.