1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain temperature, top-k and top-p (nucleus) sampling.
30-second answerSay your answer out loud first, then reveal.
Temperature: probabilities = softmax(logits / T).
- T → 0: always the most likely token (greedy).
- T = 1: the model's raw distribution.
- T > 1: flatter, more surprising choices, more errors.
Example: next-token probabilities {"blue": 0.6, "clear": 0.25, "grey": 0.1, "purple": 0.05}.
Top-k=2 → sample from {blue, clear}, renormalised.
Top-p=0.9 → {blue, clear, grey} (0.6+0.25+0.1 = 0.95 ≥ 0.9).
Top-k=2 → sample from {blue, clear}, renormalised.
Top-p=0.9 → {blue, clear, grey} (0.6+0.25+0.1 = 0.95 ≥ 0.9).
Top-p vs top-k: top-p adapts to the model's confidence. When the model is sure, the nucleus is tiny; when unsure, it widens. Top-k uses a fixed count regardless.
Other decoding controls
- Greedy: pick the argmax each step. Deterministic, but can be repetitive.
- Beam search: keep the top-B partial sequences. Common in translation, rarely used in chat LLMs.
- Repetition / frequency / presence penalties: discourage repeated tokens (Q29).
- Min-p: keep tokens with probability ≥ min_p × top probability.
- Stop sequences, max tokens.
Practical defaults: extraction / JSON / code → temperature 0–0.3; chat → ~0.7; brainstorming / creative → 0.8–1.0. Usually tune temperature or top-p, not both aggressively.
Related
Slow is fine. Stopping is the only problem.