Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q8EasyConcept

Explain temperature, top-k and top-p (nucleus) sampling.

30-second answerSay your answer out loud first, then reveal.

Temperature: probabilities = softmax(logits / T).

  • T → 0: always the most likely token (greedy).
  • T = 1: the model's raw distribution.
  • T > 1: flatter, more surprising choices, more errors.
Example: next-token probabilities {"blue": 0.6, "clear": 0.25, "grey": 0.1, "purple": 0.05}.
Top-k=2 → sample from {blue, clear}, renormalised.
Top-p=0.9 → {blue, clear, grey} (0.6+0.25+0.1 = 0.95 ≥ 0.9).

Top-p vs top-k: top-p adapts to the model's confidence. When the model is sure, the nucleus is tiny; when unsure, it widens. Top-k uses a fixed count regardless.

Other decoding controls

  • Greedy: pick the argmax each step. Deterministic, but can be repetitive.
  • Beam search: keep the top-B partial sequences. Common in translation, rarely used in chat LLMs.
  • Repetition / frequency / presence penalties: discourage repeated tokens (Q29).
  • Min-p: keep tokens with probability ≥ min_p × top probability.
  • Stop sequences, max tokens.

Practical defaults: extraction / JSON / code → temperature 0–0.3; chat → ~0.7; brainstorming / creative → 0.8–1.0. Usually tune temperature or top-p, not both aggressively.

Slow is fine. Stopping is the only problem.