top_p and top_k
top_p and top_k are sampling settings that remove unlikely tokens before the random pick: top_k keeps a fixed number of the likeliest tokens, and top_p keeps the likeliest tokens until their probabilities add up to p.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
A high Temperature gives every token a chance, the bad ones included. Cutting the unlikely tail first keeps the variety and drops the nonsense.
Syntax: client.chat.completions.create(..., temperature=1.8, top_p=0.1).
softmax and the scores, from Next-token prediction
import math
def softmax(scores):
exps = {word: math.exp(score) for word, score in scores.items()}
total = sum(exps.values())
return {word: e / total for word, e in exps.items()}
# made-up scores for the word after "The cat sat on the"
scores = {"mat": 4.0, "floor": 3.1, "sofa": 2.6, "roof": 1.2, "moon": -1.5}Cutting the list at 90% and at two words
probs = softmax(scores)
ranked = sorted(probs, key=probs.get, reverse=True) # likeliest first
kept, total = [], 0.0
for word in ranked:
if total >= 0.9:
break
kept.append(word)
total += probs[word]
print("top_p=0.9 keeps", kept, "holding", round(total, 3))
print("top_k=2 keeps", ranked[:2])top_p=0.9 keeps ['mat', 'floor', 'sofa'] holding 0.962 top_k=2 keeps ['mat', 'floor']
- top_p=0.9 keeps
mat,floorandsofa: the first two hold less than 0.9, and addingsofatakes the total to 0.962. - top_k=2 keeps the two likeliest words, whatever their share.
- roof and moon can no longer be picked with either setting.
Asking at temperature 1.8 with top_p 1.0 and 0.1
This continues the file from Temperature:
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
prompt = [{"role": "user", "content": "Suggest a name for a small online shop that sells tea. Reply with the name only."}]for top_p in (1.0, 0.1):
names = []
for _ in range(3):
response = client.chat.completions.create(
model="openai/gpt-oss-120b", messages=prompt, temperature=1.8, top_p=top_p
)
names.append(response.choices[0].message.content)
print(top_p, names)1.0 ['Steeping Serenity', 'Steeped Serenity', 'Leaf & Loop Teahouse'] 0.1 ['Steep & Sip Boutique', 'Steep & Sip Boutique', 'Steep & Sip Boutique']
- top_p 1.0 cuts nothing, so temperature 1.8 gives three different names.
- top_p 0.1 keeps only the likeliest tokens, and all three names are
Steep & Sip Boutique, the temperature-0 name, even at temperature 1.8.
Passing top_k to Groq
client.chat.completions.create(model="openai/gpt-oss-120b", messages=prompt, top_k=5)TypeError: Completions.create() got an unexpected keyword argument 'top_k'
The Groq client has no top_k argument, so the call fails before anything is sent. The model list says which sampling settings a model takes:
for model in client.models.list().data:
if model.id == "openai/gpt-oss-120b":
print(model.supported_sampling_parameters)['temperature', 'top_p', 'stop', 'seed', 'max_tokens']
top_p is there and top_k is not. Some other APIs, such as Anthropic's and Gemini's, do accept top_k.
top_p vs top_k
| top_p | top_k | |
|---|---|---|
| Keeps | The likeliest tokens up to a share of probability | A fixed number of the likeliest tokens |
| When the model is sure | Keeps very few tokens | Still keeps k tokens |
| When the model is unsure | Keeps many tokens | Still keeps k tokens |
| On Groq | Yes | No |
When to set top_p
- With a high temperature, to keep variety without rare, broken tokens.
- Lowered towards 0.1 when you want nearly the same answer every time but must keep a temperature above 0.
- Left at its default when you already control the output with temperature.
Related
- Previous: Temperature
- Next: Max tokens and finish_reason
- Reference: Groq chat completions API
- Change
0.9to0.5in both places in the first example. How many words are kept? - Run the live example with
top_pvalues(0.5, 0.9). - Print the whole model entry for
openai/gpt-oss-120band findsupported_features.
You understood something today that you didn't yesterday.