LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

top_p and top_k

top_p and top_k are sampling settings that remove unlikely tokens before the random pick: top_k keeps a fixed number of the likeliest tokens, and top_p keeps the likeliest tokens until their probabilities add up to p.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

A high Temperature gives every token a chance, the bad ones included. Cutting the unlikely tail first keeps the variety and drops the nonsense.

Syntax: client.chat.completions.create(..., temperature=1.8, top_p=0.1).

softmax and the scores, from Next-token prediction

python
import math


def softmax(scores):
    exps = {word: math.exp(score) for word, score in scores.items()}
    total = sum(exps.values())
    return {word: e / total for word, e in exps.items()}

# made-up scores for the word after "The cat sat on the"
scores = {"mat": 4.0, "floor": 3.1, "sofa": 2.6, "roof": 1.2, "moon": -1.5}

Cutting the list at 90% and at two words

Example
probs = softmax(scores)
ranked = sorted(probs, key=probs.get, reverse=True)  # likeliest first

kept, total = [], 0.0
for word in ranked:
    if total >= 0.9:
        break
    kept.append(word)
    total += probs[word]

print("top_p=0.9 keeps", kept, "holding", round(total, 3))
print("top_k=2 keeps", ranked[:2])
  • top_p=0.9 keeps mat, floor and sofa: the first two hold less than 0.9, and adding sofa takes the total to 0.962.
  • top_k=2 keeps the two likeliest words, whatever their share.
  • roof and moon can no longer be picked with either setting.

Asking at temperature 1.8 with top_p 1.0 and 0.1

This continues the file from Temperature:

python
from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment
prompt = [{"role": "user", "content": "Suggest a name for a small online shop that sells tea. Reply with the name only."}]
ExampleAPI key
for top_p in (1.0, 0.1):
    names = []
    for _ in range(3):
        response = client.chat.completions.create(
            model="openai/gpt-oss-120b", messages=prompt, temperature=1.8, top_p=top_p
        )
        names.append(response.choices[0].message.content)
    print(top_p, names)
  • top_p 1.0 cuts nothing, so temperature 1.8 gives three different names.
  • top_p 0.1 keeps only the likeliest tokens, and all three names are Steep & Sip Boutique, the temperature-0 name, even at temperature 1.8.

Passing top_k to Groq

ExampleAPI key
client.chat.completions.create(model="openai/gpt-oss-120b", messages=prompt, top_k=5)

The Groq client has no top_k argument, so the call fails before anything is sent. The model list says which sampling settings a model takes:

ExampleAPI key
for model in client.models.list().data:
    if model.id == "openai/gpt-oss-120b":
        print(model.supported_sampling_parameters)

top_p is there and top_k is not. Some other APIs, such as Anthropic's and Gemini's, do accept top_k.

top_p vs top_k

top_ptop_k
KeepsThe likeliest tokens up to a share of probabilityA fixed number of the likeliest tokens
When the model is sureKeeps very few tokensStill keeps k tokens
When the model is unsureKeeps many tokensStill keeps k tokens
On GroqYesNo

When to set top_p

  • With a high temperature, to keep variety without rare, broken tokens.
  • Lowered towards 0.1 when you want nearly the same answer every time but must keep a temperature above 0.
  • Left at its default when you already control the output with temperature.
Watch out. Change temperature or top_p, not both at once. With two settings moving, you cannot tell which one changed the answer.
Try it yourself
  • Change 0.9 to 0.5 in both places in the first example. How many words are kept?
  • Run the live example with top_p values (0.5, 0.9).
  • Print the whole model entry for openai/gpt-oss-120b and find supported_features.
PreviousTemperature

You understood something today that you didn't yesterday.