1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →
Temperature
Temperature is a sampling setting that divides the model's scores before they become probabilities: low values let the likeliest token win almost every time, and high values give unlikely tokens a real chance.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Greedy decoding always takes the top token. Most chat APIs instead pick at random, weighted by the probabilities, and temperature decides how uneven those weights are.
Syntax: client.chat.completions.create(model=..., messages=..., temperature=0.2). Groq accepts values from 0 to 2.
softmax and the scores, from Next-token prediction
import math
def softmax(scores):
exps = {word: math.exp(score) for word, score in scores.items()}
total = sum(exps.values())
return {word: e / total for word, e in exps.items()}
# made-up scores for the word after "The cat sat on the"
scores = {"mat": 4.0, "floor": 3.1, "sofa": 2.6, "roof": 1.2, "moon": -1.5}Dividing the scores by a temperature
for temperature in (0.2, 1.0, 1.8):
probs = softmax({word: score / temperature for word, score in scores.items()})
print(temperature, {word: round(p, 2) for word, p in probs.items()})Output
0.2 {'mat': 0.99, 'floor': 0.01, 'sofa': 0.0, 'roof': 0.0, 'moon': 0.0}
1.0 {'mat': 0.58, 'floor': 0.24, 'sofa': 0.14, 'roof': 0.04, 'moon': 0.0}
1.8 {'mat': 0.43, 'floor': 0.26, 'sofa': 0.2, 'roof': 0.09, 'moon': 0.02}- 0.2 multiplies the gaps between scores by five, and
mattakes 0.99 of the probability. - 1.0 leaves the scores as they are: the same probabilities as in Next-token prediction.
- 1.8 shrinks the gaps.
matdrops to 0.43, and evenmoonrises from 0.00 to 0.02.
Three tea-shop names at temperature 0 and 1.8
The same request is sent three times at each temperature:
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
prompt = [{"role": "user", "content": "Suggest a name for a small online shop that sells tea. Reply with the name only."}]
for temperature in (0, 1.8):
names = []
for _ in range(3):
response = client.chat.completions.create(
model="openai/gpt-oss-120b", messages=prompt, temperature=temperature
)
names.append(response.choices[0].message.content)
print(temperature, names)Output
0 ['Steep & Sip Boutique', 'Steep & Sip Boutique', 'Steep & Sip Boutique'] 1.8 ['Leaf & Sip', 'Steep & Sip', 'Steeped Serenity']
- At 0 the three names are identical:
Steep & Sip Boutiqueevery time. - At 1.8 the three names are all different, and none of them is the temperature-0 name.
Low vs high temperature
| Low (0 to 0.3) | High (1 and above) | |
|---|---|---|
| Same request, repeated | Same or nearly the same answer | Different answers |
| Unlikely tokens | Almost never picked | Picked now and then |
| Good for | Sorting, extracting, answering from given facts | Names, slogans, brainstorming |
Choosing a temperature
- 0 for sorting tickets, extracting fields and anything a program checks. Project: ticket sorter sorts at temperature 0.
- Around 0.7 to 1 for writing where some variety reads more naturally.
- Above 1 for ideas, where a surprising answer is welcome.
Watch out. Temperature changes which tokens are picked, not what the model knows. A higher temperature does not make an answer more creative in any useful sense past a point; it makes unlikely tokens likelier, and some of those are mistakes.
Related
- Previous: Generating text
- Next: top_p and top_k
- Reference: Groq chat completions API
Try it yourself
- Run the live example with
(0, 1.0)instead. How many different names do you get at 1.0? - In the first example, add
0.5to the temperatures and read how muchmatgets. - Try a temperature of
3.0in the first example. What happens to the probabilities?
Little by little, you're building something great.