LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Temperature

Temperature is a sampling setting that divides the model's scores before they become probabilities: low values let the likeliest token win almost every time, and high values give unlikely tokens a real chance.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

Greedy decoding always takes the top token. Most chat APIs instead pick at random, weighted by the probabilities, and temperature decides how uneven those weights are.

Syntax: client.chat.completions.create(model=..., messages=..., temperature=0.2). Groq accepts values from 0 to 2.

softmax and the scores, from Next-token prediction

python
import math


def softmax(scores):
    exps = {word: math.exp(score) for word, score in scores.items()}
    total = sum(exps.values())
    return {word: e / total for word, e in exps.items()}

# made-up scores for the word after "The cat sat on the"
scores = {"mat": 4.0, "floor": 3.1, "sofa": 2.6, "roof": 1.2, "moon": -1.5}

Dividing the scores by a temperature

Example
for temperature in (0.2, 1.0, 1.8):
    probs = softmax({word: score / temperature for word, score in scores.items()})
    print(temperature, {word: round(p, 2) for word, p in probs.items()})
  • 0.2 multiplies the gaps between scores by five, and mat takes 0.99 of the probability.
  • 1.0 leaves the scores as they are: the same probabilities as in Next-token prediction.
  • 1.8 shrinks the gaps. mat drops to 0.43, and even moon rises from 0.00 to 0.02.

Three tea-shop names at temperature 0 and 1.8

The same request is sent three times at each temperature:

ExampleAPI key
from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment
prompt = [{"role": "user", "content": "Suggest a name for a small online shop that sells tea. Reply with the name only."}]

for temperature in (0, 1.8):
    names = []
    for _ in range(3):
        response = client.chat.completions.create(
            model="openai/gpt-oss-120b", messages=prompt, temperature=temperature
        )
        names.append(response.choices[0].message.content)
    print(temperature, names)
  • At 0 the three names are identical: Steep & Sip Boutique every time.
  • At 1.8 the three names are all different, and none of them is the temperature-0 name.

Low vs high temperature

Low (0 to 0.3)High (1 and above)
Same request, repeatedSame or nearly the same answerDifferent answers
Unlikely tokensAlmost never pickedPicked now and then
Good forSorting, extracting, answering from given factsNames, slogans, brainstorming

Choosing a temperature

  • 0 for sorting tickets, extracting fields and anything a program checks. Project: ticket sorter sorts at temperature 0.
  • Around 0.7 to 1 for writing where some variety reads more naturally.
  • Above 1 for ideas, where a surprising answer is welcome.
Watch out. Temperature changes which tokens are picked, not what the model knows. A higher temperature does not make an answer more creative in any useful sense past a point; it makes unlikely tokens likelier, and some of those are mistakes.
Try it yourself
  • Run the live example with (0, 1.0) instead. How many different names do you get at 1.0?
  • In the first example, add 0.5 to the temperatures and read how much mat gets.
  • Try a temperature of 3.0 in the first example. What happens to the probabilities?

Little by little, you're building something great.