LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Generating text

Text generation is a loop: the model predicts one token, adds it to the text, and predicts again, until it writes an end token or reaches a length limit.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

Next-token prediction predicted one word. Pick it, add it to the text, predict again: every answer from every language model is written this way.

A table of next words

python
# a hand-written stand-in for a model: which word may follow which, and how likely
next_word = {
    "Your": {"refund": 0.6, "parcel": 0.4},
    "refund": {"is": 0.7, "was": 0.3},
    "is": {"on": 0.6, "late": 0.4},
    "on": {"its": 0.8, "the": 0.2},
    "its": {"way.": 0.9, "own": 0.1},
}

This dictionary stands in for a model: for a few words, it lists which word may follow and how likely each is. A real model looks at all the text so far, not only the last word, and scores every token in its vocabulary.

Writing a sentence one word at a time

Example
words = ["Your"]
while words[-1] in next_word:
    options = next_word[words[-1]]
    words.append(max(options, key=options.get))  # greedy: take the likeliest word
    print(" ".join(words))
  • Each line is one more pass of the loop. The likeliest next word is added and the loop runs again.
  • It stops at way. because the table has no entry for it, as a model stops when it writes its end token.
  • The same words every run. Taking the top word each time, greedy decoding, has nothing random in it.

Streaming a real answer piece by piece

Syntax: stream=True returns the answer in chunks as the model writes it; each chunk's delta.content is the new text.

ExampleAPI key
from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment

stream = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Write a one-sentence apology to a customer whose parcel is late."}],
    stream=True,
)
for chunk in stream:
    piece = chunk.choices[0].delta.content
    if piece:
        print(piece, end="|")
print()
  • Each | marks one chunk as it arrived from Groq: here one word with its leading space, and the full stop on its own.
  • The answer arrives in the order it was written. The model had nothing more to send until it had predicted the next token.
  • Some chunks carry no answer text. gpt-oss streams its reasoning first, in delta.reasoning, and the first and last chunks are empty, so the loop skips every piece without content.

Likely is not the same as true

LLM hallucination · from the What Is LLM Hallucination And How to Reduce It? video · 0:32 to 4:33

The video calls it hallucination: a model gives an answer that is not factually right, and gives it so confidently that people believe it. One cause is the training cutoff date, after which the model has no data. The video compares the model to an arrogant friend who answers every question, known or not, in a convincing way.

Brightcart is a shop name made up for this example. Nothing in either prompt says what its refund policy is:

ExampleAPI key
for prompt in [
    "What is the refund policy of Brightcart, the online shop? Answer in two sentences.",
    "Write the refund section of Brightcart's help page in two sentences.",
]:
    response = client.chat.completions.create(
        model="openai/gpt-oss-120b",
        messages=[{"role": "user", "content": prompt}],
    )
    print(prompt)
    print("  ->", response.choices[0].message.content)
  • Both answers state a policy, with a 30-day window and refunds within 5 to 7 business days.
  • Neither number came from anywhere. They are what refund policies usually say, so they were the likely tokens.
  • Both read like a real help page. A customer would believe them. This is how a confident wrong answer is made.
  • The defence is in the prompt and after it. Give the model the real facts, as System prompts does, and check what it returns, as Structured output does. The video's next step, connecting the model to documents and tools, is what retrieval (RAG) and agents do.

Streaming vs waiting for the whole answer

Waiting (default)Streaming (stream=True)
First text on screenWhen the answer is finishedAs soon as the first token is written
Total timeThe sameThe same
CodeOne response objectA loop over chunks

When to stream

  • Chat interfaces, where a person is watching the answer appear.
  • Long answers, so the reader can start before the end.
  • Not for a program that needs the whole answer before it can do anything, such as parsing JSON.
Watch out. Streaming does not make an answer arrive sooner in total; it shows each piece as soon as it exists. Latency times both.
Try it yourself
  • In next_word, change "its" to {"way.": 0.1, "own": 0.9} and run the loop. Why does it stop at own?
  • Stream a longer answer, such as a three-sentence apology, and count the chunks.
  • Put a policy in the prompt, for example "Brightcart refunds within 14 days. " before the question, and compare the answer.

Every expert started right here.