Generating text
Text generation is a loop: the model predicts one token, adds it to the text, and predicts again, until it writes an end token or reaches a length limit.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Next-token prediction predicted one word. Pick it, add it to the text, predict again: every answer from every language model is written this way.
A table of next words
# a hand-written stand-in for a model: which word may follow which, and how likely
next_word = {
"Your": {"refund": 0.6, "parcel": 0.4},
"refund": {"is": 0.7, "was": 0.3},
"is": {"on": 0.6, "late": 0.4},
"on": {"its": 0.8, "the": 0.2},
"its": {"way.": 0.9, "own": 0.1},
}This dictionary stands in for a model: for a few words, it lists which word may follow and how likely each is. A real model looks at all the text so far, not only the last word, and scores every token in its vocabulary.
Writing a sentence one word at a time
words = ["Your"]
while words[-1] in next_word:
options = next_word[words[-1]]
words.append(max(options, key=options.get)) # greedy: take the likeliest word
print(" ".join(words))Your refund Your refund is Your refund is on Your refund is on its Your refund is on its way.
- Each line is one more pass of the loop. The likeliest next word is added and the loop runs again.
- It stops at way. because the table has no entry for it, as a model stops when it writes its end token.
- The same words every run. Taking the top word each time, greedy decoding, has nothing random in it.
Streaming a real answer piece by piece
Syntax: stream=True returns the answer in chunks as the model writes it; each chunk's delta.content is the new text.
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
stream = client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[{"role": "user", "content": "Write a one-sentence apology to a customer whose parcel is late."}],
stream=True,
)
for chunk in stream:
piece = chunk.choices[0].delta.content
if piece:
print(piece, end="|")
print()We|’re| very| sorry| for| the| delay| in| delivering| your| parcel| and| appreciate| your| patience| while| we| work| to| get| it| to| you| as| quickly| as| possible|.|
- Each
|marks one chunk as it arrived from Groq: here one word with its leading space, and the full stop on its own. - The answer arrives in the order it was written. The model had nothing more to send until it had predicted the next token.
- Some chunks carry no answer text. gpt-oss streams its reasoning first, in
delta.reasoning, and the first and last chunks are empty, so the loop skips every piece without content.
Likely is not the same as true
The video calls it hallucination: a model gives an answer that is not factually right, and gives it so confidently that people believe it. One cause is the training cutoff date, after which the model has no data. The video compares the model to an arrogant friend who answers every question, known or not, in a convincing way.
Brightcart is a shop name made up for this example. Nothing in either prompt says what its refund policy is:
for prompt in [
"What is the refund policy of Brightcart, the online shop? Answer in two sentences.",
"Write the refund section of Brightcart's help page in two sentences.",
]:
response = client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[{"role": "user", "content": prompt}],
)
print(prompt)
print(" ->", response.choices[0].message.content)What is the refund policy of Brightcart, the online shop? Answer in two sentences. -> Brightcart accepts returns and issues full refunds for items that are returned within 30 days of delivery, provided they are unused, in their original condition and packaging, and accompanied by proof of purchase. Refunds are processed to the original payment method within 5‑7 business days after the returned merchandise is received and inspected. Write the refund section of Brightcart's help page in two sentences. -> Our refund policy allows you to request a full refund within 30 days of purchase if you’re unsatisfied with your Brightcart order, provided the items are unused and in their original condition; after this period, refunds are granted only for defective or damaged products. To initiate a refund, simply contact our support team through the “Help Center” with your order number and a brief description of the issue, and we’ll process it within 5–7 business days.
- Both answers state a policy, with a 30-day window and refunds within 5 to 7 business days.
- Neither number came from anywhere. They are what refund policies usually say, so they were the likely tokens.
- Both read like a real help page. A customer would believe them. This is how a confident wrong answer is made.
- The defence is in the prompt and after it. Give the model the real facts, as System prompts does, and check what it returns, as Structured output does. The video's next step, connecting the model to documents and tools, is what retrieval (RAG) and agents do.
Streaming vs waiting for the whole answer
| Waiting (default) | Streaming (stream=True) | |
|---|---|---|
| First text on screen | When the answer is finished | As soon as the first token is written |
| Total time | The same | The same |
| Code | One response object | A loop over chunks |
When to stream
- Chat interfaces, where a person is watching the answer appear.
- Long answers, so the reader can start before the end.
- Not for a program that needs the whole answer before it can do anything, such as parsing JSON.
Related
- Previous: Next-token prediction
- Next: Temperature
- Reference: Groq text generation and streaming
- In
next_word, change"its"to{"way.": 0.1, "own": 0.9}and run the loop. Why does it stop atown? - Stream a longer answer, such as a three-sentence apology, and count the chunks.
- Put a policy in the prompt, for example
"Brightcart refunds within 14 days. "before the question, and compare the answer.
Every expert started right here.