Chat messages and templates
A chat message is a dictionary with a role, system, user or assistant, and its content; a chat model reads the whole list of messages in every request and remembers nothing between requests.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Every example so far sent a list with one user message. A conversation is the same list with more messages in it, and it is your code, not the model, that keeps the list.
The memory section of the video starts from one fact: LLMs are stateless. Across API requests the model has no idea what it did a second ago, and the video stresses that this is how they are built, not a bug. The fix is a list: every user message and every reply is appended to it, and the whole list is sent again on the next call. Each message has a strict role: system for the instructions, user for the person, assistant for the model.
The video's notebook calls OpenAI's API. The requests below send the same kind of list to Groq.
Syntax: messages=[{"role": "system", "content": ...}, {"role": "user", "content": ...}, {"role": "assistant", "content": ...}]
The FinCoach system message
system = {"role": "system", "content": "You are FinCoach, a personal finance assistant. Reply in one short sentence."}FinCoach is the video's example: a personal finance assistant. The system message goes first in every request.
Two separate calls to FinCoach
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
system = {"role": "system", "content": "You are FinCoach, a personal finance assistant. Reply in one short sentence."}
first = client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[system, {"role": "user", "content": "My monthly take-home is ₹1,20,000."}],
)
print(first.choices[0].message.content)
second = client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[system, {"role": "user", "content": "What is my monthly take-home?"}],
)
print(second.choices[0].message.content)Great—let’s start budgeting your ₹1,20,000 take‑home income. I need your gross salary and deduction details to calculate your monthly take‑home.
- The first reply picks up the ₹1,20,000 from the message it was sent.
- The second call asks for the take-home and gets a request for salary and deduction details instead. Its list had only the system message and the question, so the ₹1,20,000 was never in front of the model.
Sending the history with the question
This continues the same file. The history is the earlier user message and FinCoach's reply, then the new question:
history = [
system,
{"role": "user", "content": "My monthly take-home is ₹1,20,000."},
{"role": "assistant", "content": first.choices[0].message.content},
{"role": "user", "content": "What is my monthly take-home?"},
]
response = client.chat.completions.create(model="openai/gpt-oss-120b", messages=history)
print(response.choices[0].message.content)Your monthly take‑home is ₹1,20,000.
With the history in the list, the answer is right. The model did not remember anything; it read the earlier turn in the request.
How the messages become one text
A model reads one sequence of tokens, not a list. The provider writes the messages into one text with special marker tokens between them, a chat template, and that text is what gets tokenized. You can see its size:
import tiktoken
encoding = tiktoken.get_encoding("o200k_harmony")
text = "Write a one-sentence apology to a customer whose parcel is late."
print(len(encoding.encode(text)), "tokens in the text")
response = client.chat.completions.create(
model="openai/gpt-oss-120b", messages=[{"role": "user", "content": text}]
)
print(response.usage.prompt_tokens, "tokens the model read")14 tokens in the text 85 tokens the model read
The prompt is 14 tokens, and the model read 85. The other 71 are the markers and the header that gpt-oss's chat format puts around the messages. You never see them, and they are billed as input tokens.
The ask helper for the next lessons
from groq import Groq
client = Groq()
MODEL = "openai/gpt-oss-120b"
def ask(messages, **settings):
response = client.chat.completions.create(model=MODEL, messages=messages, **settings)
return response.choices[0].message.contentask sends a list of messages and returns the reply's text. **settings passes any extra argument, such as temperature=0, on to the request. The prompt lessons start their files with it.
print(ask([{"role": "user", "content": "In one sentence, what does a customer support ticket record?"}]))A customer support ticket records a customer's request, issue, or inquiry—including its details, timestamps, communication history, status, and resolution—within a single, traceable record.
One call vs a conversation
| One call | A conversation | |
|---|---|---|
| Messages sent | System and one user message | System, every earlier turn, the new message |
| Who keeps the history | Not needed | Your code |
| Tokens per call | Stay the same | Grow with every turn |
When to send the history
- Chat assistants, where a question refers to something said earlier.
- Not for independent tasks such as sorting tickets, where each ticket should be judged on its own.
- Trimmed when it grows long, as Context window shows.
Related
- Previous: Seeds
- Next: System prompts
- Reference: Groq text generation
- Reference: The harmony format of gpt-oss
- In the second call, add the first user message before the question, without FinCoach's reply. Does the answer change?
- Print
response.usage.prompt_tokensfor a list with a system message and compare it with the 85 above. - Call
askwithtemperature=0passed through**settings.
Every expert started right here.