LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Chat messages and templates

A chat message is a dictionary with a role, system, user or assistant, and its content; a chat model reads the whole list of messages in every request and remembers nothing between requests.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

Every example so far sent a list with one user message. A conversation is the same list with more messages in it, and it is your code, not the model, that keeps the list.

LLMs are stateless · from the Complete AI Security Course In 8 Hours · 181:15 to 183:51

The memory section of the video starts from one fact: LLMs are stateless. Across API requests the model has no idea what it did a second ago, and the video stresses that this is how they are built, not a bug. The fix is a list: every user message and every reply is appended to it, and the whole list is sent again on the next call. Each message has a strict role: system for the instructions, user for the person, assistant for the model.

The video's notebook calls OpenAI's API. The requests below send the same kind of list to Groq.

Syntax: messages=[{"role": "system", "content": ...}, {"role": "user", "content": ...}, {"role": "assistant", "content": ...}]

The FinCoach system message

python
system = {"role": "system", "content": "You are FinCoach, a personal finance assistant. Reply in one short sentence."}

FinCoach is the video's example: a personal finance assistant. The system message goes first in every request.

Two separate calls to FinCoach

ExampleAPI keyFrom the video, run on Groq
from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment
system = {"role": "system", "content": "You are FinCoach, a personal finance assistant. Reply in one short sentence."}

first = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[system, {"role": "user", "content": "My monthly take-home is ₹1,20,000."}],
)
print(first.choices[0].message.content)

second = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[system, {"role": "user", "content": "What is my monthly take-home?"}],
)
print(second.choices[0].message.content)
  • The first reply picks up the ₹1,20,000 from the message it was sent.
  • The second call asks for the take-home and gets a request for salary and deduction details instead. Its list had only the system message and the question, so the ₹1,20,000 was never in front of the model.

Sending the history with the question

This continues the same file. The history is the earlier user message and FinCoach's reply, then the new question:

ExampleAPI key
history = [
    system,
    {"role": "user", "content": "My monthly take-home is ₹1,20,000."},
    {"role": "assistant", "content": first.choices[0].message.content},
    {"role": "user", "content": "What is my monthly take-home?"},
]
response = client.chat.completions.create(model="openai/gpt-oss-120b", messages=history)
print(response.choices[0].message.content)

With the history in the list, the answer is right. The model did not remember anything; it read the earlier turn in the request.

How the messages become one text

A model reads one sequence of tokens, not a list. The provider writes the messages into one text with special marker tokens between them, a chat template, and that text is what gets tokenized. You can see its size:

ExampleAPI key
import tiktoken

encoding = tiktoken.get_encoding("o200k_harmony")
text = "Write a one-sentence apology to a customer whose parcel is late."
print(len(encoding.encode(text)), "tokens in the text")

response = client.chat.completions.create(
    model="openai/gpt-oss-120b", messages=[{"role": "user", "content": text}]
)
print(response.usage.prompt_tokens, "tokens the model read")

The prompt is 14 tokens, and the model read 85. The other 71 are the markers and the header that gpt-oss's chat format puts around the messages. You never see them, and they are billed as input tokens.

The ask helper for the next lessons

python
from groq import Groq

client = Groq()
MODEL = "openai/gpt-oss-120b"


def ask(messages, **settings):
    response = client.chat.completions.create(model=MODEL, messages=messages, **settings)
    return response.choices[0].message.content

ask sends a list of messages and returns the reply's text. **settings passes any extra argument, such as temperature=0, on to the request. The prompt lessons start their files with it.

ExampleAPI key
print(ask([{"role": "user", "content": "In one sentence, what does a customer support ticket record?"}]))

One call vs a conversation

One callA conversation
Messages sentSystem and one user messageSystem, every earlier turn, the new message
Who keeps the historyNot neededYour code
Tokens per callStay the sameGrow with every turn

When to send the history

  • Chat assistants, where a question refers to something said earlier.
  • Not for independent tasks such as sorting tickets, where each ticket should be judged on its own.
  • Trimmed when it grows long, as Context window shows.
Watch out. Every turn you keep is sent and billed again on every later call. The list only grows, and Context window shows what happens when it gets too big.
Try it yourself
  • In the second call, add the first user message before the question, without FinCoach's reply. Does the answer change?
  • Print response.usage.prompt_tokens for a list with a system message and compare it with the 85 above.
  • Call ask with temperature=0 passed through **settings.
PreviousSeeds

Every expert started right here.