Context window
A context window is the largest number of tokens a model can handle in one request, counting the system prompt, the examples, the whole conversation and the answer it writes.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Chat messages and templates kept a conversation by sending the whole history every time. That list grows with every turn, and the window is where it stops.
The video continues from the growing list: every turn adds a user message and a reply, so the list keeps increasing until the context window blows up. In the FinCoach example on screen, the first turn sends about 80 tokens, the third about 200 and the fifth about 400, and a long conversation reaches thousands. The notebook names GPT-4o's window, 128,000 tokens, and states the rule: hit the window and the API call fails.
Syntax: the window is a property of the model, read from client.models.list(); what you control is how many tokens you send.
Reading the window from the model list
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
for model in client.models.list().data:
if model.id == "openai/gpt-oss-120b":
print("context window:", model.context_window, "tokens")
print("longest answer:", model.max_completion_tokens, "tokens")context window: 131072 tokens longest answer: 65536 tokens
- 131,072 tokens is everything gpt-oss-120b can take in one request on Groq.
- 65,536 tokens is the most it may write in one answer, and those tokens come out of the same window.
Watching FinCoach's history grow
The FinCoach system message from Chat messages and templates starts the file:
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
system = {"role": "system", "content": "You are FinCoach, a personal finance assistant. Reply in one short sentence."}turns = [
"My monthly take-home is ₹1,20,000.",
"My expenses are ₹60,000 a month.",
"What is the most important thing I should do this month?",
]
history = [system]
for turn in turns:
history.append({"role": "user", "content": turn})
response = client.chat.completions.create(model="openai/gpt-oss-120b", messages=history)
history.append({"role": "assistant", "content": response.choices[0].message.content})
print(len(history), "messages,", response.usage.prompt_tokens, "tokens sent")3 messages, 102 tokens sent 5 messages, 164 tokens sent 7 messages, 218 tokens sent
- 102, 164, 218 tokens: each turn sends everything before it again, plus the new message.
- The model answered each turn from the list it was sent, as in Chat messages and templates.
- At this rate a long chat reaches the window. A 131,072-token window is large, but documents pasted into a chat, or hundreds of turns, get there.
Keeping the newest turns
A long conversation
A support chat with 40 earlier turns, starting with structured from System prompts:
conversation = [{"role": "system", "content": structured}]
for turn in range(40):
conversation.append({"role": "user", "content": f"Ticket {turn}: my parcel has not arrived yet"})
conversation.append({"role": "assistant", "content": '{"category": "shipping", "priority": 3}'})Counting tokens with tiktoken
import tiktoken
encoding = tiktoken.get_encoding("o200k_harmony")
def count_tokens(messages):
return sum(len(encoding.encode(message["content"])) for message in messages)count_tokens adds up the tokens in each message's content. It leaves out the chat template's markers, so a real request is somewhat larger.
A trim function
def trim(messages, budget):
system, rest = messages[:1], messages[1:]
while rest and count_tokens(system + rest) > budget:
rest = rest[2:] # drop the oldest question and its answer together
return system + resttrim keeps the system message and drops the oldest user message and its reply, two at a time, until the rest fits the budget.
Trimming 81 messages to a 500-token budget
print(len(conversation), "messages,", count_tokens(conversation), "tokens")
short = trim(conversation, budget=500)
print(len(short), "messages,", count_tokens(short), "tokens")
print("oldest kept:", short[1]["content"])81 messages, 956 tokens 39 messages, 494 tokens oldest kept: Ticket 21: my parcel has not arrived yet
- 81 messages, 956 tokens before: the system message and 40 turns.
- 39 messages, 494 tokens after: the system message and the newest 19 turns fit under 500.
- The oldest turn kept is Ticket 21; tickets 0 to 20 are gone, and so is anything said in them.
Trimming vs summarising vs storing facts
| Trim the oldest turns | Summarise old turns | Store facts separately | |
|---|---|---|---|
| Keeps | The newest turns word for word | A short summary of old turns | Chosen facts, fetched when relevant |
| Loses | Everything older | Details the summary left out | Anything not stored |
| Extra model calls | None | One per summary | Some, to extract and search |
When the window matters
- Long chats, where the history grows every turn.
- Pasting documents into a prompt: a long PDF can be tens of thousands of tokens.
- Answers you expect to be long, since the answer shares the window.
Related
- Previous: Evaluating prompts
- Next: Cost
- Reference: Groq supported models
- Trim to a budget of
200and print what is left. - Change
rest[2:]torest[1:], trim to a budget of490and printshort[1]. Which question does that first message answer? - Add a fourth FinCoach turn and see how many tokens it sends.
Every expert started right here.