LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Context window

A context window is the largest number of tokens a model can handle in one request, counting the system prompt, the examples, the whole conversation and the answer it writes.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

Chat messages and templates kept a conversation by sending the whole history every time. That list grows with every turn, and the window is where it stops.

Token cost grows every turn · from the Complete AI Security Course In 8 Hours · 183:51 to 186:33

The video continues from the growing list: every turn adds a user message and a reply, so the list keeps increasing until the context window blows up. In the FinCoach example on screen, the first turn sends about 80 tokens, the third about 200 and the fifth about 400, and a long conversation reaches thousands. The notebook names GPT-4o's window, 128,000 tokens, and states the rule: hit the window and the API call fails.

Syntax: the window is a property of the model, read from client.models.list(); what you control is how many tokens you send.

Reading the window from the model list

ExampleAPI key
from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment

for model in client.models.list().data:
    if model.id == "openai/gpt-oss-120b":
        print("context window:", model.context_window, "tokens")
        print("longest answer:", model.max_completion_tokens, "tokens")
  • 131,072 tokens is everything gpt-oss-120b can take in one request on Groq.
  • 65,536 tokens is the most it may write in one answer, and those tokens come out of the same window.

Watching FinCoach's history grow

The FinCoach system message from Chat messages and templates starts the file:

python
from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment
system = {"role": "system", "content": "You are FinCoach, a personal finance assistant. Reply in one short sentence."}
ExampleAPI keyFrom the video, run on Groq
turns = [
    "My monthly take-home is ₹1,20,000.",
    "My expenses are ₹60,000 a month.",
    "What is the most important thing I should do this month?",
]
history = [system]
for turn in turns:
    history.append({"role": "user", "content": turn})
    response = client.chat.completions.create(model="openai/gpt-oss-120b", messages=history)
    history.append({"role": "assistant", "content": response.choices[0].message.content})
    print(len(history), "messages,", response.usage.prompt_tokens, "tokens sent")
  • 102, 164, 218 tokens: each turn sends everything before it again, plus the new message.
  • The model answered each turn from the list it was sent, as in Chat messages and templates.
  • At this rate a long chat reaches the window. A 131,072-token window is large, but documents pasted into a chat, or hundreds of turns, get there.

Keeping the newest turns

A long conversation

A support chat with 40 earlier turns, starting with structured from System prompts:

python
conversation = [{"role": "system", "content": structured}]
for turn in range(40):
    conversation.append({"role": "user", "content": f"Ticket {turn}: my parcel has not arrived yet"})
    conversation.append({"role": "assistant", "content": '{"category": "shipping", "priority": 3}'})

Counting tokens with tiktoken

python
import tiktoken

encoding = tiktoken.get_encoding("o200k_harmony")


def count_tokens(messages):
    return sum(len(encoding.encode(message["content"])) for message in messages)

count_tokens adds up the tokens in each message's content. It leaves out the chat template's markers, so a real request is somewhat larger.

A trim function

python
def trim(messages, budget):
    system, rest = messages[:1], messages[1:]
    while rest and count_tokens(system + rest) > budget:
        rest = rest[2:]  # drop the oldest question and its answer together
    return system + rest

trim keeps the system message and drops the oldest user message and its reply, two at a time, until the rest fits the budget.

Trimming 81 messages to a 500-token budget

Example
print(len(conversation), "messages,", count_tokens(conversation), "tokens")

short = trim(conversation, budget=500)
print(len(short), "messages,", count_tokens(short), "tokens")
print("oldest kept:", short[1]["content"])
  • 81 messages, 956 tokens before: the system message and 40 turns.
  • 39 messages, 494 tokens after: the system message and the newest 19 turns fit under 500.
  • The oldest turn kept is Ticket 21; tickets 0 to 20 are gone, and so is anything said in them.

Trimming vs summarising vs storing facts

Trim the oldest turnsSummarise old turnsStore facts separately
KeepsThe newest turns word for wordA short summary of old turnsChosen facts, fetched when relevant
LosesEverything olderDetails the summary left outAnything not stored
Extra model callsNoneOne per summarySome, to extract and search

When the window matters

  • Long chats, where the history grows every turn.
  • Pasting documents into a prompt: a long PDF can be tens of thousands of tokens.
  • Answers you expect to be long, since the answer shares the window.
Watch out. Leave room for the answer and the template. A request whose input fills the window leaves nothing for the reply, and one over the window fails with an error, as the video says.
Try it yourself
  • Trim to a budget of 200 and print what is left.
  • Change rest[2:] to rest[1:], trim to a budget of 490 and print short[1]. Which question does that first message answer?
  • Add a fourth FinCoach turn and see how many tokens it sends.

Every expert started right here.