1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is the context window, and what happens when you exceed it?
30-second answerSay your answer out loud first, then reveal.
Key points
- Input + output share the budget. A 128K model with a 120K-token prompt has at most 8K tokens for the answer (and providers often cap output separately).
- Why there's a limit: training length, positional encoding limits, and attention memory/compute that grows with sequence length (quadratic for standard attention, plus KV cache memory growing linearly per token).
- Effective vs advertised context: models may accept 1M tokens but use information less reliably deep in long contexts, especially with many similar distractors ("lost in the middle").
Handling long inputs
- Retrieve only the relevant parts (RAG).
- Summarise / compact older conversation turns.
- Chunk and map-reduce: process parts separately, then combine.
- Choose a long-context model when the task truly needs holistic reading.
- Prompt caching for repeatedly used long prefixes.
Common mistakes
- Building chatbots that keep appending history until they crash or silently truncate the system prompt.
Related
This is what real progress feels like.