1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you manage the context window in a long-running agent?
30-second answerSay your answer out loud first, then reveal.
Why it matters: even with 200K–1M token windows, (1) cost and latency grow with every token re-sent on each step, and (2) model attention degrades as context fills. Important instructions get lost in the middle, and stale or contradictory information confuses the model. This is often called context rot. "Context engineering" is the discipline of curating what goes in.
Techniques (name several)
- Compaction / summarisation: when history passes a threshold, summarise older turns into key decisions, open questions and facts, and keep recent turns verbatim.
- Tool-output hygiene: return concise results; truncate with a note ("showing 20 of 340 rows; refine your query").
- Offloading to files / scratchpad: write large results to a file or store and keep a pointer. The agent re-reads only what it needs.
- Just-in-time retrieval: keep identifiers (file paths, doc IDs) and load the content only when needed, instead of pre-loading everything.
- Sub-agents: delegate "read these 30 files and report" to a sub-agent and get back a 500-token summary (Q34).
- Clearing old tool results: once a tool result has been used, older raw outputs can be dropped.
- Structured state: keep a compact state object (plan, findings, decisions) instead of relying on a long transcript.
Trade-off: aggressive compaction can drop a detail you later need. Mitigate by keeping originals retrievable (files, logs) and testing compaction prompts on real long traces.
Follow-ups to expect
- How would you know compaction caused a failure? Trace analysis: the agent asks for or re-fetches information it already had.
Related
You understood something today that you didn't yesterday.