Context offloading: large results saved to files
Context offloading is how a deep agent keeps a huge tool result out of the conversation: the result is saved to a file under /large_tool_results/, and the model gets the path and a short preview instead.
Last updated: 29 Sep, 2026 · Deep Agents 0.7
The files the video did not ask for
After the first run in the video, result["files"] holds a file the video never asked for, under /large_tool_results/, with the full Tavily search response inside: query, results, URLs and titles. The video explains it as the agent preserving context: when a result is very large, the agent stores it in a file instead of carrying it in the conversation.
When web_search ran on Groq in create_deep_agent: your first deep agent, result["files"] stayed empty: that version leaves out include_raw_content, so its results were small enough to stay in the conversation. With include_raw_content on, as the video's model chose, a test run for this course got "Tool result too large" and the result saved under /large_tool_results/, as in the video.
How offloading and summarization work
Two built-in mechanisms keep the context small. Offloading: when a tool result is larger than a limit (a large default, set by tool_token_limit_before_evict), FilesystemMiddleware writes it to a file and replaces it with the path and a preview of the first and last lines. The agent can then read parts of it or search it with grep. Summarization: when the whole conversation nears the model's context window, older messages are replaced by a summary, and the original text is kept in a file.
Lowering the limit to see it happen
A result above the default limit would not fit Groq's free tier, so this example sets the limit to 500 tokens. To change a built-in middleware, pass your own instance with the same name; it replaces the default one. FilesystemMiddleware needs the backend passed to it as well.
A tool with a long result
hotel_reviews returns 80 reviews, one per line. Every fourth one mentions a noisy street, so the right answer to "how many mention noise" is 20. Save it as reviews.py:
from langchain.tools import tool
@tool
def hotel_reviews(hotel: str) -> str:
"""Return every guest review for a hotel, one per line."""
lines = []
for n in range(1, 81):
note = "noisy street at night" if n % 4 == 0 else "clean room, friendly staff"
lines.append(f"Review {n} of {hotel}: {note}, stayed {n % 5 + 1} nights.")
return "\n".join(lines)The agent with a 500-token limit
from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)from deepagents.backends import StateBackend
from deepagents.middleware import FilesystemMiddleware
backend = StateBackend()
agent = create_deep_agent(
model=model,
tools=[hotel_reviews],
backend=backend,
middleware=[FilesystemMiddleware(backend=backend, tool_token_limit_before_evict=500)],
system_prompt="You answer questions about hotel reviews. When a result is saved to a file, "
"count matching lines with grep. Reply in one sentence.",
)Counting noisy reviews in an offloaded result
question = "How many reviews of the Seine Budget Inn mention noise?"
result = agent.invoke({"messages": [{"role": "user", "content": question}]})
for message in result["messages"]:
text = message.text[:160].replace("\n", " | ")
print(f"{message.type:<5}", text or [(c["name"], c["args"]) for c in message.tool_calls])
print("files:", list(result["files"]))human How many reviews of the Seine Budget Inn mention noise?
ai [('hotel_reviews', {'hotel': 'Seine Budget Inn'})]
tool Tool result too large, the result of this tool call fc_431e7d45-b08c-4ab7-8926-ff48f... was saved in the filesystem at this path: /large_tool_results/fc_431e7d4
ai [('grep', {'output_mode': 'count', 'path': '/large_tool_results/fc_431e7d45-b08c-4ab7-8926-ff48fcc28d5e', 'pattern': 'noisy'})]
tool /large_tool_results/fc_431e7d45-b08c-4ab7-8926-ff48fcc28d5e: 20
ai 20 reviews mention noise.
files: ['/large_tool_results/fc_431e7d45-b08c-4ab7-8926-ff48fcc28d5e']What happened to the 80 reviews
- The model never saw all 80 lines: the tool message says the result was too large, gives its path under
/large_tool_results/and shows a short preview of the first and last lines. - The agent searched the file with
grepin count mode for "noisy", as the prompt suggested, and got 20. - The answer is right: 20 of 80 reviews mention noise, and the conversation stayed small.
Offloading vs summarization
| Offloading | Summarization | |
|---|---|---|
| Triggers when | One tool result is over the limit | The conversation nears the context window |
| What is saved | The full result, in a file | The old messages, in a file |
| What the model sees | A path and a preview | A summary of the old messages |
Where offloading matters
- Web search, which can return whole pages of text.
- Database queries and logs with thousands of rows.
- Long documents the agent should search, not read in full.
FilesystemMiddleware replaces the default one entirely. Forget backend=backend in your instance and the file tools no longer use the backend you gave create_deep_agent.Related
- Previous: Checkpointer: a thread that keeps its files
- Next: FilesystemBackend: files on your disk
- Reference: Context engineering: offloading
- Raise the limit to 5,000 and see whether the result is still offloaded.
- Ask "Which review numbers mention noise?" and read which file tool the agent uses.
- Print
result["files"]'s first path and the first 200 characters of its content.
Every expert started right here.