Deep AgentsDeep Agents 0.7 · Python 3.11+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

Context offloading: large results saved to files

Context offloading is how a deep agent keeps a huge tool result out of the conversation: the result is saved to a file under /large_tool_results/, and the model gets the path and a short preview instead.

Last updated: 29 Sep, 2026 · Deep Agents 0.7

Files the agent created on its own · from the Complete Deep Agents Course With LangChain · 40:40 to 43:10
The video links this file to summarization and says it may be written to disk. It is offloading, and with the default backend the file lives in the agent's state, not on disk.

The files the video did not ask for

After the first run in the video, result["files"] holds a file the video never asked for, under /large_tool_results/, with the full Tavily search response inside: query, results, URLs and titles. The video explains it as the agent preserving context: when a result is very large, the agent stores it in a file instead of carrying it in the conversation.

When web_search ran on Groq in create_deep_agent: your first deep agent, result["files"] stayed empty: that version leaves out include_raw_content, so its results were small enough to stay in the conversation. With include_raw_content on, as the video's model chose, a test run for this course got "Tool result too large" and the result saved under /large_tool_results/, as in the video.

How offloading and summarization work

Two built-in mechanisms keep the context small. Offloading: when a tool result is larger than a limit (a large default, set by tool_token_limit_before_evict), FilesystemMiddleware writes it to a file and replaces it with the path and a preview of the first and last lines. The agent can then read parts of it or search it with grep. Summarization: when the whole conversation nears the model's context window, older messages are replaced by a summary, and the original text is kept in a file.

Lowering the limit to see it happen

A result above the default limit would not fit Groq's free tier, so this example sets the limit to 500 tokens. To change a built-in middleware, pass your own instance with the same name; it replaces the default one. FilesystemMiddleware needs the backend passed to it as well.

A tool with a long result

hotel_reviews returns 80 reviews, one per line. Every fourth one mentions a noisy street, so the right answer to "how many mention noise" is 20. Save it as reviews.py:

python
from langchain.tools import tool

@tool
def hotel_reviews(hotel: str) -> str:
    """Return every guest review for a hotel, one per line."""
    lines = []
    for n in range(1, 81):
        note = "noisy street at night" if n % 4 == 0 else "clean room, friendly staff"
        lines.append(f"Review {n} of {hotel}: {note}, stayed {n % 5 + 1} nights.")
    return "\n".join(lines)

The agent with a 500-token limit

python
from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)
python
from deepagents.backends import StateBackend
from deepagents.middleware import FilesystemMiddleware

backend = StateBackend()
agent = create_deep_agent(
    model=model,
    tools=[hotel_reviews],
    backend=backend,
    middleware=[FilesystemMiddleware(backend=backend, tool_token_limit_before_evict=500)],
    system_prompt="You answer questions about hotel reviews. When a result is saved to a file, "
                  "count matching lines with grep. Reply in one sentence.",
)

Counting noisy reviews in an offloaded result

ExampleAPI keyreviews.py, continued
question = "How many reviews of the Seine Budget Inn mention noise?"
result = agent.invoke({"messages": [{"role": "user", "content": question}]})

for message in result["messages"]:
    text = message.text[:160].replace("\n", " | ")
    print(f"{message.type:<5}", text or [(c["name"], c["args"]) for c in message.tool_calls])
print("files:", list(result["files"]))

What happened to the 80 reviews

  • The model never saw all 80 lines: the tool message says the result was too large, gives its path under /large_tool_results/ and shows a short preview of the first and last lines.
  • The agent searched the file with grep in count mode for "noisy", as the prompt suggested, and got 20.
  • The answer is right: 20 of 80 reviews mention noise, and the conversation stayed small.

Offloading vs summarization

OffloadingSummarization
Triggers whenOne tool result is over the limitThe conversation nears the context window
What is savedThe full result, in a fileThe old messages, in a file
What the model seesA path and a previewA summary of the old messages

Where offloading matters

  • Web search, which can return whole pages of text.
  • Database queries and logs with thousands of rows.
  • Long documents the agent should search, not read in full.
Watch out. Overriding FilesystemMiddleware replaces the default one entirely. Forget backend=backend in your instance and the file tools no longer use the backend you gave create_deep_agent.
Try it yourself
  • Raise the limit to 5,000 and see whether the result is still offloaded.
  • Ask "Which review numbers mention noise?" and read which file tool the agent uses.
  • Print result["files"]'s first path and the first 200 characters of its content.

Every expert started right here.