create_deep_agent: your first deep agent
create_deep_agent is the function that builds a deep agent: it takes a model, your tools and a system prompt, adds the built-in file and subagent tools, and returns an agent you run with invoke.
Last updated: 29 Sep, 2026 · Deep Agents 0.7
groq:qwen/qwen3-32b, which Groq has since retired. The page runs the same code on groq:openai/gpt-oss-120b, with temperature=0 and max_retries=6 as in every lesson.The video's first deep agent
The video imports create_deep_agent from deepagents and passes three things: tools=[web_search], a system prompt, "Act as a researcher", and a model. The model comes from init_chat_model, which loads any provider's chat model from a provider:model string; the video uses Groq's qwen3-32b. A first try fails with an unexpected keyword models; the argument is model.
Start research.py with web_search, the Tavily tool from Tools: a travel search the agent can call. It needs tavily-python and a TAVILY_API_KEY, both set up there. The agent below also needs GROQ_API_KEY from Installation and setup.
import os
from tavily import TavilyClient
from typing import Literal
tavily_client = TavilyClient(api_key=os.getenv("TAVILY_API_KEY"))
def web_search(query: str, max_results: int = 5,
topic: Literal["general", "sports", "news", "finance"] = "general"):
"""Run a web search"""
return tavily_client.search(query, max_results=min(max_results, 5), topic=topic)Then the video's agent, on the course model:
from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)
deepagent = create_deep_agent(
model=model,
tools=[web_search],
system_prompt="Act as a researcher",
)write_todos tool unless you add TodoListMiddleware, which write_todos: planning with TodoListMiddleware does.It runs the agent with invoke and a dictionary whose messages key holds the conversation, one message with "role": "user" and the question "What is deepagent?". The run takes a while because the agent searches the web. result["messages"][-1].content is the answer; in the saved notebook it describes DeepAgent, a 2025 research paper found through the search.
The same agent runs below on Groq. Add the call to the end of research.py. The loop prints each step shortened, then the answer in full, as the video prints it, then the files in the agent's state.
result = deepagent.invoke({"messages": [{"role": "user", "content": "What is deepagent?"}]})
for message in result["messages"][:-1]: # the steps, shortened
print(f"{message.type:<5}", message.text[:120] or [(c["name"], c["args"]) for c in message.tool_calls])
print(result["messages"][-1].content) # the answer, as the video prints it
print("files:", list(result["files"]))human What is deepagent?
ai [('web_search', {'query': 'DeepAgent framework deep reinforcement learning', 'topic': 'general', 'topn': 10})]
tool {"query": "DeepAgent framework deep reinforcement learning", "follow_up_questions": null, "answer": null, "images": [],
**DeepAgent** is a research‑grade, end‑to‑end autonomous‑agent framework that equips a large language model (LLM) with the ability to **reason, discover and invoke tools, and learn from its own actions**.
Below is a concise, self‑contained description drawn from the authors’ papers, the project’s GitHub page, and related write‑ups.
---
## 1. Core Idea
Traditional “tool‑use” agents (e.g., ReAct) follow a fixed **Reason → Act → Observe** loop and rely on a pre‑specified set of tools. DeepAgent removes that rigidity:
| Traditional agents | DeepAgent |
|--------------------|-----------|
| **Fixed workflow** (reason → call a tool → observe) | **Unified reasoning process** that can interleave thinking, tool discovery, and execution arbitrarily |
| **Static tool list** (hard‑coded) | **Dynamic tool discovery** – the model can search a large toolbox and decide on‑the‑fly which tool(s) to use |
| **Limited memory** (few recent turns) | **Memory folding** – a brain‑inspired schema that compresses episodic, working‑, and tool‑memories, keeping long‑horizon context while staying within token limits |
| **No learning of tool use** | **ToolPO reinforcement learning** – an RL regime that gives credit to the exact tokens that trigger a successful tool call, allowing the agent to improve its tool‑use policy over time |
In short, DeepAgent is a **general‑reasoning agent with scalable toolsets** that can handle complex, long‑horizon tasks more robustly than earlier approaches.
---
## 2. Architectural Pillars
### 2.1. Large‑Reasoning Model (the “brain”)
* A standard LLM (e.g., LLaMA‑2, GPT‑4) serves as the core reasoning engine.
* The model receives a **structured JSON‑like context** that contains:
* **Episodic memory** – a compressed history of what has happened.
* **Working memory** – the current sub‑goal and any intermediate results.
* **Tool memory** – a list of available tools, their signatures, and recent usage statistics.
### 2.2. Dynamic Tool Discovery & Dense Retrieval
* Instead of a static toolbox, DeepAgent maintains a **large repository of “tool descriptors”** (APIs, code snippets, CLI commands, etc.).
* When the model decides a tool might be useful, it performs a **dense vector retrieval** (or a simple keyword match) to fetch the most relevant tool description, which is then inserted into the reasoning context.
### 2.3. Memory Folding (Context Management)
* Token limits are a major bottleneck for LLMs. DeepAgent introduces a **folding algorithm** that:
1. **Identifies redundant or low‑utility portions** of the conversation.
2. **Compresses them into a concise summary** (using the LLM itself).
3. **Re‑injects the summary** into the context, preserving essential dependencies while freeing up space for new reasoning steps.
* This mimics how the brain “takes a breath” and reorganizes information after a failed attempt.
### 2.4. ToolPO – Reinforcement Learning for Tool Use
* **ToolPO** (Tool‑Policy Optimization) treats each token that triggers a tool call as an **action**.
* The agent receives a **reward** when the tool call leads to a successful sub‑task (e.g., correct API response, completed search).
* By applying **policy‑gradient methods** (e.g., PPO) directly on the token‑level logits, the model learns to **attribute credit** precisely to the tool‑invocation tokens, improving both *when* and *which* tools to call.
---
## 3. What Problems Does DeepAgent Solve?
| Problem | How DeepAgent Addresses It |
|---------|----------------------------|
| **Long‑horizon planning** (many steps, large context) | Memory folding keeps the relevant history while staying within token limits. |
| **Tool‑selection brittleness** (hard‑coded tool lists) | Dynamic discovery lets the agent pick from a massive, extensible toolbox. |
| **Error accumulation** (once a tool call fails, the agent drifts) | The unified reasoning loop and memory folding allow the agent to “re‑think” and backtrack. |
| **Lack of learning in tool use** | ToolPO RL gives the model a way to improve its tool‑calling policy from experience. |
---
## 4. Current Status (as of 2024‑09)
| Artifact | Status |
|----------|--------|
| **Paper** | “DeepAgent: A General Reasoning Agent with Scalable Toolsets” (arXiv 2510.21618) – peer‑reviewed at WWW 2026 (oral). |
| **Code** | Open‑source repository: <https://github.com/RUC-NLPIR/DeepAgent>. Includes training scripts for ToolPO, a toolbox of >200 APIs, and a memory‑folding library. |
| **Benchmarks** | Evaluated on eight diverse tool‑use suites: ToolBench, API‑Bank, TMDB, Spotify, ToolHop, plus downstream tasks such as ALFWorld, WebShop, GAIA, and HLE. DeepAgent consistently outperforms ReAct, Voyager, and other baselines. |
| **Demo** | Interactive demo on the project site (EmergentMind) shows the agent solving a multi‑step travel‑planning problem by searching maps, calling a flight‑API, and booking a hotel—all without any hard‑coded plan. |
---
## 5. Quick “Elevator Pitch”
> **DeepAgent** is a next‑generation autonomous AI agent that **thinks, discovers tools, and learns** in a single, coherent process. By folding memory, dynamically retrieving the right tool, and training with reinforcement learning (ToolPO), it can tackle long, complex tasks that would overwhelm conventional LLM‑plus‑tool pipelines.
---
### TL;DR
- **What:** A general‑reasoning autonomous agent framework.
- **Key innovations:** Dynamic tool discovery, memory folding, and ToolPO RL for tool use.
- **Why it matters:** Enables LLMs to handle long‑horizon, multi‑tool problems more reliably and to improve over time.
If you need deeper technical details (e.g., the exact folding algorithm, loss functions for ToolPO, or how to plug in a new toolbox), let me know and I can walk you through the relevant sections of the paper or the source code.
files: []What the research agent did
- One search: the model called
web_searchwith a query it wrote itself, "DeepAgent framework deep reinforcement learning". Its extratopnargument is not one the function takes, and the call ran with the defaults. - The tool message is Tavily's JSON: the query, then a list of results, each with a URL, a title and an extract.
- The answer describes DeepAgent, the research paper arXiv 2510.21618, which is also what the video's run found. It is long, and not every detail in it can be traced to the results; the prompt "Act as a researcher" does not ask the model to stay with its sources, as the course's own prompts do.
filesis empty: the results were small enough to stay in the conversation, so nothing was offloaded to a file.- Search results change from day to day, so your run will find other pages and word its answer differently.
The rest of this lesson builds the same shape on the travel catalog, whose answers can be checked line by line.
The create_deep_agent call
agent = create_deep_agent(model=model, tools=[my_tool], system_prompt="...")
result = agent.invoke({"messages": [{"role": "user", "content": "..."}]})
result["messages"][-1].text # the final answerThe model
Start trip.py with search_travel, the catalog tool from Tools: a travel search the agent can call. Everything below goes in the same file, under it. Then create the model. temperature=0 keeps the replies steady from run to run.
from langchain.tools import tool
CATALOG = {
"paris": {
"flight": ["Return flight Delhi to Paris: 42,000 rupees"],
"hotel": ["Seine Budget Inn, Latin Quarter: 5,200 rupees a night",
"Hotel Lumiere, Montmartre: 7,500 rupees a night",
"Le Grand Opera Hotel: 16,000 rupees a night"],
"sight": ["Eiffel Tower summit: 3,100 rupees", "Louvre Museum: 2,000 rupees",
"Seine river cruise: 1,500 rupees", "Versailles day trip: 2,600 rupees",
"Montmartre walking tour: free"],
"food": ["Cafe breakfast and bistro dinner: 3,000 rupees a day"],
},
}
@tool
def search_travel(city: str, kind: str) -> str:
"""Search the travel catalog. kind is "flight", "hotel", "sight" or "food". Prices are in rupees."""
entries = CATALOG.get(city.lower(), {}).get(kind)
return "\n".join(entries) if entries else f"The catalog has no {kind} entries for {city}."from deepagents import create_deep_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0, max_retries=6)The agent
The system prompt makes this agent a travel planner and tells it to use only what the catalog returns, so the answer cannot drift into prices the model remembers from training.
agent = create_deep_agent(
model=model,
tools=[search_travel],
system_prompt="You are a travel planner. Look up every price with search_travel and use only "
"what it returns. Answer in two short sentences.",
)Running the agent on a hotel question
Add the call and a loop that prints every message of the run: its type, then its text, or the tool it asked for.
result = agent.invoke({"messages": [{"role": "user", "content": "What is the cheapest hotel in Paris?"}]})
for message in result["messages"]:
print(f"{message.type:<5}", message.text or [(c["name"], c["args"]) for c in message.tool_calls])human What is the cheapest hotel in Paris?
ai [('search_travel', {'city': 'Paris', 'kind': 'hotel'})]
tool Seine Budget Inn, Latin Quarter: 5,200 rupees a night
Hotel Lumiere, Montmartre: 7,500 rupees a night
Le Grand Opera Hotel: 16,000 rupees a night
ai The cheapest hotel in Paris is the Seine Budget Inn in the Latin Quarter, costing 5,200 rupees per night. It’s the most affordable option among the listed hotels.Reading the four messages
- human: the question you sent.
- ai with a tool call: the model did not answer yet. It asked for
search_travelwithcityParis andkindhotel, arguments it chose from the docstring. - tool: the agent ran the function and added its result to the conversation.
- ai with text: the model read the result and answered. A reply with no tool call ends the run.
- The loop that ran is the same one
create_agent, LangChain's plain agent builder compared in the next lesson, runs. What makes the agent "deep" is whatcreate_deep_agentadds around it, the subject of the next lesson.
create_deep_agent vs calling the model yourself
| model.invoke | create_deep_agent | |
|---|---|---|
| Tools | The model can only ask for them | The agent runs them and loops |
| Input | A string or a list of messages | A dictionary with a messages key |
| Output | One AI message | The whole conversation, plus files and other state |
When to reach for create_deep_agent
- A job that needs more than one tool call, such as comparing hotels and sights before answering.
- Work that should leave files behind: notes, a plan, a report.
- Tasks you may later split between subagents or put behind approval.
{"messages": [...]}. Passing the question string on its own fails, because the agent's input is its whole state, not one message.Related
- Previous: Tools: a travel search the agent can call
- Next: Built-in tools: files and the task tool
- Reference: Deep Agents quickstart
- Ask "Which Paris sight is free?" and read the tool call's
kindargument. - Ask about Tokyo and check that the answer says the catalog has nothing.
- Remove "Answer in two short sentences" from the prompt and compare the length of the reply.
Slow is fine. Stopping is the only problem.