Call limits with ModelCallLimitMiddleware
A call limit is a middleware setting that ends an agent run after a set number of model or tool calls, so a stuck agent costs a few calls instead of thousands.
Last updated: 27 Sep, 2026 · LangChain 1.4
Nothing in the agent loop counts by itself. As long as the model keeps asking for tools, it keeps going. This stand-in model asks for the same lookup every time.
ModelCallLimitMiddleware and ToolCallLimitMiddleware
from langchain.agents.middleware import ModelCallLimitMiddleware, ToolCallLimitMiddleware
# cap model calls in one run (or across a thread with thread_limit)
limit = ModelCallLimitMiddleware(run_limit=3)
# cap how often one named tool may be called
tool_limit = ToolCallLimitMiddleware(tool_name="lookup_order", run_limit=2)A model stuck in a loop
StuckModel extends ShopModel, the stand-in from earlier lessons, and always asks for the same tool, so it never stops on its own.
from langchain.messages import AIMessage
from shop_model import ShopModel
class StuckModel(ShopModel):
def decide(self, messages):
# always ask for the same lookup, so the loop never ends
call = {"name": "lookup_order", "args": {"order_id": "A17"}, "id": f"call_{len(messages)}"}
return AIMessage("", tool_calls=[call])The imports
Bring in the pieces for a plain agent, no limits yet.
from langchain.agents import create_agent
from langchain.agents.middleware import ModelCallLimitMiddleware, ToolCallLimitMiddleware
from stuck_model import StuckModel
from tools import lookup_orderWithout a limit
With no limit, only LangGraph's recursion_limit stops it. Set it low to see the run end with an error.
agent = create_agent(StuckModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]}, {"recursion_limit": 10})recursion_limit caps the number of steps, and here the run stopped with an error after ten. Left at the agent's default it goes on for 9,999 steps before this error. With a hosted model, that is thousands of paid calls for one question.
A limit on model calls
ModelCallLimitMiddleware ends the run cleanly once the model has been called a set number of times.
limit = ModelCallLimitMiddleware(run_limit=3)
agent = create_agent(StuckModel(), tools=[lookup_order], middleware=[limit])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
print(len(result["messages"]))
print(result["messages"][-1].text)Three model calls, then the run ended with an AI message saying why. run_limit counts calls in one invoke; thread_limit counts across a whole thread when there is a checkpointer. The default exit_behavior of "end" finishes with that message; "error" raises instead.
A limit on one tool
ToolCallLimitMiddleware caps one named tool rather than the model. Pair it with a model limit to be sure the run ends.
limits = [ToolCallLimitMiddleware(tool_name="lookup_order", run_limit=2),
ModelCallLimitMiddleware(run_limit=4)]
agent = create_agent(StuckModel(), tools=[lookup_order], middleware=limits)
for message in agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"]:
if message.type == "tool":
print(message.text)Two lookups ran. After that the tool call limit answered each request itself, telling the model to stop, and the model call limit ended the run. A tool limit's default behaviour is "continue": block the tool, keep the agent going.
How each limit stopped the run
- Nothing counts on its own. Without a limit the loop runs to
recursion_limit, 9,999 steps by default, then errors. - A model limit ends the run. After three calls the run finished with an AI message naming the limit it hit.
- A tool limit blocks one tool. Two lookups ran, then each further call was answered with a message telling the model to stop.
- They pair well. The tool limit stops the tool, the model limit ends the run.
ModelCallLimitMiddleware vs ToolCallLimitMiddleware
| ModelCallLimitMiddleware | ToolCallLimitMiddleware | |
|---|---|---|
| Counts | Model calls | Calls to one named tool |
| Default on limit | Ends the run (exit_behavior='end') | Blocks the tool, keeps going ('continue') |
| Scope | run_limit per invoke, thread_limit per thread | run_limit per invoke, thread_limit per thread |
| Use for | Capping the total cost of a run | Stopping one tool from being overused |
Where call limits fit
- Capping the cost of a single request so a stuck loop cannot run up thousands of calls.
- Stopping one expensive tool from being called over and over.
- Ending a run with a clear message rather than a raised error.
Related
- Previous: Tool errors, caught and retried
- Next: Retries and a fallback model
- Reference: LangChain agent middleware
- Set
exit_behavior="error"on the model call limit and read the exception. - Give the tool limit no
tool_name, so it counts every tool. - Run lesson 7's agent with a model call limit of 1 and ask about A17.
Slow is fine. Stopping is the only problem.