Call limits with ModelCallLimitMiddleware
A call limit is a middleware setting that ends an agent run after a set number of model or tool calls, so a stuck agent costs a few calls instead of thousands.
Last updated: 27 Sep, 2026 · LangChain 1.4
Nothing in the agent loop counts by itself. As long as the model keeps asking for tools, it keeps going. A real model stops once it has its answer, so a stuck loop cannot be produced on cue; this lesson uses a stand-in model that asks for the same lookup every time.
ModelCallLimitMiddleware and ToolCallLimitMiddleware
from langchain.agents.middleware import ModelCallLimitMiddleware, ToolCallLimitMiddleware
# cap model calls in one run (or across a thread with thread_limit)
limit = ModelCallLimitMiddleware(run_limit=3)
# cap how often one named tool may be called
tool_limit = ToolCallLimitMiddleware(tool_name="lookup_order", run_limit=2)A model stuck in a loop
The agent here runs on ShopModel, the stand-in chat model built in Several tool calls at once. Start the file with it.
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult
class ShopModel(BaseChatModel):
tools: list = []
@property
def _llm_type(self):
return "shop"
def bind_tools(self, tools, **kwargs):
return self.model_copy(update={"tools": tools}) # a copy holding the tools
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
message = self.decide(messages) # the reply comes from decide
return ChatResult(generations=[ChatGeneration(message=message)])
def decide(self, messages):
results = [] # the tool results at the end
for m in reversed(messages):
if not isinstance(m, ToolMessage):
break
results.insert(0, m.text)
if results: # results are back: answer with them
return AIMessage(" ".join(results))
text = messages[-1].text
orders = re.findall(r"\b[A-Z]\d+\b", text)
tool = "refund_order" if "refund" in text.lower() else "lookup_order"
if orders and tool in [t.name for t in self.tools]: # one call per order id
calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
for o in orders]
return AIMessage("", tool_calls=calls)
if orders: # that tool is not bound
return AIMessage(f"I have no way to look up {orders[0]} yet.")
return AIMessage("Hello. Which order is this about?")StuckModel extends ShopModel and overrides its decide method to always ask for the same tool, so it never stops on its own.
from langchain.messages import AIMessage
class StuckModel(ShopModel):
def decide(self, messages):
# always ask for the same lookup, so the loop never ends
call = {"name": "lookup_order", "args": {"order_id": "A17"}, "id": f"call_{len(messages)}"}
return AIMessage("", tool_calls=[call])The tool and the agent
This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Add it below the models.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Without a limit
With no limit, only LangGraph's recursion_limit stops it. Set it low to see the run end with an error.
from langchain.agents import create_agent
agent = create_agent(StuckModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]}, {"recursion_limit": 10})Traceback (most recent call last):
File "main.py", line 4, in <module>
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]}, {"recursion_limit": 10})
langgraph.errors.GraphRecursionError: Recursion limit of 10 reached without hitting a stop condition. You can increase the limit by setting the `recursion_limit` config key.
For troubleshooting, visit: https://docs.langchain.com/oss/python/langgraph/errors/GRAPH_RECURSION_LIMITrecursion_limit caps the number of steps, and here the run stopped with an error after ten. Left at the agent's default it goes on for 9,999 steps before this error. With a hosted model, that is thousands of paid calls for one question.
A limit on model calls
ModelCallLimitMiddleware ends the run cleanly once the model has been called a set number of times.
limit = ModelCallLimitMiddleware(run_limit=3)
agent = create_agent(StuckModel(), tools=[lookup_order], middleware=[limit])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})
print(len(result["messages"]))
print(result["messages"][-1].text)8 Model call limits exceeded: run limit (3/3)
Three model calls, then the run ended with an AI message saying why. The 8 counts the messages: the question, three pairs of tool request and tool result, and the limit message at the end. run_limit counts calls in one invoke; thread_limit counts across a whole thread when there is a checkpointer. The default exit_behavior of "end" finishes with that message; "error" raises instead.
A limit on one tool
ToolCallLimitMiddleware caps one named tool rather than the model. Pair it with a model limit to be sure the run ends.
limits = [ToolCallLimitMiddleware(tool_name="lookup_order", run_limit=2),
ModelCallLimitMiddleware(run_limit=4)]
agent = create_agent(StuckModel(), tools=[lookup_order], middleware=limits)
for message in agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"]:
if message.type == "tool":
print(message.text)A17 shipped on 3 March. A17 shipped on 3 March. Tool call limit exceeded. Do not call 'lookup_order' again. Tool call limit exceeded. Do not call 'lookup_order' again.
Two lookups ran. After that the tool call limit answered each request itself, telling the model to stop, and the model call limit ended the run. A tool limit's default behaviour is "continue": block the tool, keep the agent going.
How each limit stopped the run
- Nothing counts on its own. Without a limit the loop runs to
recursion_limit, 9,999 steps by default, then errors. - A model limit ends the run. After three calls the run finished with an AI message naming the limit it hit.
- A tool limit blocks one tool. Two lookups ran, then each further call was answered with a message telling the model to stop.
- They pair well. The tool limit stops the tool, the model limit ends the run.
ModelCallLimitMiddleware vs ToolCallLimitMiddleware
| ModelCallLimitMiddleware | ToolCallLimitMiddleware | |
|---|---|---|
| Counts | Model calls | Calls to one named tool |
| Default on limit | Ends the run (exit_behavior='end') | Blocks the tool, keeps going ('continue') |
| Scope | run_limit per invoke, thread_limit per thread | run_limit per invoke, thread_limit per thread |
| Use for | Capping the total cost of a run | Stopping one tool from being overused |
Where call limits fit
- Capping the cost of a single request so a stuck loop cannot run up thousands of calls.
- Stopping one expensive tool from being called over and over.
- Ending a run with a clear message rather than a raised error.
Related
- Previous: Tool errors, caught and retried
- Next: Retries and a fallback model
- Reference: LangChain agent middleware
- Set
exit_behavior="error"on the model call limit and read the exception. - Give the tool limit no
tool_name, so it counts every tool. - Run the create_agent lesson's agent with a model call limit of 1 and ask about A17.
Slow is fine. Stopping is the only problem.