LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Call limits with ModelCallLimitMiddleware

A call limit is a middleware setting that ends an agent run after a set number of model or tool calls, so a stuck agent costs a few calls instead of thousands.

Last updated: 27 Sep, 2026 · LangChain 1.4

Nothing in the agent loop counts by itself. As long as the model keeps asking for tools, it keeps going. A real model stops once it has its answer, so a stuck loop cannot be produced on cue; this lesson uses a stand-in model that asks for the same lookup every time.

ModelCallLimitMiddleware and ToolCallLimitMiddleware

python
from langchain.agents.middleware import ModelCallLimitMiddleware, ToolCallLimitMiddleware

# cap model calls in one run (or across a thread with thread_limit)
limit = ModelCallLimitMiddleware(run_limit=3)
# cap how often one named tool may be called
tool_limit = ToolCallLimitMiddleware(tool_name="lookup_order", run_limit=2)

A model stuck in a loop

The agent here runs on ShopModel, the stand-in chat model built in Several tool calls at once. Start the file with it.

python
import re

from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult


class ShopModel(BaseChatModel):
    tools: list = []

    @property
    def _llm_type(self):
        return "shop"

    def bind_tools(self, tools, **kwargs):
        return self.model_copy(update={"tools": tools})   # a copy holding the tools

    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        message = self.decide(messages)                   # the reply comes from decide
        return ChatResult(generations=[ChatGeneration(message=message)])

    def decide(self, messages):
        results = []                                # the tool results at the end
        for m in reversed(messages):
            if not isinstance(m, ToolMessage):
                break
            results.insert(0, m.text)
        if results:                                 # results are back: answer with them
            return AIMessage(" ".join(results))
        text = messages[-1].text
        orders = re.findall(r"\b[A-Z]\d+\b", text)
        tool = "refund_order" if "refund" in text.lower() else "lookup_order"
        if orders and tool in [t.name for t in self.tools]:   # one call per order id
            calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
                     for o in orders]
            return AIMessage("", tool_calls=calls)
        if orders:                                  # that tool is not bound
            return AIMessage(f"I have no way to look up {orders[0]} yet.")
        return AIMessage("Hello. Which order is this about?")

StuckModel extends ShopModel and overrides its decide method to always ask for the same tool, so it never stops on its own.

python
from langchain.messages import AIMessage

class StuckModel(ShopModel):
    def decide(self, messages):
        # always ask for the same lookup, so the loop never ends
        call = {"name": "lookup_order", "args": {"order_id": "A17"}, "id": f"call_{len(messages)}"}
        return AIMessage("", tool_calls=[call])

The tool and the agent

This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Add it below the models.

python
from langchain.tools import tool

ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}


@tool
def lookup_order(order_id: str) -> str:
    """Look up an order's shipping status by its id, such as A17."""
    status = ORDERS.get(order_id)
    return f"{order_id} {status}." if status else f"{order_id} is not an order we have."

Without a limit

With no limit, only LangGraph's recursion_limit stops it. Set it low to see the run end with an error.

Example
from langchain.agents import create_agent

agent = create_agent(StuckModel(), tools=[lookup_order])
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]}, {"recursion_limit": 10})

recursion_limit caps the number of steps, and here the run stopped with an error after ten. Left at the agent's default it goes on for 9,999 steps before this error. With a hosted model, that is thousands of paid calls for one question.

A limit on model calls

ModelCallLimitMiddleware ends the run cleanly once the model has been called a set number of times.

Example
limit = ModelCallLimitMiddleware(run_limit=3)
agent = create_agent(StuckModel(), tools=[lookup_order], middleware=[limit])
result = agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})

print(len(result["messages"]))
print(result["messages"][-1].text)

Three model calls, then the run ended with an AI message saying why. The 8 counts the messages: the question, three pairs of tool request and tool result, and the limit message at the end. run_limit counts calls in one invoke; thread_limit counts across a whole thread when there is a checkpointer. The default exit_behavior of "end" finishes with that message; "error" raises instead.

A limit on one tool

ToolCallLimitMiddleware caps one named tool rather than the model. Pair it with a model limit to be sure the run ends.

Example
limits = [ToolCallLimitMiddleware(tool_name="lookup_order", run_limit=2),
          ModelCallLimitMiddleware(run_limit=4)]
agent = create_agent(StuckModel(), tools=[lookup_order], middleware=limits)

for message in agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]})["messages"]:
    if message.type == "tool":
        print(message.text)

Two lookups ran. After that the tool call limit answered each request itself, telling the model to stop, and the model call limit ended the run. A tool limit's default behaviour is "continue": block the tool, keep the agent going.

How each limit stopped the run

  • Nothing counts on its own. Without a limit the loop runs to recursion_limit, 9,999 steps by default, then errors.
  • A model limit ends the run. After three calls the run finished with an AI message naming the limit it hit.
  • A tool limit blocks one tool. Two lookups ran, then each further call was answered with a message telling the model to stop.
  • They pair well. The tool limit stops the tool, the model limit ends the run.

ModelCallLimitMiddleware vs ToolCallLimitMiddleware

ModelCallLimitMiddlewareToolCallLimitMiddleware
CountsModel callsCalls to one named tool
Default on limitEnds the run (exit_behavior='end')Blocks the tool, keeps going ('continue')
Scoperun_limit per invoke, thread_limit per threadrun_limit per invoke, thread_limit per thread
Use forCapping the total cost of a runStopping one tool from being overused

Where call limits fit

  • Capping the cost of a single request so a stuck loop cannot run up thousands of calls.
  • Stopping one expensive tool from being called over and over.
  • Ending a run with a clear message rather than a raised error.
Watch out. A tool limit alone will not end the run; its default is to block the tool and let the agent continue. Add a model call limit so the loop is brought to a stop.
Try it yourself
  • Set exit_behavior="error" on the model call limit and read the exception.
  • Give the tool limit no tool_name, so it counts every tool.
  • Run the create_agent lesson's agent with a model call limit of 1 and ask about A17.

Slow is fine. Stopping is the only problem.