Pydantic AIPydantic AI 2.51 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
29 small wins to finish your pathNext lesson →

Usage limits: stopping a run that loops

UsageLimits is an object you pass to a run to cap its requests, tool calls and tokens. When a cap is reached the run ends with UsageLimitExceeded instead of looping.

Last updated: 28 Sep, 2026 · Pydantic AI 2.51

A model can keep calling a tool without ever writing an answer. Nothing in the model stops that, so the cost and the wait grow with every call. A usage limit is the backstop.

A model that never stops asking

This model asks whether a refund has been paid, hears that it is pending, and asks again, forever.

python
from pydantic_ai import Agent, ModelResponse, ToolCallPart, UsageLimits
from pydantic_ai.models.function import FunctionModel


def impatient(messages, info):
    return ModelResponse(parts=[ToolCallPart("check_refund", {"order_id": "A-1001"})])


agent = Agent(FunctionModel(impatient))
checks = 0
python
@agent.tool_plain
def check_refund(order_id: str) -> str:
    """Check whether a refund has been paid."""
    global checks
    checks += 1
    return "still pending"

The built-in limit ends the loop

Run it with no limit of your own and count how many times the tool ran.

Example
try:
    agent.run_sync("Has my refund been paid?")
except UsageLimitExceeded as error:
    print(error)
print("checks:", checks)

Even with no limit set, a run stops itself before looping forever: a built-in request limit applies. The exact number can change between versions of Pydantic AI, so read it from the message rather than memorising it, and set your own for anything you ship. At a real model's prices and speed, this many calls for one ticket is already far too much.

Setting your own limit

Pass UsageLimits to the run to cap it lower. The request limit is checked before each request, so the one that would go over is never sent.

Example
try:
    agent.run_sync("Has my refund been paid?", usage_limits=UsageLimits(request_limit=4))
except UsageLimitExceeded as error:
    print(error)
print("checks:", checks)

What each limit counts

LimitCounts
request_limitRequests to the model. A built-in default applies when you set none.
tool_calls_limitTool calls the agent runs.
input_tokens_limit, output_tokens_limitTokens across the run.
total_tokens_limitInput and output tokens together.

Token limits are checked against the counts a response reports, so the request that goes over has already been sent. count_tokens_before_request=True counts input tokens first, with an extra call to the provider. Pick a per-ticket limit from what a normal ticket uses, which result.usage tells you.

When you reach for a usage limit

  • A tool loop that could run away, like polling a status that stays the same.
  • A per-ticket budget, so one conversation cannot cost more than you allow.
  • A guard in tests, so a stand-in that misbehaves fails fast instead of hanging.
Watch out. The token limits are checked after a response comes back, so the request that crosses the line has already been paid for. Only request_limit and tool_calls_limit stop a call before it is sent, unless you turn on count_tokens_before_request.
Try it yourself
  • Set tool_calls_limit=2 instead and read the message.
  • Make check_refund return "paid" on the third call, and give the model a rule to answer with text then.
  • Print result.usage for a ticket from the tools lesson and choose a limit for it.

This is what real progress feels like.