Usage limits: stopping a run that loops
UsageLimits is an object you pass to a run to cap its requests, tool calls and tokens. When a cap is reached the run ends with UsageLimitExceeded instead of looping.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
A model can keep calling a tool without ever writing an answer. Nothing in the model stops that, so the cost and the wait grow with every call. A usage limit is the backstop.
A model that never stops asking
This model asks whether a refund has been paid, hears that it is pending, and asks again, forever.
from pydantic_ai import Agent, ModelResponse, ToolCallPart, UsageLimits
from pydantic_ai.models.function import FunctionModel
def impatient(messages, info):
return ModelResponse(parts=[ToolCallPart("check_refund", {"order_id": "A-1001"})])
agent = Agent(FunctionModel(impatient))
checks = 0@agent.tool_plain
def check_refund(order_id: str) -> str:
"""Check whether a refund has been paid."""
global checks
checks += 1
return "still pending"The built-in limit ends the loop
Run it with no limit of your own and count how many times the tool ran.
try:
agent.run_sync("Has my refund been paid?")
except UsageLimitExceeded as error:
print(error)
print("checks:", checks)The next request would exceed the request_limit of 50. Consider raising the limit, or see the docs on usage limits for budget-aware patterns: https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits checks: 50
Even with no limit set, a run stops itself before looping forever: a built-in request limit applies. The exact number can change between versions of Pydantic AI, so read it from the message rather than memorising it, and set your own for anything you ship. At a real model's prices and speed, this many calls for one ticket is already far too much.
Setting your own limit
Pass UsageLimits to the run to cap it lower. The request limit is checked before each request, so the one that would go over is never sent.
try:
agent.run_sync("Has my refund been paid?", usage_limits=UsageLimits(request_limit=4))
except UsageLimitExceeded as error:
print(error)
print("checks:", checks)The next request would exceed the request_limit of 4. Consider raising the limit, or see the docs on usage limits for budget-aware patterns: https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits checks: 4
What each limit counts
| Limit | Counts |
|---|---|
request_limit | Requests to the model. A built-in default applies when you set none. |
tool_calls_limit | Tool calls the agent runs. |
input_tokens_limit, output_tokens_limit | Tokens across the run. |
total_tokens_limit | Input and output tokens together. |
Token limits are checked against the counts a response reports, so the request that goes over has already been sent. count_tokens_before_request=True counts input tokens first, with an extra call to the provider. Pick a per-ticket limit from what a normal ticket uses, which result.usage tells you.
When you reach for a usage limit
- A tool loop that could run away, like polling a status that stays the same.
- A per-ticket budget, so one conversation cannot cost more than you allow.
- A guard in tests, so a stand-in that misbehaves fails fast instead of hanging.
request_limit and tool_calls_limit stop a call before it is sent, unless you turn on count_tokens_before_request.Related
- Previous: Tool errors: ModelRetry and crashes
- Next: Message history: continuing a conversation
- Reference: Usage limits
- Set
tool_calls_limit=2instead and read the message. - Make
check_refundreturn"paid"on the third call, and give the model a rule to answer with text then. - Print
result.usagefor a ticket from the tools lesson and choose a limit for it.
This is what real progress feels like.