1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How should an agent handle tool failures and retries?
30-second answerSay your answer out loud first, then reveal.
Error taxonomy
| Type | Example | Who handles it | How |
|---|---|---|---|
| Transient | 429, timeout, 503 | Code | Retry w/ exponential backoff + jitter, cap attempts |
| Model-fixable | Invalid date format, unknown enum value | Model | Return descriptive error and let it retry |
| Business rule | "Refund exceeds policy limit" | Model, then human | Explain the rule; maybe escalate |
| Permanent / auth | 403, account deleted | Stop | Report to user / escalate, don't loop |
| Unknown side effect | Timeout on a payment call | Careful! | Check status before retrying (idempotency, Q32) |
Good practices
- Circuit breakers: if a dependency is down, stop calling it for a while and tell the agent the tool is unavailable so it can plan around it.
- Fallback tools: e.g. a secondary search provider.
- Error budgets in the loop: after N tool errors in a run, escalate.
- Log everything: error type, args and retries in the trace for later analysis.
Common mistakes
- Sending a full stack trace to the model. It wastes tokens and can leak internals.
- Retrying a "send email" or "charge card" after a timeout and causing a duplicate.
Related
Every expert started right here.