1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Users see intermittent timeouts because the LLM provider's tail latency is high. How do you make the system robust?
30-second answerSay your answer out loud first, then reveal.
Techniques
- Timeouts: TTFT timeout (e.g. 5s) + total timeout scaled to max_tokens; don't use one fixed 30s for everything.
- Retries: only for timeouts, 429s and 5xx; exponential backoff + jitter; cap the attempts and the total time budget.
- Hedging: for short, idempotent calls (classification, routing), fire a backup request after the p95 latency elapses; cancel the loser. It increases cost slightly and cuts tail latency significantly.
- Health-aware routing: track latency and error rates per region and provider; shift traffic away from degraded endpoints automatically.
- Circuit breakers: stop hammering a failing endpoint; fail over or degrade.
- Graceful degradation: a smaller or faster model, cached answers, "We're taking longer than usual" messaging, async completion with notification.
- Observability: p99 latency per provider, region and model; timeout and retry rates; alerting.
Caution: retries and hedges multiply load during incidents. Use retry budgets to avoid amplifying an outage.
Related
Every expert started right here.