Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q25IntermediateScenario

Users see intermittent timeouts because the LLM provider's tail latency is high. How do you make the system robust?

30-second answerSay your answer out loud first, then reveal.

Techniques

  1. Timeouts: TTFT timeout (e.g. 5s) + total timeout scaled to max_tokens; don't use one fixed 30s for everything.
  2. Retries: only for timeouts, 429s and 5xx; exponential backoff + jitter; cap the attempts and the total time budget.
  3. Hedging: for short, idempotent calls (classification, routing), fire a backup request after the p95 latency elapses; cancel the loser. It increases cost slightly and cuts tail latency significantly.
  4. Health-aware routing: track latency and error rates per region and provider; shift traffic away from degraded endpoints automatically.
  5. Circuit breakers: stop hammering a failing endpoint; fail over or degrade.
  6. Graceful degradation: a smaller or faster model, cached answers, "We're taking longer than usual" messaging, async completion with notification.
  7. Observability: p99 latency per provider, region and model; timeout and retry rates; alerting.

Caution: retries and hedges multiply load during incidents. Use retry budgets to avoid amplifying an outage.

Every expert started right here.