1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your primary LLM provider has a major outage during peak hours. How should the system have been designed to handle it?
30-second answerSay your answer out loud first, then reveal.

Design elements
- Multi-region / multi-provider: the same model may be available via several clouds or regions; a second provider guards against vendor-wide outages.
- Per-model prompt variants: prompts tuned for each fallback model, with evals run regularly (a nightly CI eval on fallbacks).
- Circuit breakers: trip on error rate or latency thresholds; half-open probing to recover.
- Graceful degradation tiers:
• Critical flows (customer support) → fallback model.
• Non-critical (summaries, suggestions) → disable temporarily or queue.
• Async jobs → pause and resume later. - Idempotency + checkpoints so in-flight agent tasks resume rather than restart.
- Observability and runbooks: alerts, dashboards per provider, an on-call runbook, status page communication.
- Chaos testing: periodically force failover in staging and production game days.
Trade-offs: multi-provider means more integration work, inconsistent behaviour between models, and data agreements with multiple vendors.
Related
This is what real progress feels like.