Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q24IntermediateScenario

Your primary LLM provider has a major outage during peak hours. How should the system have been designed to handle it?

30-second answerSay your answer out loud first, then reveal.
Provider failover: the app calls a gateway that checks primary health by error rate and latency; if healthy it uses the primary provider in region A, otherwise the circuit opens to the same model in region B, then a secondary provider with model-specific prompts, then a degraded mode of a self-hosted small model, cached answers and a queue, with a user notice of limited functionality.

Design elements

  1. Multi-region / multi-provider: the same model may be available via several clouds or regions; a second provider guards against vendor-wide outages.
  2. Per-model prompt variants: prompts tuned for each fallback model, with evals run regularly (a nightly CI eval on fallbacks).
  3. Circuit breakers: trip on error rate or latency thresholds; half-open probing to recover.
  4. Graceful degradation tiers:
    • Critical flows (customer support) → fallback model.
    • Non-critical (summaries, suggestions) → disable temporarily or queue.
    • Async jobs → pause and resume later.
  5. Idempotency + checkpoints so in-flight agent tasks resume rather than restart.
  6. Observability and runbooks: alerts, dashboards per provider, an on-call runbook, status page communication.
  7. Chaos testing: periodically force failover in staging and production game days.

Trade-offs: multi-provider means more integration work, inconsistent behaviour between models, and data agreements with multiple vendors.

This is what real progress feels like.