Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q42HardConcept

How do you plan disaster recovery and business continuity for AI features?

30-second answerSay your answer out loud first, then reveal.

Continuity planning table

ComponentFailureContinuity measure
Model providerOutage / account suspension / deprecationSecondary provider or self-hosted fallback; tested prompts
GPU clusterRegion or capacity lossMulti-zone; reserved capacity; API overflow
Vector indexCorruption / lossSnapshots; rebuild pipeline with a known duration; previous version retained
Prompt / config registryUnavailableCached last-known-good config in apps
Conversation storeRegion outageReplication; graceful "history unavailable" mode
Whole AI featureUnusableManual process: human agents, standard forms, FAQ pages

Key principles

  • Critical business processes must not depend solely on AI without a manual fallback (e.g. claims can still be processed by adjusters).
  • Test it: an untested failover plan doesn't work. Run game days.
  • Know the rebuild times: "We can rebuild the 40M-chunk index in 9 hours" is a real RTO input.
  • Vendor risk: contractual terms, exit plans, data export.

You understood something today that you didn't yesterday.