1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you plan disaster recovery and business continuity for AI features?
30-second answerSay your answer out loud first, then reveal.
Continuity planning table
| Component | Failure | Continuity measure |
|---|---|---|
| Model provider | Outage / account suspension / deprecation | Secondary provider or self-hosted fallback; tested prompts |
| GPU cluster | Region or capacity loss | Multi-zone; reserved capacity; API overflow |
| Vector index | Corruption / loss | Snapshots; rebuild pipeline with a known duration; previous version retained |
| Prompt / config registry | Unavailable | Cached last-known-good config in apps |
| Conversation store | Region outage | Replication; graceful "history unavailable" mode |
| Whole AI feature | Unusable | Manual process: human agents, standard forms, FAQ pages |
Key principles
- Critical business processes must not depend solely on AI without a manual fallback (e.g. claims can still be processed by adjusters).
- Test it: an untested failover plan doesn't work. Run game days.
- Know the rebuild times: "We can rebuild the 40M-chunk index in 9 hours" is a real RTO input.
- Vendor risk: contractual terms, exit plans, data export.
Related
You understood something today that you didn't yesterday.