1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your company wants to move a high-volume feature from a commercial LLM API to a self-hosted open model. Plan the migration.
30-second answerSay your answer out loud first, then reveal.
Phases
| Phase | Work | Exit criteria |
|---|---|---|
| 1. Business case | Volume, cost model (GPU TCO at realistic utilisation incl. engineering), data/compliance drivers | Clear savings or a strategic reason |
| 2. Model selection | Eval 2–4 open models on the golden set; consider fine-tuning or distillation from API outputs (check terms) | Within agreed quality tolerance |
| 3. Platform | Inference cluster, gateway integration, observability, autoscaling, load tests | SLOs met under 1.5x peak |
| 4. Shadow | Mirror 10–100% of traffic; judge outputs side by side | Quality parity confirmed on real traffic |
| 5. Canary | 5% → 25% → 50% → 100%, API fallback on errors/latency | Stable SLOs, quality, cost |
| 6. Operate | On-call, runbooks, upgrade process, capacity reviews | Smooth for 1–2 months before reducing API commitments |
Risks and mitigations
- Quality gaps on edge cases: route hard cases to the API (hybrid routing).
- Underestimated ops cost: include on-call, GPU procurement, upgrades.
- Utilisation: peaks may require API overflow rather than idle GPUs.
- Model lifecycle: you now own upgrades and security patches.
Interview signal. Treating it as a business decision plus a careful, reversible engineering migration.
Related
Every expert started right here.