Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q25IntermediateScenario

Traffic will spike 10x during a festival sale. Your AI shopping assistant must stay up. How do you prepare?

30-second answerSay your answer out loud first, then reveal.

Preparation plan

  1. Demand forecast: previous sale traffic × growth; convert to peak QPS and tokens per minute per model.
  2. Capacity:
    • Request API quota increases weeks in advance; consider provisioned or reserved throughput.
    • Self-hosted: reserve GPUs, pre-scale replicas (model loading is slow, so don't rely on reactive autoscaling).
  3. Reduce load per request:
    • Pre-generate answers for predictable questions (offer terms, return policy during sale) and cache them.
    • Route simple intents (order status, coupon FAQs) to rules or a small model.
    • Trim conversation history and retrieved context.
  4. Admission control: token-aware rate limits per user; queue with timeouts; priority for checkout-related help over browsing chat.
  5. Degradation ladder: full assistant → smaller model → FAQ-only mode → static help page, triggered automatically by latency and error thresholds.
  6. Load testing: realistic traffic replay at 10–12x, including dependent services (search, order APIs, vector DB).
  7. Operational readiness: a change freeze, dashboards (QPS, tokens/min vs quota, p95 latency, 429s, cost), alerts, runbooks, a war room.
  8. Cost guardrail: budget alerts, since a 10x spike in tokens is a 10x spike in spend.

Every expert started right here.