Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q46HardScenario

The interviewer asks: "Your AI feature must respond in under 500ms at p95. How do you get there?" Walk through your approach.

30-second answerSay your answer out loud first, then reveal.

Step 1 — Clarify: first token vs complete response? What output length? Which percentile, and measured where (server vs client)?

Step 2 — Measure the breakdown (example before)

Stagep95
Network + auth60ms
Retrieval (embed + search)120ms
LLM TTFT (large model, 3K-token prompt)900ms
Generation (150 tokens)2,000ms

Step 3 — Levers

  1. Model: a small model (fine-tuned or distilled for the task) self-hosted near the app; quantized (FP8/INT4); speculative decoding.
  2. Prompt: cut tokens (shorter instructions, fewer chunks); prompt caching for the static prefix; avoid long histories.
  3. Output: constrain length (labels, short answers); structured outputs; generate only what's needed.
  4. Pipeline: remove sequential LLM calls (no separate rewrite step, or use a tiny model for it); run retrieval in parallel with other work; precompute embeddings for known queries.
  5. Caching: exact and semantic caches for frequent requests; precomputed answers for the head of the distribution.
  6. Infrastructure: warm GPU capacity, low queueing (headroom, admission control), persistent connections, same-region deployment.
  7. Tail latency: timeouts with fallback (cached or simpler answer), hedged requests to a second replica for slow ones, isolation from batch workloads.
  8. UX: stream the first token within 500ms and fill in the rest; optimistic UI; precompute when user intent is predictable (prefetch on hover or focus).

Step 4 — Verify: load test at peak, track p95/p99 continuously, and add latency regression checks to CI.

Little by little, you're building something great.