1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
The interviewer asks: "Your AI feature must respond in under 500ms at p95. How do you get there?" Walk through your approach.
30-second answerSay your answer out loud first, then reveal.
Step 1 — Clarify: first token vs complete response? What output length? Which percentile, and measured where (server vs client)?
Step 2 — Measure the breakdown (example before)
| Stage | p95 |
|---|---|
| Network + auth | 60ms |
| Retrieval (embed + search) | 120ms |
| LLM TTFT (large model, 3K-token prompt) | 900ms |
| Generation (150 tokens) | 2,000ms |
Step 3 — Levers
- Model: a small model (fine-tuned or distilled for the task) self-hosted near the app; quantized (FP8/INT4); speculative decoding.
- Prompt: cut tokens (shorter instructions, fewer chunks); prompt caching for the static prefix; avoid long histories.
- Output: constrain length (labels, short answers); structured outputs; generate only what's needed.
- Pipeline: remove sequential LLM calls (no separate rewrite step, or use a tiny model for it); run retrieval in parallel with other work; precompute embeddings for known queries.
- Caching: exact and semantic caches for frequent requests; precomputed answers for the head of the distribution.
- Infrastructure: warm GPU capacity, low queueing (headroom, admission control), persistent connections, same-region deployment.
- Tail latency: timeouts with fallback (cached or simpler answer), hedged requests to a second replica for slow ones, isolation from batch workloads.
- UX: stream the first token within 500ms and fill in the rest; optimistic UI; precompute when user intent is predictable (prefetch on hover or focus).
Step 4 — Verify: load test at peak, track p95/p99 continuously, and add latency regression checks to CI.
Related
- Previous: Q45. Design an AI system to process insurance claims with photos, forms and documents, with human adjusters in the loop.
- Next: Q47. Your company debates building its own LLM stack vs buying a vendor platform (or using a single model provider end to end). How do you guide the decision?
- LLM Fundamentals Interview Questions
Little by little, you're building something great.