Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q23IntermediateScenario

After a deploy, p95 latency doubled. How do you investigate?

30-second answerSay your answer out loud first, then reveal.

Investigation steps

  1. Mitigate first: if the SLO error budget is burning fast, roll back (flag or config), then debug.
  2. Diff the release bundle: what changed in code, prompt, model, retrieval or guardrails?
  3. Per-span comparison (before vs after):
ObservationLikely cause
Input tokens +60%Bigger prompt, more retrieved chunks, history handling bug
Output tokens +80%Prompt change made answers verbose; max_tokens raised
TTFT up, same tokensPrompt caching broken (dynamic content moved to the start); different model/region
Extra LLM spanNew step (rewrite, judge) added sequentially
Retry spansJSON validation failures with the new prompt
Tool spans slowerNew tool or changed API usage

4. Check external factors: provider status, traffic mix shift, a canary on different infrastructure.

5. Fix and verify with a load test or canary; add a latency regression check (tokens and p95) to the CI gates.

Slow is fine. Stopping is the only problem.